How it works, honestly

AI voice agents:
what happens in those two seconds

Every vendor in New Zealand will tell you their voice agent sounds natural. Fewer will tell you what it is doing between the caller finishing a sentence and the reply starting, or what makes it fall over. Here is both.

Four things, in a loop, every turn

A voice agent is not one model. It is a pipeline, and it runs end to end every time the caller stops talking.

Hears

Speech to text, streaming

The caller’s audio is transcribed continuously as they speak, not after they stop. Streaming is what makes the reply feel immediate rather than walkie-talkie.

Decides

Reasoning against your business

A language model works out what the caller wants, checked against your services, pricing, hours and service area — not general internet knowledge.

Acts

Function calling

When it needs to do something real — check availability, create a job, look up an address — it calls a function against your systems mid-conversation and uses the answer.

Speaks

Text to speech

The reply is spoken in a consistent voice. Generation starts before the full sentence is ready, which is most of the difference between natural and stilted.

The three things that actually decide if it works

1. Latency, and where it hides

Every stage above adds delay, and they add up. Past roughly a second of silence a caller assumes the line has dropped and starts talking again. The fix is not one fast model, it is streaming at every stage so the reply begins before the whole thought is finished.

2. Interruptions and turn-taking

Real callers interrupt, trail off, and pause mid-sentence to read a job number off a docket. An agent that treats any silence as its turn will talk over people constantly. Getting this right is unglamorous tuning work and it is the single most common reason a demo feels great and the live line does not.

3. Local names, accents and addresses

Generic speech recognition mangles local suburb and street names, and a booking with a wrong address is worse than a missed call. AnswR feeds the agent your actual service area and, in some regions, a street-level gazetteer, then has it read the address back to the caller to confirm. Confirming beats guessing every time.

What we will not claim

No voice agent is indistinguishable from a person for the length of a whole call, and any vendor who says otherwise is describing a demo. Some callers will work out it is AI. In practice most do not mind, because the alternative they were about to get was voicemail.

What matters is that it is polite, it gets the details right, and it hands over to a human when it is out of its depth.

Judge it on your own line

Demos are easy. Put it on real calls for seven days and listen to the recordings.

AnswR Support

Powered by AI - replies instantly

Chat with our AI assistant

Get instant answers about AnswR

By chatting, you agree to our privacy policy.