How voice AI understands a phone call: speech recognition, turn-taking and memory

What happens inside a voice AI agent on a phone call: how it turns speech into text, decides when you have finished talking, keeps track of the conversation and remembers returning callers — and why each part goes wrong.

6 min read Updated 13 September 2026 Sonorch knowledge base
On this page

Talking to a voice AI feels like one thing: you speak, it answers. Underneath, it is several separate systems handing work to one another in fractions of a second, and each has its own way of getting a phone call wrong.

This guide walks through those systems in the order a call uses them. It is written from building one: the examples of what goes wrong are real failures found on test calls to Sonorch’s own receptionist, and since fixed.

Key takeaways

  • A voice agent is a pipeline — speech recognition, a language model with tools, and speech synthesis — plus the timing logic that decides when to talk.
  • Phone audio is narrow and noisy. Recognition improves sharply when it is told the business’s own words in advance.
  • The hardest problem is not understanding words. It is knowing whether a pause means the caller has finished.
  • Memory should come only from what the caller said, and be treated as possibly out of date.

The pipeline behind a voice agent

A voice agent on a phone call is usually built from three models and the logic that joins them:

  1. Speech recognition (speech-to-text)

    Streams the caller’s audio into words, revising its guess as more sound arrives.

  2. A language model with tools

    Reads the conversation so far and decides what to say or do. When it needs facts — openings, prices, the caller’s bookings — it calls tools connected to the business’s systems instead of guessing.

  3. Speech synthesis (text-to-speech)

    Turns the reply into a voice, streamed so the first words play before the whole sentence is ready.

Around those sits the part callers notice most without knowing it exists: turn-taking, the logic that decides when the caller has finished, and when the agent should stop because the caller has started again.

Hearing: speech recognition on a phone line

A phone call carries much less of the voice than a recording made on the same phone: calls are squeezed into a narrow band of frequencies, and they arrive with road noise, kitchen noise and whoever else is in the room. Recognition loses accuracy fastest on exactly the words a business cares about — names, services and times.

Three things make the difference in practice:

  • Vocabulary. Giving the recogniser the business’s service names, staff names and scheduling words in advance. On one of our own test calls, a caller asking for “the pedicure one” was transcribed as “The Pedicure 1”.
  • Times, the way people say them. “Two”, “two thirty”, “after lunch”. A misheard AM or PM books the right time on the wrong half of the day, so a good system listens hard for those words and reads the time back.
  • Other voices. A phone in a busy room hears everyone. Suppressing background voices and separating speakers helps the agent answer the caller rather than the television.

Knowing when you have finished

People pause mid-sentence all the time: to think, to check a diary, to breathe. A person on the phone uses grammar, tone and context to tell “I’d like Saturday…” from “I’d like Saturday.” Software has to judge it, and both mistakes are costly:

  • Answering too early cuts the caller off, and the agent replies to half a request.
  • Waiting too long leaves dead air, and the caller wonders whether the line dropped.

Modern systems combine two signals. Voice-activity detection hears that the sound has stopped; a turn-detection model reads the words so far and judges whether the thought is complete. The agent speaks when both agree.

Interruptions are the reverse problem. When a caller says “actually, wait” over the agent, it should stop and listen. On one of our test calls, “actually, wait” was heard as “Ashley Way” and taken as the caller’s name. Failures like that only show up on real calls, which is why every change to a voice agent has to be tested on them.

Deciding: the language model and its tools

The language model is the part that sounds intelligent, and the part that most needs a short leash. Left alone, a model will happily produce a plausible appointment time. Connected properly, it can only offer times the scheduling system returned, and only confirm a booking the system accepted.

  • Tools, not memory, for facts. Availability, prices and bookings come from live lookups on every call.
  • Confirmation before action. The agent reads back the service, the day and the time before it books.
  • Vague requests made specific. “Afternoon” should become a spread of real afternoon openings, not the first slot after noon. How many appointment times should you offer a caller? explains why.

Remembering: caller recognition and memory

Recognising a returning caller starts with the phone number. Matched against the business’s customer records, it gives the agent a name, any upcoming appointments and the services that caller usually books — enough to ask “Are you calling about Thursday?” without being told.

Memory goes a step further: short facts kept from past calls, like a preference for mornings. Two rules keep it trustworthy:

  • Only from the caller. Facts should come from what the caller said, never from what the agent said — otherwise the agent ends up remembering its own mistakes.
  • Treated as possibly stale. A preference from March may not hold in September. Memory is a hint to confirm, not a fact to act on.

Speaking: voice, pace and speed of reply

Synthetic voices are now good enough that the voice itself is rarely the problem. Two settings still matter: a voice that suits the business, and a pace a caller can follow on speakerphone. Slower is better for confirmations — the moment the caller checks the day and time.

How quickly the agent replies matters more than how it sounds. Every stage above adds a little delay, and a gap of more than a second or so starts to feel like a bad line. That is why each stage streams its work to the next instead of waiting for a finished sentence.

Questions

Voice AI is software that understands spoken language and replies in speech. On a phone call it combines speech recognition, a language model that decides what to say or do, and speech synthesis, plus timing logic that decides when to speak.