Skip to content
All articles

Realtime Voice AI

A voice agent is a stack, not a model

A demo voice bot is speech-to-text, a language model and text-to-speech. A production voice agent adds media, state, tools, people and controls, and keeps media apart from business actions.

· 4 min read

Voice agent stack: phone or app, media edge (WebRTC, SIP, codecs, jitter), speech layer (voice activity, ASR, TTS, speech-to-speech) and orchestrator (turns, state, prompts, policy), which calls tools and RAG or hands off to a human agent. Around every turn: identity, consent, authorization, session state, tracing, evaluation and retention. Design rule: a reconnect must never repeat a payment.

The first voice demo most teams build is three calls in a row: speech-to-text, a language model, text-to-speech. It works on a laptop. A production voice agent answering real calls is a stack, and the model is one dependency inside it.

The layers a call passes through

  • The client or phone captures audio and plays the reply, sometimes with its own processing such as echo cancellation.
  • A media edge terminates WebRTC or telephony (SIP), handles codecs, jitter and session signaling, and forwards normalized audio frames.
  • The speech layer detects when someone is speaking, recognizes words, synthesizes the reply, or does both through a speech-to-speech model.
  • The orchestrator owns the conversation: it keeps state, applies prompts and policies, retrieves knowledge, runs tools, and decides when to hand off.
  • A human agent receives a bounded summary and verified business context when the call transfers, not an unreviewed transcript of everything the model considered.

Around all of it sit identity, consent, authorization, policy and audit. Observability has to join the media sequence, transcript revisions, model events, tool calls, generated audio and the reason each session ended, or nobody can explain a bad call afterwards.

Keep media apart from business actions

This is the rule that separates a voice product from a voice demo. Audio is fragile and gets retried; business actions must happen once. So the two are controlled separately:

  • A dropped connection can reconnect the audio without repeating a payment the agent already made.
  • A slow or failed tool call must not corrupt the media session.
  • When a caller interrupts, the agent cancels the audio it was generating and the model work in progress, and keeps the actions it has already committed.

Think Like an Engineer

A caller says "yes, pay it", the payment tool succeeds, and the caller interrupts while the agent is still saying "Your payment of…". What gets cancelled? The rest of that sentence and any model work queued behind it. What survives? The payment, and the record that it happened, so the next thing the agent says is consistent with it.

Give every event a name

Interruptions and reconnects produce stale events: audio from a response the caller already talked over, a tool result arriving after the turn moved on. Explicit identifiers for the session, the turn, the response, each generated piece of audio, each tool call and each handoff let the orchestrator recognize events that belong to the past and drop them. Without those identifiers, a late event can quietly overwrite the present.

The session store keeps only the state that must survive; high-rate media buffers stay ephemeral.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy