> Realtime speech-to-speech or a cascaded pipeline? Latency budget per stage, turn-taking, tool calls, SIP, German and Hungarian quality, and AI Act disclosure.
>
> Web page: https://balazscsorba.com/blog/voice-agents-realtime-latency · Language: English · Also available in: [Deutsch](https://balazscsorba.com/de/blog/voice-agents-realtime-latency.md) · [Magyar](https://balazscsorba.com/hu/blog/voice-agents-realtime-latency.md)
> Author: Balázs Csorba · Published: 2026-10-02 · Keywords: voice agents, speech-to-speech vs STT LLM TTS, OpenAI Realtime API, Gemini Live API, voice agent latency budget, turn detection and barge-in, Pipecat vs LiveKit, AI voice agent SIP telephony, German Hungarian voice AI, AI Act Article 50 voice bot disclosure

[Blog](https://balazscsorba.com/blog)/AI agents

# Building voice agents: realtime speech-to-speech or STT, LLM and TTS?

Realtime speech-to-speech or a cascaded pipeline? Latency budget per stage, turn-taking, tool calls, SIP, German and Hungarian quality, and AI Act disclosure.

[Balázs Csorba](https://balazscsorba.com/about)·October 2, 2026·13 min read

-   Voice agents
-   Realtime API
-   Latency
-   Telephony
-   AI Act

![Diagram: a caller reaches a voice agent over SIP or WebRTC, which fans out to turn detection, speech recognition, an LLM with tools and speech synthesis.](https://balazscsorba.com/images/blog/voice-agents-realtime-latency/cover.webp?v=66ccbf6239)

## Key takeaways

-   A cascaded pipeline has a realistic best case of about 725 ms to the first spoken word and typically lands at 1.1 to 2.1 seconds, with the LLM first token as the largest slice.
-   Speech-to-speech models remove two hops and handle interruptions natively, but you lose the text between the stages: no transcript checks before the answer, and fewer components to swap.
-   Turn-taking decides how a voice agent feels more than raw speed does. Use semantic end-of-turn detection, and check that it supports your languages: LiveKit, Pipecat Smart Turn and Deepgram Flux cover German but not Hungarian.
-   Tool calls are where voice agents break: a slow API means silence. Use asynchronous tools where the model supports it, and cover every wait over a second with a short spoken filler.
-   Since 2 August 2026, Article 50 of the EU AI Act requires that people are told they are talking to an AI unless it is obvious. Say it in the first sentence of every call.

On this page

1.  [Speech-to-speech or cascaded pipeline?](https://balazscsorba.com/#speech-to-speech-vs-cascaded)
2.  [The latency budget, stage by stage](https://balazscsorba.com/#latency-budget)
3.  [Turn-taking and barge-in](https://balazscsorba.com/#turn-taking-barge-in)
4.  [Tool calls while the agent is talking](https://balazscsorba.com/#tool-calls)
5.  [Telephony: SIP, WebRTC and the phone network](https://balazscsorba.com/#telephony)
6.  [German and Hungarian: test the whole chain](https://balazscsorba.com/#german-hungarian)
7.  [Consent, disclosure and the AI Act](https://balazscsorba.com/#consent-ai-act)
8.  [What I would build first](https://balazscsorba.com/#what-i-would-build)
9.  [The bigger picture](https://balazscsorba.com/#the-bigger-picture)
10.  [Sources](https://balazscsorba.com/#sources)

Voice is the interface where an AI agent gets judged in two seconds. In a chat window, a slow answer is an inconvenience; on a phone call, a pause of one second feels like a dropped line and callers start talking over the agent or hang up. That is why the architecture questions look different from a text agent: you are spending a time budget, not a token budget.

The first decision is between two shapes. A **speech-to-speech** model, such as the OpenAI Realtime API or the Gemini Live API, consumes audio and produces audio inside one model. A **cascaded pipeline** chains speech-to-text, a text LLM and text-to-speech, and orchestration frameworks such as Pipecat and LiveKit Agents wire the pieces together. Both work in production, and I use the choice to decide where I need control.

This article goes through the decision with numbers from the vendors' documentation: the latency budget per stage, turn-taking and barge-in, tool calls while the agent is talking, telephony, German and Hungarian quality, and what the EU AI Act asks of you. It ends with the checklist I would start from. Prices and model names in this field change every few months, so treat the details as a snapshot of 2 October 2026.

## Speech-to-speech or cascaded pipeline?

OpenAI's own [voice agents guide](https://developers.openai.com/api/docs/guides/voice-agents) frames the choice cleanly. The Realtime API is for "speech, reasoning, and tool use in one session", with one model interpreting audio, deciding what to do and answering in speech. The chained path is for when you need "control over each speech and text stage": you store the transcript, run policy checks before the text agent responds, and replace each component independently. The guide does not publish latency numbers for either, only advice to compare the median and the 95th percentile on similar calls.

That framing matches what I see in projects. Speech-to-speech wins when the conversation itself is the product and the logic is light. The cascaded pipeline wins when the answer must be checked, logged or grounded before it is spoken, which is most B2B work: order status, appointment changes, support triage with a knowledge base. Here is how the two compare:

Aspect

Speech-to-speech (realtime model)

Cascaded (STT, LLM, TTS)

Hops per turn

One model, one session

Three services plus orchestration

Latency floor

Lower: no transcription or synthesis hop; the practitioner figure is about 800 ms voice to voice including network

Higher: roughly 725 ms best case, typically 1.1 to 2.1 s

Interruptions

Handled inside the session; you truncate what was not played

You own the logic: cancel TTS, flush audio, update the history

Intermediate text

Transcripts are a side output, not a gate

Text between every stage: store, redact, run guardrails

Swapping parts

You take the model as it is, including its voices

Choose the best STT, LLM and voice for each language

Language coverage

Check per language; I could not confirm Hungarian in the documentation

Mix and match: Nova-3 and Flash v2.5 both list Hungarian

Cost control

Audio tokens: the OpenAI model page lists $32 per million input, $64 per million output and $0.40 cached

Per-minute STT and TTS plus text tokens; cheaper models for easy turns

Debugging

Harder: reasoning and speech are fused

Easier: every stage has a log line and a timing

Two details from the documentation are worth knowing. The `gpt-realtime` model page lists a 32,000-token context window and a maximum of 4,096 output tokens, and it supports WebRTC, WebSocket and SIP. Google's Live API is a stateful WebSocket connection, and its guide caps audio-only sessions at 15 minutes and audio with video at 2 minutes unless you use its session management techniques. Both are session-shaped, not request-shaped, which changes how you handle reconnects and long calls.

**My default**

For a business process with tools and compliance needs, I start with a cascaded pipeline in Pipecat or LiveKit, because I can inspect the text between the stages and swap each part. I reach for a realtime model when the call is open-ended and the experience matters more than the audit trail, and I keep the framework in place so I can switch.

## The latency budget, stage by stage

A cascaded turn is a relay race, and the caller hears only the total. One practitioner [playbook](https://www.forasoft.com/blog/article/voice-ai-agents-livekit-guide) puts numbers on every leg, from a best case to what deployments typically show. I treat these as planning figures, not guarantees, but the shape holds across the sources I read:

The LLM first token is the biggest slice in both rows. Cutting it, through a smaller model, prompt caching or starting generation early, pays off more than shaving the other stages combined.

Three conclusions follow. First, **the LLM is the budget**. Its first token takes 400 ms in the best case and 600 to 1,200 ms typically, so model choice, prompt size and caching matter more than any audio tweak. I covered the levers in [LLM cost and latency: prompt caching and routing](https://balazscsorba.com/blog/llm-cost-latency-prompt-caching-routing). Route simple turns to a small, fast model and keep the strong model for the hard ones.

Second, **start early and stream everything**. Deepgram's Flux speech-to-text model detects the end of a turn itself, in about 260 ms, and emits an `EagerEndOfTurn` event that lets you start the LLM before the user has definitely finished, saving hundreds of milliseconds. The price is wasted calls when the user continues, so you must be able to cancel a speculative generation cleanly. Streaming TTS has the same logic: start speaking at the first sentence, not the last.

Third, **the phone network costs you extra**. The same playbook adds 80 to 150 ms for the public telephone network on top of WebRTC. And measure the way OpenAI suggests: median and 95th percentile of the time a caller waits for a useful spoken answer, on real calls, because the tail is what people remember.

## Turn-taking and barge-in

The hardest part of a voice agent is not speaking but knowing when to speak. Simple voice activity detection ends a turn after a fixed silence. That fails twice: it cuts people off when they pause to think, and it waits too long when they are done. OpenAI's [turn detection guide](https://developers.openai.com/api/docs/guides/realtime-vad) describes both modes. Server VAD "uses periods of silence to automatically chunk the audio", with a threshold, a silence duration and prefix padding to tune. Semantic VAD "uses a semantic classifier to detect when the user has finished speaking, based on the words they have uttered", and an `eagerness` setting from low to high controls how quickly it answers.

The cascaded frameworks have their own answers. LiveKit's audio turn detector combines "semantic understanding with acoustic cues like intonation, pitch, and rhythm" and, with it enabled, moves the default endpointing window to a 0.3 to 2.5 second range. Pipecat's open-source Smart Turn model (version 3.2, BSD 2-clause) analyses the whole user turn once a lightweight VAD such as Silero reports silence, with inference "as little as 10ms on some CPUs, and under 100ms on most cloud instances". Google's Live API exposes start and end sensitivity and a silence duration, and warns that very low silence values of 100 to 200 ms split utterances into fragments.

**Barge-in** is the other half. When the caller speaks over the agent, three things must happen quickly: stop the audio you are playing (the playbook asks for under about 150 ms), clear every buffered audio chunk on the client, and tell the model what the caller actually heard. In the OpenAI Realtime API the client sends a `conversation.item.truncate` event with the audio position, so the unplayed part is removed from the conversation. In the Gemini Live API the server signals the interruption and the application should stop playback and clear queued audio. If you skip this, the model believes it said things the caller never heard, and the next answer refers to them.

Two failure modes deserve a test each. **False barge-in**: line noise, a cough or the caller's own echo from a speakerphone cancels the agent mid-sentence. Raise the VAD threshold, use echo cancellation, and consider requiring a minimum of speech before interrupting. **Backchannels**: a caller saying "mhm" is not a new turn. I would not let those cut the agent off.

## Tool calls while the agent is talking

A voice agent that can only talk is a toy. The value is in the tools: look up the order, move the appointment, create the ticket. This is the agent loop from [the agent loop explained](https://balazscsorba.com/blog/agent-loop-explained), with a stopwatch attached. Every tool call is a gap in the conversation, and a gap of silence on a phone line reads as a failure.

The mechanics differ by platform. In the OpenAI Realtime API, the model emits a `function_call` item in `response.done`; your code runs the function, returns a `function_call_output` item with the `call_id`, and triggers a new response. Google's guide is more explicit about concurrency: Gemini 3.8 Live supports asynchronous (`NON_BLOCKING`) function execution by default, with scheduling options (`SILENT`, `WHEN_IDLE`, `INTERRUPTED`) that decide whether the result is spoken right away or when the model is idle, while the 3.1 Flash model runs calls sequentially and waits for the response. In a cascaded pipeline you control all of this yourself, which is more work and more flexibility.

What I do in practice:

-   **Speak before you fetch.** Start a short filler ("one moment, I am checking that") as the call begins, whenever the expected wait is over about a second. Pre-written audio costs nothing.
-   **Keep tools fast and small.** Give the agent a handful of narrow tools with timeouts, not a general database client. Return only the fields the model needs to say aloud. See the lessons in [MCP tool design](https://balazscsorba.com/blog/mcp-tool-design-lessons-jira-server).
-   **Confirm before you change anything.** Read back the date, the amount or the address and wait for a spoken yes before a write. Speech recognition errors turn into real actions otherwise.
-   **Treat the transcript as untrusted input.** Anyone can say "ignore your instructions" on a call. The same [injection patterns](https://balazscsorba.com/blog/prompt-injection-lethal-trifecta-patterns) apply, and spoken input is even harder to filter.

## Telephony: SIP, WebRTC and the phone network

Browsers and apps use WebRTC. Phones use SIP and the public telephone network, and most B2B voice use cases are phone calls. The good news is that the platforms now meet you at the SIP layer. With the OpenAI Realtime API you configure a webhook in your project; when a call arrives, OpenAI sends a `realtime.call.incoming` event with the call ID and SIP headers, and you accept, reject, transfer (refer) or hang up through four endpoints. The provider side needs TLS signalling on port 5061 and SRTP media, and after accepting you attach a WebSocket with the call ID to stream events and send commands as usual.

ElevenLabs lists SIP trunk integration and a native Twilio integration for its agents platform. LiveKit's documentation says its agents have full telephony support, so a phone caller joins a room like any other participant. Pipecat runs the same pipeline behind a telephony transport. Whichever you choose, plan for the practical costs: phone audio is narrowband, so recognition quality drops compared with a studio microphone, and the network adds latency to every turn, as the budget above shows.

Two design points save pain later. Keep a **human handoff** path from day one: a SIP transfer to a person when the agent is unsure, the caller asks, or a tool fails twice. And store the **call ID** in every log line, so that a complaint about a call becomes a trace you can replay, with timings per stage.

## German and Hungarian: test the whole chain

Voice quality in a language is a chain, and the weakest link sets the experience. German is well covered everywhere I looked. Hungarian is where the chain breaks, and not in the place people expect: the recognition and the voice exist, but the turn detection often does not. Here is what the documentation says:

Component

German

Hungarian

Deepgram Nova-3 (speech-to-text)

Supported (`de`)

Supported (`hu`)

Deepgram Flux (turn detection in STT)

Supported in `flux-general-multi`, which covers 10 languages

Not listed

LiveKit audio turn detector

Supported

Not listed

Pipecat Smart Turn v3.2

Supported (23 languages)

Not listed

ElevenLabs Flash v2.5 (TTS)

Supported; claimed latency about 75 ms

Supported; Flash v2.5 added it over v2

Realtime models (OpenAI, Gemini)

Not verified per language in what I read

Not verified per language in what I read

For Hungarian this has a concrete consequence. A cascaded stack can use Nova-3 for recognition and Flash v2.5 for the voice, but without a semantic turn detector it falls back to silence-based endpointing, which means longer pauses or more interruptions. The Gemini guide adds that its native audio models pick the language automatically and can switch mid-conversation, with no explicit language setting. That is convenient for a bilingual caller and risky when you need to guarantee Hungarian.

My advice is to build a small evaluation set per language before you choose a vendor: 30 to 50 real recordings with accents, numbers, names and street addresses, scored on word error for the slots you care about and on how natural the endpointing feels. The approach from [evals for LLM features](https://balazscsorba.com/blog/llm-evals-for-product-features) carries over. Vendor language lists tell you what is possible; only your own recordings tell you what is good enough.

## Consent, disclosure and the AI Act

A synthetic voice that talks to people triggers the transparency rules. Article 50(1) of the EU AI Act requires providers to design AI systems that interact directly with people so that those people "are informed that they are interacting with an AI system", unless this is obvious to a reasonably well-informed, observant and circumspect person. Article 50(2) adds that synthetic audio must be marked in a machine-readable, detectable way, where technically feasible. According to a [law firm summary](https://www.joneswalker.com/en/insights/blogs/ai-law-blog/yes-august-2-still-matters-the-eu-approved-a-high-risk-ai-delay-but-most-trans.html?id=102nbon), these duties applied from 2 August 2026 and were not postponed by the Digital Omnibus, which moved the high-risk obligations instead. Only providers of systems placed on the market before that date get until 2 December 2026 for the technical marking, and fines can reach 15 million euros or 3 percent of worldwide turnover.

In practice, I would do four things. **Disclose at the start**: the first sentence of every call says that this is an AI assistant, and the voice does not pretend otherwise. **Offer a human**: tell callers how to reach a person. **Record deliberately**: if you store audio or transcripts, you need a legal basis, a retention period and a clear notice, and a voice recording can identify a person, so it is personal data. My notes on [the GDPR side of LLM APIs](https://balazscsorba.com/blog/gdpr-llm-api-eu-data-residency) cover where audio may be processed. **Be careful with cloned voices**: Article 50(4) requires deployers to disclose deepfake audio, so never imitate a real person's voice without their consent.

**Where legal advice starts**

This is an engineering summary, not legal advice. Call recording rules, consent for outbound calls and sector rules differ between countries and use cases. For the full list of developer obligations, see my [Article 50 developer checklist](https://balazscsorba.com/blog/eu-ai-act-article-50-developer-checklist), and ask your data protection officer before a pilot goes live.

## What I would build first

If a team asked me to start a voice agent next week, this is the sequence I would follow:

1.  Pick one narrow, phone-based process with clear success criteria, such as appointment changes or order status. Define the target: median under about a second to the first spoken word, and a 95th percentile you can live with.
2.  Start with a cascaded pipeline on Pipecat or LiveKit, so that every stage has a log, a timing and a replaceable provider. Run a realtime model as a second variant on the same calls.
3.  Instrument every turn from day one: end of turn, STT final, LLM first token, TTS first byte, and playback start. Plot the median and the 95th percentile per stage.
4.  Choose a semantic turn detector that supports your languages, and test false barge-ins with noisy lines and speakerphones. Where your language has no detector, tune silence thresholds per language.
5.  Build the tools narrow and fast, with timeouts and spoken fillers, and add read-back confirmation before every write.
6.  Add the AI disclosure to the first sentence, a human handoff path and a retention policy for audio and transcripts.
7.  Build the per-language evaluation set from real recordings, and re-run it whenever you change a model or a voice.
8.  Only then optimise cost: smaller models for easy turns, prompt caching, and cheaper voices where quality allows.

The framework choice matters less than discipline about the budget. Pipecat is an open-source (BSD-2) Python framework that orchestrates more than 150 services, and LiveKit Agents is open source under Apache 2.0 and puts the agent into a WebRTC room as a participant. Both let you change providers without rewriting your logic, which is exactly what you want in a field where the best model changes every quarter.

## The bigger picture

Speech-to-speech models will keep closing the gap on latency and quality, and the cascaded pipeline will keep its place wherever text must be inspected. I expect most production systems to become hybrids: a realtime model for small talk and simple turns, and a text path with tools and guardrails for anything that touches a system of record.

The constant is the budget. A voice agent that answers in a second, knows when to stay quiet, says who it is and hands over to a person when it should will beat a cleverer one that does not. If you are planning one and want a second pair of eyes on the architecture, that is exactly the kind of work I do as an [AI engineer](https://balazscsorba.com/expertise/ai-engineer).

## Sources

1.  [OpenAI: Voice agents guide](https://developers.openai.com/api/docs/guides/voice-agents)
2.  [OpenAI: gpt-realtime model page](https://developers.openai.com/api/docs/models/gpt-realtime)
3.  [OpenAI: Voice activity detection in the Realtime API](https://developers.openai.com/api/docs/guides/realtime-vad)
4.  [OpenAI: Realtime API with SIP](https://developers.openai.com/api/docs/guides/realtime-sip)
5.  [OpenAI: Realtime conversations (function calling, interruption)](https://developers.openai.com/api/docs/guides/realtime-conversations)
6.  [Google: Gemini Live API overview](https://ai.google.dev/gemini-api/docs/live)
7.  [Google: Gemini Live API capabilities guide](https://ai.google.dev/gemini-api/docs/live-guide)
8.  [Deepgram: Flux quickstart](https://developers.deepgram.com/docs/flux/quickstart)
9.  [Deepgram: Models and languages overview](https://developers.deepgram.com/docs/models-languages-overview)
10.  [ElevenLabs: Agents platform overview](https://elevenlabs.io/docs/eleven-agents/overview)
11.  [ElevenLabs: Text to speech models and languages](https://elevenlabs.io/docs/overview/capabilities/text-to-speech)
12.  [LiveKit: Agents overview](https://docs.livekit.io/agents/)
13.  [LiveKit: Turn detector](https://docs.livekit.io/agents/logic/turns/turn-detector/)
14.  [Pipecat: Introduction](https://docs.pipecat.ai/getting-started/introduction)
15.  [Pipecat: Smart Turn model (GitHub)](https://github.com/pipecat-ai/smart-turn)
16.  [Fora Soft: Voice AI agents on LiveKit, 2026 engineer playbook](https://www.forasoft.com/blog/article/voice-ai-agents-livekit-guide)
17.  [EU AI Act: Article 50, transparency obligations](https://artificialintelligenceact.eu/article/50/)
18.  [Jones Walker: Yes, August 2 still matters (AI Act delay and Article 50)](https://www.joneswalker.com/en/insights/blogs/ai-law-blog/yes-august-2-still-matters-the-eu-approved-a-high-risk-ai-delay-but-most-trans.html?id=102nbon)

## Frequently asked questions

What is the difference between a speech-to-speech model and an STT, LLM and TTS pipeline?

A speech-to-speech model such as the OpenAI Realtime API or the Gemini Live API takes audio in and produces audio out in one model and one session. A cascaded pipeline chains three components: speech-to-text, a text LLM and text-to-speech. The pipeline gives you text between the stages, which you can store, check and transform, and lets you replace each part independently.

How fast does a voice agent need to be?

Human conversation has response gaps of roughly 200 to 300 ms, callers consciously notice lag at about 500 ms and start to interrupt or hang up around one second, according to one practitioner playbook. A tuned cascaded stack can reach 550 to 700 ms at the median, which is good enough for most business calls.

Should I use the OpenAI Realtime API or Pipecat and LiveKit?

They are not alternatives on the same level. The Realtime API is a model endpoint with WebRTC, WebSocket and SIP connections. Pipecat and LiveKit Agents are open-source frameworks that orchestrate the audio transport, turn detection and your choice of models, and they can use a realtime model or a cascaded pipeline underneath.

Do voice agents work well in German and Hungarian?

German is well supported across the stack. Hungarian is patchier: Deepgram Nova-3 and ElevenLabs Flash v2.5 support it, but the LiveKit turn detector, Pipecat Smart Turn and Deepgram Flux do not list it. I could not confirm Hungarian for the realtime models in vendor documentation, so test with real callers before you commit.

Do I have to tell callers that they are talking to an AI?

Yes, in most cases in the EU. Article 50(1) of the AI Act, applicable since 2 August 2026, requires that people interacting with an AI system are informed of it, unless this is obvious to a reasonably well-informed person. A synthetic voice on a phone line is rarely obvious, so say it at the start of the call.

How do I connect a voice agent to the phone network?

Through SIP or a telephony provider. The OpenAI Realtime API accepts SIP calls and notifies your server by webhook, ElevenLabs offers SIP trunking and a native Twilio integration, and LiveKit supports telephony so a caller joins like any room participant. Expect extra latency on the phone network compared with WebRTC.

Written by Balázs Csorba

Senior fullstack & AI engineer in Styria, Austria – 10+ years of Vue, Nuxt, Node.js and PHP, now building tooling for AI agents.

[AI engineering & MCP servers →](https://balazscsorba.com/expertise/ai-engineer)[About me →](https://balazscsorba.com/about)

## More articles

-   [OpenAI dots: what always-on agents will change, and what they will not](https://balazscsorba.com/blog/openai-dots-always-on-agents-impact)
-   [Human in the loop for AI agents: where to put approval gates](https://balazscsorba.com/blog/human-in-the-loop-ai-agents)
-   [Designing memory for AI agents: tiers, write rules, poisoning and GDPR](https://balazscsorba.com/blog/ai-agent-memory-design)
-   [Multi-agent systems: when they beat one agent, and when they do not](https://balazscsorba.com/blog/multi-agent-systems-when-worth-it)

## Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.

[Book a call](mailto:contact@balazscsorba.com) [Connect on LinkedIn](https://www.linkedin.com/in/balazs-csorba)
