Voice agent provider with the lowest latency
Fastest provider measured so far: Telnyx at 1,270 ms median, over 5 providers and 195 usable turns. If you are picking a provider on speed, the number that matters is not an API round trip. It is the silence your caller sits through: the gap between them stopping and the agent's first audio, on a real phone call, timed from the recording rather than from the provider's own clock.
Two numbers per platform, and they can disagree
TTFAB is reported twice for every platform. A platform can lead one reading without leading the other, so the table below carries both.
| Reading | The question it answers | What it means |
|---|---|---|
| Median (p50) | How long is the pause on a normal turn? | Half of all turns are faster than this. It is the wait a caller gets most of the time, and the number to compare if the calls you care about are ordinary ones. |
| Tail (p95) | How bad is it when it is bad? | One turn in twenty is at least this slow. It matters more than it looks on a phone call: past roughly two seconds of silence a caller assumes the line dropped and starts talking again, which collides with the agent's reply and derails the turn. A platform can win the median and still do this to callers twice a conversation. |
TTFAB data, lowest median first
Every figure read off the call's own audio, pooled over usable turns. Discarded turns are counted, not hidden — a platform that fails to answer is worse than one slightly slower.
| # | Platform | Median TTFAB | p95 | p95 / median | Usable turns | Discarded |
|---|---|---|---|---|---|---|
| 1 | Telnyx | 1,270 ms | 2,015 ms | 1.59× | 40 / 40 | 0 |
| 2 | ElevenLabs | 1,362 ms | 1,633 ms | 1.20× | 40 / 40 | 0 |
| 3 | Bland AI | 1,493 ms | 2,162 ms | 1.45× | 40 / 40 | 0 |
| 4 | Vapi | 1,508 ms | 1,912 ms | 1.27× | 37 / 40 | 3 |
| 5 | Retell AI | 1,814 ms | 2,329 ms | 1.28× | 38 / 39 | 1 |
The p95 / median column is the two read together — the tail penalty. At 1.2× a platform's bad turns feel close to its normal ones. At 3× a caller regularly waits three times what the median promised.
There is no single fastest voice agent platform — which one wins depends on which pause your callers actually notice. Pick the workflow that matches yours:
- Most of your calls are ordinary back-and-forthTelnyxmedian 1,270 ms
Every turn is a normal question with a normal answer, so what matters is the wait a caller gets most of the time rather than the worst case.
- A caller must never think the line droppedElevenLabsp95 1,633 ms
Long pauses make callers talk over the agent or hang up, and that is a property of the slow turns, not the typical ones. Judge on the tail.
Full ranking, both readings, and the discard counts: the voice agent latency benchmark →
Read off the call's audio, never from a platform's timestamp
A platform measures from where it stands, and the caller is not standing there. We checked that two ways on this bench. A platform's own recording of a call reads roughly 550 ms earlier than our recording of the same call. The latency a platform reports for itself runs roughly 490 ms below what we measure from that call's audio. The two figures agree, and that agreement is the finding: both describe the moment a reply was produced, not the moment a caller heard it. Neither is dishonest — they answer a different question than the one a caller is asking. This board answers the caller's.
1. A caller robot dials the platform's agent over a real phone call and reads a fixed script — a greeting, several scripted questions, and a goodbye — one measured turn each.
2. Both sides of the call are recorded on one clock (a dual-channel recording).
3. Both endpoints are then found in that recording: our speech-end (t1) and the agent's reply start (t2), each by a speech detector (Silero VAD) with an energy refinement and an independent cross-check. No timestamp reported by any platform is used.
4. TTFAB = t2 − t1, per turn. Turns failing quality gates (the two sides talking over each other, detectors disagreeing, no reply) are discarded and the discard counts are published.
Recording-path overhead sits inside every figure here. We have not characterised the current path against a known-delay reference, so we quote no overhead figure and subtract none.
We always call from Plivo, which is not a platform under test, but the leg that answers belongs to whoever ships the number: Telnyx on its own network, Retell's and Bland's Twilio-backed inside their own accounts, Vapi's upstream undisclosed, and ElevenLabs — which sells no numbers — on a Twilio number we bought for it. Twilio was a deliberate choice there: a Telnyx number would have worked, but Telnyx is itself on this board, and one platform's network should not carry another platform's row.
A discarded turn is one we refuse to time, not one the platform failed. The two sides talked over each other, our detectors disagreed, or no reply came. Discards are published per reason, because a platform that never answers is worse than one that answers slowly.
Our speech-end is found by a detector, not by matching a known waveform. It carries a few milliseconds of error. Differences smaller than that are not resolvable, and we do not claim them.
Endpointing — how long a platform waits after you stop talking before it decides you are finished — is pinned to 0.1 s on Telnyx, Vapi and Retell. It is the one setting we do not leave at the default, and the reason is that it is a timer sitting inside the number being measured: a platform that ships a 1.5 s wait posts a slower TTFAB without its stack being any slower, and the board would be comparing configuration choices rather than engineering. Vapi shipped 0.4 s (1.5 s after speech ending without punctuation) and Retell 1000 ms; both now run 0.1 s. Bland and ElevenLabs expose no equivalent fixed-wait knob, so they run whatever they ship and their figures still contain a wait we could not equalise.
Full method, the per-turn latency curve, and the instrument's own limits are on the voice agent latency benchmark.
Fastest voice agent provider — common questions
Which voice agent provider is fastest?
Telnyx, on median TTFAB, at 1,270 ms measured from real phone-call audio. Treat that as the current measured leader on a small sample rather than a settled result — the per-provider turn count is published with every figure, and more providers are being added. Check the tail as well as the median: a provider can win the median and still make callers wait far longer on its slow turns.
What counts as latency when comparing voice agent providers?
Not request-to-response on an HTTP call. The wait a caller experiences is the whole chain after they stop speaking: speech recognition has to decide the turn ended, the model has to produce a reply, and speech synthesis has to emit its first audio. TTFAB measures that end to end from the call's own recording, so no part of the path is excluded by definition — which is also why a provider's internal API timing can look much faster than what a caller hears.
Why not use the latency figures each provider publishes?
We measured the same calls twice — from our own recording and from the provider's — and the provider's read faster. That is consistent with a timestamp taken when a reply is generated rather than when a caller can hear it. Both readings can be honest and still differ, which is why every figure here comes from audio a caller would actually have heard.
Does the fastest provider depend on which model it runs?
Partly, and we do not separate the two. Each provider runs the model, speech recognition and voice it ships to a new signup, because that is the product you would buy. So a fast figure may be fast turn-taking, a fast model, or aggressive endpointing — the configuration each provider chose is recorded and published with its run, which is what makes the figure interpretable rather than merely low.
What is TTFAB (Time To First Audio Byte)?
Time from the moment the caller stops speaking to the moment the agent's audio starts — the silence a real caller sits through on every turn. Measured from a saved recording of the actual phone call, not from any API timestamp. Lower is better.
How is this different from vendor-reported latency?
A platform measures from where it stands, and the caller is not standing there. We checked that two ways on this bench. A platform's own recording of a call reads roughly 550 ms earlier than our recording of the same call. The latency a platform reports for itself runs roughly 490 ms below what we measure from that call's audio. The two figures agree, and that agreement is the finding: both describe the moment a reply was produced, not the moment a caller heard it. Neither is dishonest — they answer a different question than the one a caller is asking. This board answers the caller's.
How is the latency actually measured?
A caller robot dials the platform's agent over a real phone call and reads a fixed script — a greeting, several scripted questions, and a goodbye — one measured turn each. Both sides of the call are recorded on one clock (a dual-channel recording). Both endpoints are then found in that recording: our speech-end (t1) and the agent's reply start (t2), each by a speech detector (Silero VAD) with an energy refinement and an independent cross-check. No timestamp reported by any platform is used. TTFAB = t2 − t1, per turn. Turns failing quality gates (the two sides talking over each other, detectors disagreeing, no reply) are discarded and the discard counts are published.
Why not pin the same model, speech recognition and voice on every platform?
The stack is not pinned. Each agent runs the model, speech recognition and voice the platform gives a new signup, and we record what it chose. Endpointing is the single exception — it is a timer inside the number we report, so where a platform exposes it we set it to 0.1 s. That one override is stated in full below. Pinning a stack measures a platform you would have to configure to match, not the one you would buy — and it excludes the vertically integrated platforms outright, since you cannot drop a third-party speech recogniser into a platform that owns its own.
What does this benchmark NOT measure?
Not measured: answer quality, voice quality, price, and platform features — this board measures response latency only. Treat those as vendor claims until measured.