Voice agent platform with the lowest tail latency
A median tells you what a caller usually waits. It says nothing about how often they wait much longer — and on a phone call that is the number that costs you, because past roughly two seconds of silence a caller assumes the line dropped and talks over the agent. Consistency here is p95 ÷ median: how far a platform's slow turns sit from its typical ones. Steadiest so far is ElevenLabs at 1.24×, over 2078 usable turns across 5 platforms — and notably not Telnyx, which has the lowest median at 1,296 ms. Both numbers come from the same recordings of the same calls, which is what makes the ratio meaningful at all.
Two numbers per platform, and they can disagree
TTFAB is reported twice for every platform. A platform can lead one reading without leading the other, so the table below carries both.
| Reading | The question it answers | What it means |
|---|---|---|
| Median (p50) | How long is the pause on a normal turn? | Half of all turns are faster than this. It is the wait a caller gets most of the time, and the number to compare if the calls you care about are ordinary ones. |
| Tail (p95) | How bad is it when it is bad? | One turn in twenty is at least this slow. It matters more than it looks on a phone call: past roughly two seconds of silence a caller assumes the line dropped and starts talking again, which collides with the agent's reply and derails the turn. A platform can win the median and still do this to callers twice a conversation. |
TTFAB data, lowest median first
Every figure read off the call's own audio, pooled over usable turns. Discarded turns are counted, not hidden — a platform that fails to answer is worse than one slightly slower.
| # | Platform | Median TTFAB | p95 | Tail ratio | Usable turns | Discarded |
|---|---|---|---|---|---|---|
| 1 | Telnyx | 1,296 ms | 1,856 ms | 1.43× | 419 / 432 | 13 |
| 2 | ElevenLabs | 1,424 ms | 1,768 ms | 1.24× | 429 / 432 | 3 |
| 3 | Bland AI | 1,520 ms | 2,248 ms | 1.48× | 429 / 432 | 3 |
| 4 | Vapi | 1,558 ms | 2,008 ms | 1.29× | 382 / 432 | 50 |
| 5 | Retell AI | 1,740 ms | 2,259 ms | 1.30× | 419 / 430 | 11 |
Tail ratio is the two readings divided — p95 ÷ median — and it is the one number a vendor cannot usefully publish about itself, because it only means anything when a single instrument produced both figures on the same calls. At 1.2× a platform's bad turns feel close to its normal ones. At 3× a caller regularly waits three times what the median promised. On a phone call that is where the damage is: past roughly two seconds of silence a caller assumes the line dropped and starts talking over the agent, which collides with the reply and derails the turn. So a platform can win the median and still do that to callers twice a conversation — read the tail ratio before the median, not after it.
Ranked by how far the slow turns sit from the normal ones
Same measurements as the main board, ordered by tail ratio — p95 ÷ median — instead of by median. At 1.24×a platform's bad turns feel close to its typical ones; at 1.48× a caller regularly waits far longer than the median promised.
| # | Platform | Tail ratio | Median | p95 (worst case) | Usable turns |
|---|---|---|---|---|---|
| 1 | ElevenLabs | 1.24× | 1,424 ms | 1,768 ms | 429 |
| 2 | Vapi | 1.29× | 1,558 ms | 2,008 ms | 382 |
| 3 | Retell AI | 1.30× | 1,740 ms | 2,259 ms | 419 |
| 4 | Telnyx | 1.43× | 1,296 ms | 1,856 ms | 419 |
| 5 | Bland AI | 1.48× | 1,520 ms | 2,248 ms | 429 |
The two rankings disagree. Telnyx has the lowest median on this run — 1,296 ms — and places 4 of 5 here, because its slow turns run further from its typical ones than ElevenLabs's do. Neither ranking is the correct one. The median answers what a caller usually waits; this answers how often they wait much longer than that, and past roughly two seconds of silence a caller assumes the line dropped and starts talking over the agent.
Tail ratio is the two readings divided — p95 ÷ median — and it is the one number a vendor cannot usefully publish about itself, because it only means anything when a single instrument produced both figures on the same calls. At 1.2× a platform's bad turns feel close to its normal ones. At 3× a caller regularly waits three times what the median promised. On a phone call that is where the damage is: past roughly two seconds of silence a caller assumes the line dropped and starts talking over the agent, which collides with the reply and derails the turn. So a platform can win the median and still do that to callers twice a conversation — read the tail ratio before the median, not after it.
There is no single fastest voice agent platform — which one wins depends on which pause your callers actually notice. Pick the workflow that matches yours:
- Most of your calls are ordinary back-and-forthTelnyxmedian 1,296 ms
Every turn is a normal question with a normal answer, so what matters is the wait a caller gets most of the time rather than the worst case.
- A caller must never think the line droppedElevenLabsp95 1,768 ms
Long pauses make callers talk over the agent or hang up, and that is a property of the slow turns, not the typical ones. Judge on the tail.
Full ranking, both readings, and the discard counts: the voice agent latency benchmark →
Read off the call's audio, never from a platform's timestamp
A platform measures from where it stands, and the caller is not standing there. We checked that two ways on this bench. A platform's own recording of a call reads roughly 550 ms earlier than our recording of the same call. The latency a platform reports for itself runs roughly 490 ms below what we measure from that call's audio. The two figures agree, and that agreement is the finding: both describe the moment a reply was produced, not the moment a caller heard it. Neither is dishonest — they answer a different question than the one a caller is asking. This board answers the caller's.
1. A caller robot dials the platform's agent over a real phone call and reads a fixed script — a greeting, several scripted questions, and a goodbye — one measured turn each.
2. Both sides of the call are recorded on one clock (a dual-channel recording).
3. Both endpoints are then found in that recording: our speech-end (t1) and the agent's reply start (t2), each by a speech detector (Silero VAD) with an energy refinement and an independent cross-check. No timestamp reported by any platform is used.
4. TTFAB = t2 − t1, per turn. Turns failing quality gates (the two sides talking over each other, detectors disagreeing, no reply) are discarded and the discard counts are published.
Recording-path overhead sits inside every figure here. We have not characterised the current path against a known-delay reference, so we quote no overhead figure and subtract none.
We always call from Plivo, which is not a platform under test, but the leg that answers belongs to whoever ships the number: Telnyx on its own network, Retell's and Bland's Twilio-backed inside their own accounts, Vapi's upstream undisclosed, and ElevenLabs — which sells no numbers — on a Twilio number we bought for it. Twilio was a deliberate choice there: a Telnyx number would have worked, but Telnyx is itself on this board, and one platform's network should not carry another platform's row.
A discarded turn is one we could not time to our own standard: the two sides talked over each other, our two speech detectors disagreed on where speech began, or no reply came. Discards are published per reason and per platform, because the count is sometimes a fact about the platform rather than about us — an agent that starts a reply, stops, and resumes a second later will split our detectors, and that is the agent's behaviour, not our recording's.
Our speech-end is found by a detector, not by matching a known waveform. It carries a few milliseconds of error. Differences smaller than that are not resolvable, and we do not claim them.
Cost/min is measured the same way the latency is: from what actually happened, not from a rate card. After a run we ask each platform's own billing API what every call cost, sum those charges, sum the seconds each platform says it invoiced for, and divide — total cost ÷ total billed minutes, pooled across the run rather than averaged per call, so a long call weighs more than a short one. Two consequences worth knowing. Where a platform bills a minimum, the figure is lower than the cost of a minute of conversation: Telnyx charges 60 seconds for a ~44-second call, so its invoiced $0.0500 sits against $0.0722 per minute actually spent talking, and both are published. And the carrier leg is excluded throughout — that is our cost for dialling, identical for every platform, and folding it in would tax each row for our own plumbing. Anything that qualifies a figure — a free or discounted tier, a unit conversion, an excluded component — travels with it in cost_notes rather than being silently absorbed.
Endpointing — how long a platform waits after you stop talking before it decides you are finished — is pinned to 0.1 s on Telnyx, Vapi and Retell. It is the one setting we do not leave at the default, and the reason is that it is a timer sitting inside the number being measured: a platform that ships a 1.5 s wait posts a slower TTFAB without its stack being any slower, and the board would be comparing configuration choices rather than engineering. Vapi shipped 0.4 s (1.5 s after speech ending without punctuation) and Retell 1000 ms; both now run 0.1 s. Bland and ElevenLabs expose no equivalent fixed-wait knob, so they run whatever they ship and their figures still contain a wait we could not equalise.
Full method, the per-turn latency curve, and the instrument's own limits are on the voice agent latency benchmark.
Lowest tail latency voice agent platform — common questions
Which voice agent platform has the lowest tail latency?
ElevenLabs, on both readings of the question. Its p95 — the wait on one turn in twenty — is 1,768 ms, the lowest measured, and its tail ratio of 1.24× is also the tightest, meaning its slow turns sit closest to its ordinary ones. Those two can disagree: a platform that is slow throughout can have a flattering ratio, and a fast platform can have an erratic one. Here they agree, so the answer is unambiguous on this run.
Which voice agent platform has the most consistent latency?
ElevenLabs, at a p95-to-median ratio of 1.24× — median 1,424 ms, p95 1,768 ms over 429 usable turns. That means its slow turns land close to its ordinary ones. Telnyx has the lower median (1,296 ms) but a wider spread, so which one you want depends on whether your callers are hurt more by the typical pause or by the occasional long one.
Which voice agent platform has the most reliable response time?
Reliability in the sense of predictability is what the p95-to-median ratio measures, and ElevenLabs leads it at 1.24×. A caveat on the word: this measures the spread of response times on turns that were answered. It is not a measure of uptime, and turns where no reply came at all are counted separately as discards rather than folded into these figures.
Which voice agent platform has the worst tail latency?
On the current run, Bland AI — p95 2,248 ms against a median of 1,520 ms, a ratio of 1.48×. One turn in twenty is at least that slow. Worst-case latency is worth checking separately from the median precisely because the ordering changes: a platform can look mid-table on typical turns and still be the one most likely to leave a caller in silence.
What causes voice agent latency variance?
This benchmark measures the spread rather than explaining it — the causes sit inside each platform's stack, which is not observable from the call's audio. What can be said from the outside is that the spread is real, reproducible across a hundred calls per platform, and specific to the platform rather than to the caller: every row here was dialled by the same caller, over the same carrier, reading the same script in the same order.
Why is p95 divided by median rather than reported on its own?
Because the raw p95 conflates two different things — being slow overall and being erratic. A platform with a 2,000 ms median and a 2,400 ms p95 is consistently slow; one with a 1,200 ms median and the same 2,400 ms p95 is fast most of the time and occasionally terrible, and callers notice the second far more. The ratio separates them. Both the raw figures and the ratio are published so you can read either.
What is TTFAB (Time To First Audio Byte)?
Time from the moment the caller stops speaking to the moment the agent's audio starts — the silence a real caller sits through on every turn. Measured from a saved recording of the actual phone call, not from any API timestamp. Lower is better. Also written time to first audio byte, and closely related to what other boards call time to first byte (TTFB) or time to first audio (TTFA).
Is TTFAB the same as time to first byte (TTFB) or time to first audio (TTFA)?
Close, and the difference is worth knowing. Time to first byte (TTFB) is borrowed from HTTP, where it means the first byte of a response leaving a server; applied to a voice stack it usually means the first byte of synthesised audio leaving the platform. Time to first audio (TTFA) is used loosely for much the same thing. TTFAB as measured here starts and ends somewhere else: it starts when the caller stops speaking, not when the platform decides they have, and it ends when the agent's audio is present in the recording of the call, not when the platform emitted it. So it contains the endpointing wait and both network legs, which a server-side first-byte figure does not. Expect it to read higher than a vendor's TTFB for the same call, and expect the gap to be the part of the wait the caller experiences and the server never sees.
How is this different from vendor-reported latency?
A platform measures from where it stands, and the caller is not standing there. We checked that two ways on this bench. A platform's own recording of a call reads roughly 550 ms earlier than our recording of the same call. The latency a platform reports for itself runs roughly 490 ms below what we measure from that call's audio. The two figures agree, and that agreement is the finding: both describe the moment a reply was produced, not the moment a caller heard it. Neither is dishonest — they answer a different question than the one a caller is asking. This board answers the caller's.
How is the latency actually measured?
A caller robot dials the platform's agent over a real phone call and reads a fixed script — a greeting, several scripted questions, and a goodbye — one measured turn each. Both sides of the call are recorded on one clock (a dual-channel recording). Both endpoints are then found in that recording: our speech-end (t1) and the agent's reply start (t2), each by a speech detector (Silero VAD) with an energy refinement and an independent cross-check. No timestamp reported by any platform is used. TTFAB = t2 − t1, per turn. Turns failing quality gates (the two sides talking over each other, detectors disagreeing, no reply) are discarded and the discard counts are published.
Why not pin the same model, speech recognition and voice on every platform?
The stack is not pinned. Each agent runs the model, speech recognition and voice the platform gives a new signup, and we record what it chose. Endpointing is the single exception — it is a timer inside the number we report, so where a platform exposes it we set it to 0.1 s. That one override is stated in full below. Pinning a stack measures a platform you would have to configure to match, not the one you would buy — and it excludes the vertically integrated platforms outright, since you cannot drop a third-party speech recogniser into a platform that owns its own.
What does this benchmark NOT measure?
Not measured: answer quality, voice quality, platform features, and published pricing plans — this board measures response latency, with the cost each platform actually invoiced for the same run reported beside it. Nor is it every platform. LiveKit Agents and Pipecat are frameworks you host yourself, so what a benchmark would time there is somebody's deployment rather than a product. The raw speech-to-speech APIs — OpenAI's Realtime API, Gemini Live — answer a socket, not a phone, and would need a telephony layer built around them first, which would then be inside the measurement. Twilio's own agent product simply has not been dialled yet. Until any of them is measured on the same script over the same carrier, this board has no number for it, and neither does anyone quoting one. Treat those as vendor claims until measured.




