{
  "slug": "voice-agent-latency",
  "name": "Voice Agent Latency Benchmark",
  "description": "First-party benchmark measuring Time To First Audio Byte (TTFAB) of voice AI agent platforms over real phone calls, from saved call audio — never from platform-reported timestamps. Time from the moment the caller stops speaking to the moment the agent's audio starts — the silence a real caller sits through on every turn. Measured from a saved recording of the actual phone call, not from any API timestamp. Lower is better. Also written time to first audio byte, and closely related to what other boards call time to first byte (TTFB) or time to first audio (TTFA). Not measured: answer quality, voice quality, platform features, and published pricing plans — this board measures response latency, with the cost each platform actually invoiced for the same run reported beside it. Nor is it every platform. LiveKit Agents and Pipecat are frameworks you host yourself, so what a benchmark would time there is somebody's deployment rather than a product. The raw speech-to-speech APIs — OpenAI's Realtime API, Gemini Live — answer a socket, not a phone, and would need a telephony layer built around them first, which would then be inside the measurement. Twilio's own agent product simply has not been dialled yet. Until any of them is measured on the same script over the same carrier, this board has no number for it, and neither does anyone quoting one.",
  "page_url": "https://openbenchmarks.com/voice-agent-latency",
  "last_updated": "2026-08-22T18:00:00.000Z",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "metrics": [
    {
      "key": "ttfab_onset_p50",
      "label": "TTFAB (Time To First Audio Byte), median",
      "unit": "milliseconds",
      "direction": "lower_is_better",
      "definition": "HEADLINE ranking metric. Time from the moment the caller stops speaking to the moment the agent's audio starts — the silence a real caller sits through on every turn. Measured from a saved recording of the actual phone call, not from any API timestamp. Lower is better. Also written time to first audio byte, and closely related to what other boards call time to first byte (TTFB) or time to first audio (TTFA)."
    },
    {
      "key": "ttfab_onset_p95",
      "label": "TTFAB, p95",
      "unit": "milliseconds",
      "direction": "lower_is_better",
      "definition": "95th percentile of TTFAB over usable turns — the tail a caller hits one turn in twenty."
    },
    {
      "key": "cost_per_minute",
      "label": "Cost per minute",
      "unit": "USD",
      "direction": "lower_is_better",
      "definition": "What a minute of conversation costs, from each platform's own billing API rather than its price list. Pooled: total charged over total call duration, so short calls do not weigh the same as long ones. cost_per_billed_minute is the same money over the seconds actually invoiced, and differs wherever a platform applies a minimum."
    },
    {
      "key": "consistency_p95_over_p50",
      "label": "Consistency (p95/p50)",
      "unit": "ratio",
      "direction": "lower_is_better",
      "definition": "Tail ratio: how much worse the slow turns are than the typical turn. 1.0 is perfectly consistent."
    },
    {
      "key": "turns_usable",
      "label": "Usable turns",
      "unit": "count",
      "direction": "higher_is_better",
      "definition": "Turns that survived every measurement-quality gate. Discarded turns are counted and published by reason."
    }
  ],
  "method_notes": [
    "Recording-path overhead sits inside every figure here. We have not characterised the current path against a known-delay reference, so we quote no overhead figure and subtract none.",
    "We always call from Plivo, which is not a platform under test, but the leg that answers belongs to whoever ships the number: Telnyx on its own network, Retell's and Bland's Twilio-backed inside their own accounts, Vapi's upstream undisclosed, and ElevenLabs — which sells no numbers — on a Twilio number we bought for it. Twilio was a deliberate choice there: a Telnyx number would have worked, but Telnyx is itself on this board, and one platform's network should not carry another platform's row.",
    "A discarded turn is one we could not time to our own standard: the two sides talked over each other, our two speech detectors disagreed on where speech began, or no reply came. Discards are published per reason and per platform, because the count is sometimes a fact about the platform rather than about us — an agent that starts a reply, stops, and resumes a second later will split our detectors, and that is the agent's behaviour, not our recording's.",
    "Our speech-end is found by a detector, not by matching a known waveform. It carries a few milliseconds of error. Differences smaller than that are not resolvable, and we do not claim them.",
    "Cost/min is measured the same way the latency is: from what actually happened, not from a rate card. After a run we ask each platform's own billing API what every call cost, sum those charges, sum the seconds each platform says it invoiced for, and divide — total cost ÷ total billed minutes, pooled across the run rather than averaged per call, so a long call weighs more than a short one. Two consequences worth knowing. Where a platform bills a minimum, the figure is lower than the cost of a minute of conversation: Telnyx charges 60 seconds for a ~44-second call, so its invoiced $0.0500 sits against $0.0722 per minute actually spent talking, and both are published. And the carrier leg is excluded throughout — that is our cost for dialling, identical for every platform, and folding it in would tax each row for our own plumbing. Anything that qualifies a figure — a free or discounted tier, a unit conversion, an excluded component — travels with it in cost_notes rather than being silently absorbed.",
    "Endpointing — how long a platform waits after you stop talking before it decides you are finished — is pinned to 0.1 s on Telnyx, Vapi and Retell. It is the one setting we do not leave at the default, and the reason is that it is a timer sitting inside the number being measured: a platform that ships a 1.5 s wait posts a slower TTFAB without its stack being any slower, and the board would be comparing configuration choices rather than engineering. Vapi shipped 0.4 s (1.5 s after speech ending without punctuation) and Retell 1000 ms; both now run 0.1 s. Bland and ElevenLabs expose no equivalent fixed-wait knob, so they run whatever they ship and their figures still contain a wait we could not equalise."
  ],
  "results": [
    {
      "rank": 1,
      "provider": "Telnyx",
      "provider_slug": "telnyx",
      "ttfab_onset_p50": 1296,
      "ttfab_onset_p90": 1726,
      "ttfab_onset_p95": 1856,
      "ttfab_onset_p99": 2126,
      "cost_per_minute": 0.071877,
      "cost_per_billed_minute": 0.05,
      "cost_per_call": 0.051695,
      "cost_currency": "USD",
      "cost_calls_priced": 59,
      "cost_notes": [
        "billed 120s for a 66s call (Telnyx bills a 60-second minimum)",
        "billed 120s for a 70s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 16s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 35s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 38s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 39s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 40s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 41s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 42s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 43s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 44s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 45s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 46s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 47s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 48s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 50s call (Telnyx bills a 60-second minimum)",
        "billed 60s for a 51s call (Telnyx bills a 60-second minimum)",
        "cost is the ai-voice-assistant detail record; a separate PSTN leg may be billed under another record type"
      ],
      "consistency_p95_over_p50": 1.432,
      "turns_usable": 419,
      "turn_attempts": 432,
      "discards": {
        "no_response": 4,
        "vad_disagree": 9
      },
      "barge_in_turns": 6,
      "carrier": "plivo",
      "latest_run_id": "bench-telnyx-20260731-152959",
      "latest_run_at": "2026-07-31T15:29:59+00:00",
      "analyzer_version": "2.4.0",
      "vendor_config_sha256": "5277d96e7b0c2e5c95641a0f276bba058fb43354f778e89fe7392c8cb839a876",
      "vendor_defaults_used": {
        "model": "moonshotai/Kimi-K2.6",
        "tools": [
          "hangup"
        ],
        "voice": "Telnyx.Ultra.f786b574-daa5-4673-aa0c-cbe3e8534c02",
        "stt_model": "deepgram/flux",
        "version_id": "20260730T211656915334",
        "endpointing": {
          "eot_threshold": 0.8,
          "eot_timeout_ms": 5000,
          "eager_eot_threshold": 0.8,
          "start_speaking_wait_seconds": 0.1,
          "interrupt_prediction_threshold": 0.55,
          "transcription_endpointing_plan": {
            "on_number_seconds": 0.1,
            "on_punctuation_seconds": 0.1,
            "on_no_punctuation_seconds": 0.1
          }
        },
        "voice_speed": 1,
        "assistant_id": "assistant-0e01469e-4458-441b-be3d-a4f97e6ff2d0",
        "stt_language": "en",
        "time_limit_secs": 1800,
        "background_audio": {
          "type": "predefined_media",
          "value": "silence",
          "volume": 0.5
        },
        "noise_suppression": "disabled",
        "user_idle_reply_secs": 10
      },
      "vendor_unsupported": [
        "llm_temperature"
      ],
      "honesty": [
        "The instrument is UNCHARACTERISED for this measurement path. The known-delay reference that produced our overhead and noise figures replied over a carrier media websocket, and that path no longer exists -- a characterisation is only valid for the host, carrier and audio path it was measured on, so no overhead figure is quoted here and none is subtracted (V1.md D4). Recording-path overhead is inside every number below, unmeasured.",
        "t1 -- the end of our own speech -- is found by a speech detector on the near channel, not by matching a known waveform. It is accurate to roughly a few tens of ms rather than sample-exact, so small differences between vendors are not resolvable; the per-turn error budget is the one Gate A measures on synthetic dialogs.",
        "Turn-taking is Plivo's speech endpointing, not ours. It decides when the conversation moves on, never when a measurement starts or ends: both endpoints of every TTFAB are read off the recording afterwards. Eager or slow endpointing therefore costs discarded turns, not shifted numbers.",
        "Our speech is Polly.Joanna text-to-speech reading a fixed script, not the human recording V1.md 11.1 requires (TTS endings trip endpointing differently). It is re-rendered per call rather than played from one committed file, so the stimulus is identical in wording across vendors but not byte-identical across calls; t1 is measured from each call's own tape, so this costs no accuracy.",
        "Answer accuracy is a keyword match against the carrier's transcript of the vendor's reply -- a coarse check that the agent answered from its prompt, reported alongside latency and never mixed into it.",
        "Historical, on the DELETED websocket path: a vendor's own recording of the same calls read 547 ms lower than ours (bench-telnyx-20260729-182644, n=12 turns) -- more than that path could account for. Read vendor-reported latency with care. Not re-measured on this path.",
        "The turn curve has as few as n=91 measurements at some turn indices. Pooled figures are the sturdier read; per-turn percentiles (especially p95) are directional until each index has many more.",
        "Discarded calls are reported above, never silently dropped -- a vendor that fails to respond is worse than one that is slightly slower."
      ]
    },
    {
      "rank": 2,
      "provider": "ElevenLabs",
      "provider_slug": "elevenlabs",
      "ttfab_onset_p50": 1424,
      "ttfab_onset_p90": 1673.6,
      "ttfab_onset_p95": 1768,
      "ttfab_onset_p99": 2254.3,
      "cost_per_minute": 0.079384,
      "cost_per_billed_minute": 0.079384,
      "cost_per_call": 0.056622,
      "cost_currency": "USD",
      "cost_calls_priced": 108,
      "cost_notes": [
        "account tier is 'creator' -- cost_fiat is the list-price equivalent, not an amount billed",
        "account tier is 'free' -- cost_fiat is the list-price equivalent, not an amount billed",
        "cost_fiat is the USD figure; `cost` is in credits"
      ],
      "consistency_p95_over_p50": 1.242,
      "turns_usable": 429,
      "turn_attempts": 432,
      "discards": {
        "vad_disagree": 3
      },
      "barge_in_turns": 0,
      "carrier": "plivo",
      "latest_run_id": "bench-elevenlabs-20260731-152959",
      "latest_run_at": "2026-07-31T15:29:59+00:00",
      "analyzer_version": "2.4.0",
      "vendor_config_sha256": "f697ccc773d68033c7e35f5ed5d3e3b6ddc6f0b6fefa98c2dcf7cebbc0856976",
      "vendor_defaults_used": {
        "model": "gemini-2.5-flash",
        "tools": [],
        "voice": "eleven_flash_v2/cjVigY5qzO86Huf0OWal",
        "agent_id": "agent_1101kyt41p3hft4aaepgwt7rc4sj",
        "language": "en",
        "branch_id": "agtbrch_8701kyt41pr5fgm96a6vjp1zzj95",
        "stt_model": "scribe_realtime/high",
        "backup_llm": null,
        "version_id": "agtvrsn_9901kyx2sbxke7vr14bjf6wbfsz9",
        "endpointing": {
          "turn_mode": "turn",
          "turn_model": "turn_v3",
          "turn_eagerness": "normal",
          "turn_timeout_s": 10,
          "speculative_turn": false,
          "initial_wait_time": null,
          "background_voice_detection": false,
          "disable_first_message_interruptions": false
        },
        "voice_speed": 1,
        "filler_armed": false,
        "built_in_tools": [],
        "filler_message": "Hhmmmm...yeah.",
        "max_duration_s": 600,
        "model_max_tokens": -1,
        "reasoning_effort": null,
        "stt_audio_format": "pcm_16000",
        "model_temperature": 0,
        "filler_timeout_seconds": -1,
        "tts_optimize_streaming_latency": 3
      },
      "vendor_unsupported": [],
      "honesty": [
        "The instrument is UNCHARACTERISED for this measurement path. The known-delay reference that produced our overhead and noise figures replied over a carrier media websocket, and that path no longer exists -- a characterisation is only valid for the host, carrier and audio path it was measured on, so no overhead figure is quoted here and none is subtracted (V1.md D4). Recording-path overhead is inside every number below, unmeasured.",
        "t1 -- the end of our own speech -- is found by a speech detector on the near channel, not by matching a known waveform. It is accurate to roughly a few tens of ms rather than sample-exact, so small differences between vendors are not resolvable; the per-turn error budget is the one Gate A measures on synthetic dialogs.",
        "Turn-taking is Plivo's speech endpointing, not ours. It decides when the conversation moves on, never when a measurement starts or ends: both endpoints of every TTFAB are read off the recording afterwards. Eager or slow endpointing therefore costs discarded turns, not shifted numbers.",
        "Our speech is Polly.Joanna text-to-speech reading a fixed script, not the human recording V1.md 11.1 requires (TTS endings trip endpointing differently). It is re-rendered per call rather than played from one committed file, so the stimulus is identical in wording across vendors but not byte-identical across calls; t1 is measured from each call's own tape, so this costs no accuracy.",
        "Answer accuracy is a keyword match against the carrier's transcript of the vendor's reply -- a coarse check that the agent answered from its prompt, reported alongside latency and never mixed into it.",
        "Historical, on the DELETED websocket path: a vendor's own recording of the same calls read 547 ms lower than ours (bench-telnyx-20260729-182644, n=12 turns) -- more than that path could account for. Read vendor-reported latency with care. Not re-measured on this path.",
        "The turn curve has as few as n=95 measurements at some turn indices. Pooled figures are the sturdier read; per-turn percentiles (especially p95) are directional until each index has many more.",
        "Discarded calls are reported above, never silently dropped -- a vendor that fails to respond is worse than one that is slightly slower."
      ]
    },
    {
      "rank": 3,
      "provider": "Bland AI",
      "provider_slug": "bland",
      "ttfab_onset_p50": 1520,
      "ttfab_onset_p90": 2009.6,
      "ttfab_onset_p95": 2247.6,
      "ttfab_onset_p99": 2958.1,
      "cost_per_minute": 0.140802,
      "cost_per_billed_minute": 0.140802,
      "cost_per_call": 0.105167,
      "cost_currency": "USD",
      "cost_calls_priced": 108,
      "cost_notes": [
        "Bland publishes no cost breakdown -- the figure is a single all-in price",
        "call_length is reported in minutes; converted to seconds"
      ],
      "consistency_p95_over_p50": 1.479,
      "turns_usable": 429,
      "turn_attempts": 432,
      "discards": {
        "no_response": 1,
        "vad_disagree": 2
      },
      "barge_in_turns": 0,
      "carrier": "plivo",
      "latest_run_id": "bench-bland-20260731-152959",
      "latest_run_at": "2026-07-31T15:29:59+00:00",
      "analyzer_version": "2.4.0",
      "vendor_config_sha256": "e82a64ec97b1a8362db4c9d1a59ed73cc9eb1fbfbac8d4736ae7baaf5e9e4451",
      "vendor_defaults_used": {
        "tools": [],
        "voice": null,
        "record": false,
        "language": "ENG",
        "created_at": "2026-07-30T18:00:14.148Z",
        "model_tier": "enhanced",
        "pathway_id": null,
        "endpointing": {
          "keywords": null,
          "reduce_latency": null,
          "interruptibility": null,
          "resumption_speed": null,
          "block_interruptions": null,
          "interruption_threshold": 500
        },
        "temperature": 0,
        "phone_number": "+16282653857",
        "voice_settings": null,
        "background_track": null,
        "max_duration_min": 30,
        "noise_cancellation": true,
        "silence_end_message": null
      },
      "vendor_unsupported": [
        "llm_identity",
        "stt_provider",
        "config_version"
      ],
      "honesty": [
        "The instrument is UNCHARACTERISED for this measurement path. The known-delay reference that produced our overhead and noise figures replied over a carrier media websocket, and that path no longer exists -- a characterisation is only valid for the host, carrier and audio path it was measured on, so no overhead figure is quoted here and none is subtracted (V1.md D4). Recording-path overhead is inside every number below, unmeasured.",
        "t1 -- the end of our own speech -- is found by a speech detector on the near channel, not by matching a known waveform. It is accurate to roughly a few tens of ms rather than sample-exact, so small differences between vendors are not resolvable; the per-turn error budget is the one Gate A measures on synthetic dialogs.",
        "Turn-taking is Plivo's speech endpointing, not ours. It decides when the conversation moves on, never when a measurement starts or ends: both endpoints of every TTFAB are read off the recording afterwards. Eager or slow endpointing therefore costs discarded turns, not shifted numbers.",
        "Our speech is Polly.Joanna text-to-speech reading a fixed script, not the human recording V1.md 11.1 requires (TTS endings trip endpointing differently). It is re-rendered per call rather than played from one committed file, so the stimulus is identical in wording across vendors but not byte-identical across calls; t1 is measured from each call's own tape, so this costs no accuracy.",
        "Answer accuracy is a keyword match against the carrier's transcript of the vendor's reply -- a coarse check that the agent answered from its prompt, reported alongside latency and never mixed into it.",
        "Historical, on the DELETED websocket path: a vendor's own recording of the same calls read 547 ms lower than ours (bench-telnyx-20260729-182644, n=12 turns) -- more than that path could account for. Read vendor-reported latency with care. Not re-measured on this path.",
        "The turn curve has as few as n=97 measurements at some turn indices. Pooled figures are the sturdier read; per-turn percentiles (especially p95) are directional until each index has many more.",
        "Discarded calls are reported above, never silently dropped -- a vendor that fails to respond is worse than one that is slightly slower."
      ]
    },
    {
      "rank": 4,
      "provider": "Vapi",
      "provider_slug": "vapi",
      "ttfab_onset_p50": 1558,
      "ttfab_onset_p90": 1850,
      "ttfab_onset_p95": 2007.8,
      "ttfab_onset_p99": 2688.6,
      "cost_per_minute": 0.083625,
      "cost_per_billed_minute": 0.083625,
      "cost_per_call": 0.062669,
      "cost_currency": "USD",
      "cost_calls_priced": 103,
      "cost_notes": [],
      "consistency_p95_over_p50": 1.289,
      "turns_usable": 382,
      "turn_attempts": 432,
      "discards": {
        "no_response": 1,
        "vad_disagree": 49
      },
      "barge_in_turns": 4,
      "carrier": "plivo",
      "latest_run_id": "bench-vapi-20260731-152959",
      "latest_run_at": "2026-07-31T15:29:59+00:00",
      "analyzer_version": "2.4.0",
      "vendor_config_sha256": "4b8a525c1f804d6a7ea1c1c86784c7bcad0cb253c151a039a7301c5135c18595",
      "vendor_defaults_used": {
        "model": "openai/gpt-4.1",
        "tools": [],
        "voice": "vapi/Elliot",
        "stt_model": "soniox/stt-rt-v5",
        "version_id": "2026-07-30T23:36:18.688Z",
        "endpointing": {
          "wait_seconds": 0.1,
          "stop_num_words": 0,
          "on_number_seconds": 0.5,
          "stop_voice_seconds": 0.2,
          "stop_backoff_seconds": 1,
          "on_punctuation_seconds": 0.1,
          "custom_endpointing_rules": 0,
          "on_no_punctuation_seconds": 0.1,
          "smart_endpointing_enabled": null,
          "smart_endpointing_provider": null
        },
        "voice_speed": null,
        "assistant_id": "9d66d943-e554-42a7-ac08-6327680dade9",
        "stt_language": "en",
        "background_sound": null,
        "max_duration_secs": null,
        "model_temperature": 0,
        "endpointing_source": {
          "stop_speaking_plan": "vapi-documented-default",
          "start_speaking_plan": "api"
        },
        "first_message_mode": "assistant-speaks-first",
        "voicemail_detection": false,
        "user_idle_reply_secs": null,
        "background_speech_denoising": false,
        "first_message_interruptions_enabled": null
      },
      "vendor_unsupported": [
        "idle_reply_threshold"
      ],
      "honesty": [
        "The instrument is UNCHARACTERISED for this measurement path. The known-delay reference that produced our overhead and noise figures replied over a carrier media websocket, and that path no longer exists -- a characterisation is only valid for the host, carrier and audio path it was measured on, so no overhead figure is quoted here and none is subtracted (V1.md D4). Recording-path overhead is inside every number below, unmeasured.",
        "t1 -- the end of our own speech -- is found by a speech detector on the near channel, not by matching a known waveform. It is accurate to roughly a few tens of ms rather than sample-exact, so small differences between vendors are not resolvable; the per-turn error budget is the one Gate A measures on synthetic dialogs.",
        "Turn-taking is Plivo's speech endpointing, not ours. It decides when the conversation moves on, never when a measurement starts or ends: both endpoints of every TTFAB are read off the recording afterwards. Eager or slow endpointing therefore costs discarded turns, not shifted numbers.",
        "Our speech is Polly.Joanna text-to-speech reading a fixed script, not the human recording V1.md 11.1 requires (TTS endings trip endpointing differently). It is re-rendered per call rather than played from one committed file, so the stimulus is identical in wording across vendors but not byte-identical across calls; t1 is measured from each call's own tape, so this costs no accuracy.",
        "Answer accuracy is a keyword match against the carrier's transcript of the vendor's reply -- a coarse check that the agent answered from its prompt, reported alongside latency and never mixed into it.",
        "Historical, on the DELETED websocket path: a vendor's own recording of the same calls read 547 ms lower than ours (bench-telnyx-20260729-182644, n=12 turns) -- more than that path could account for. Read vendor-reported latency with care. Not re-measured on this path.",
        "n=99 usable call(s). V1.md wants n>=100 spread over >=5 days and 4 time-of-day buckets; this is a sample, not a distribution.",
        "The turn curve has as few as n=82 measurements at some turn indices. Pooled figures are the sturdier read; per-turn percentiles (especially p95) are directional until each index has many more.",
        "Discarded calls are reported above, never silently dropped -- a vendor that fails to respond is worse than one that is slightly slower."
      ]
    },
    {
      "rank": 5,
      "provider": "Retell AI",
      "provider_slug": "retell",
      "ttfab_onset_p50": 1740,
      "ttfab_onset_p90": 2104.8,
      "ttfab_onset_p95": 2258.6,
      "ttfab_onset_p99": 2867.8,
      "cost_per_minute": 0.135494,
      "cost_per_billed_minute": 0.134122,
      "cost_per_call": 0.105212,
      "cost_currency": "USD",
      "cost_calls_priced": 104,
      "cost_notes": [
        "combined_cost is reported in cents; divided by 100",
        "telephony excluded (0.0045 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0047 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0068 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0100 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0102 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0105 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0107 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0110 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0112 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0115 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0118 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0120 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0123 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0125 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0127 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0130 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0132 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one",
        "telephony excluded (0.0135 USD): the PSTN leg is the bench's own carrier cost, and no other platform's figure includes one"
      ],
      "consistency_p95_over_p50": 1.298,
      "turns_usable": 419,
      "turn_attempts": 430,
      "discards": {
        "double_talk": 1,
        "vad_disagree": 10
      },
      "barge_in_turns": 0,
      "carrier": "plivo",
      "latest_run_id": "bench-retell-20260731-152959",
      "latest_run_at": "2026-07-31T15:29:59+00:00",
      "analyzer_version": "2.4.0",
      "vendor_config_sha256": "3fbe40adfec2aeea3e8100517811faedc8837c2ac17564bfcc995281b1256e0a",
      "vendor_defaults_used": {
        "model": "gpt-4.1",
        "tools": [],
        "voice": "retell-Cimo",
        "llm_id": "llm_07a33cfd375e8e43ee628ff3195e",
        "agent_id": "agent_0ffe15ecf1e57f2503b34658ac",
        "language": "en-US",
        "endpointing": {
          "stt_mode": "custom",
          "responsiveness": null,
          "enable_backchannel": null,
          "vocab_specialization": null,
          "begin_message_delay_ms": 0,
          "interruption_sensitivity": null
        },
        "llm_version": 3,
        "voice_model": null,
        "voice_speed": null,
        "agent_version": 3,
        "ambient_sound": null,
        "start_speaker": "agent",
        "denoising_mode": null,
        "llm_last_modified": 1785454948707,
        "model_temperature": 0,
        "reminder_max_count": null,
        "agent_last_modified": 1785454948714,
        "model_high_priority": false,
        "reminder_trigger_ms": null,
        "max_call_duration_ms": 3600000,
        "end_call_after_silence_ms": null
      },
      "vendor_unsupported": [
        "stt_provider"
      ],
      "honesty": [
        "The instrument is UNCHARACTERISED for this measurement path. The known-delay reference that produced our overhead and noise figures replied over a carrier media websocket, and that path no longer exists -- a characterisation is only valid for the host, carrier and audio path it was measured on, so no overhead figure is quoted here and none is subtracted (V1.md D4). Recording-path overhead is inside every number below, unmeasured.",
        "t1 -- the end of our own speech -- is found by a speech detector on the near channel, not by matching a known waveform. It is accurate to roughly a few tens of ms rather than sample-exact, so small differences between vendors are not resolvable; the per-turn error budget is the one Gate A measures on synthetic dialogs.",
        "Turn-taking is Plivo's speech endpointing, not ours. It decides when the conversation moves on, never when a measurement starts or ends: both endpoints of every TTFAB are read off the recording afterwards. Eager or slow endpointing therefore costs discarded turns, not shifted numbers.",
        "Our speech is Polly.Joanna text-to-speech reading a fixed script, not the human recording V1.md 11.1 requires (TTS endings trip endpointing differently). It is re-rendered per call rather than played from one committed file, so the stimulus is identical in wording across vendors but not byte-identical across calls; t1 is measured from each call's own tape, so this costs no accuracy.",
        "Answer accuracy is a keyword match against the carrier's transcript of the vendor's reply -- a coarse check that the agent answered from its prompt, reported alongside latency and never mixed into it.",
        "Historical, on the DELETED websocket path: a vendor's own recording of the same calls read 547 ms lower than ours (bench-telnyx-20260729-182644, n=12 turns) -- more than that path could account for. Read vendor-reported latency with care. Not re-measured on this path.",
        "The turn curve has as few as n=91 measurements at some turn indices. Pooled figures are the sturdier read; per-turn percentiles (especially p95) are directional until each index has many more.",
        "Discarded calls are reported above, never silently dropped -- a vendor that fails to respond is worse than one that is slightly slower."
      ]
    }
  ],
  "methodology_url": "https://openbenchmarks.com/voice-agent-latency#methodology"
}