Best LLM inference provider for tool-calling agents
What the benchmark found. Baseten led end-to-end p99 latency at 3.23 seconds; Telnyx and Fireworks AI led task success at 99.5%. The scorecard keeps latency, tokens per second, accuracy, and failures separate rather than blending them into one synthetic score. For tool-calling agents the complete structured result is the finish line; time to first token matters only when text is shown to a person first. Measured across 10 LLM inference providers serving GLM 5.3 Flash, 600 requests each, Sep 3, 2026.
Providers measured. Baseten, DeepInfra, Fireworks AI, Modal, Nebius, Novita AI, Parasail, Telnyx, Together AI, and Z.AI, each serving GLM 5.3 Flash behind an OpenAI-compatible chat-completions endpoint.
What this page compares. A tool-calling agent cannot proceed until the model returns a complete, schema-valid object it can pass to the next tool. This page applies the full scorecard to that step: complete-answer latency and exact task success, with failures kept separate.
How to read the result. Set a task-success floor first, because a plausible but wrong argument triggers the wrong action. Then compare end-to-end p95 and p99 among the providers that clear it; sequential tool calls compound tail latency.
Workload and configuration. Every provider received the same 600 deliberately easy structured tasks: 200 meeting-note lookups, 200 ticket-triage classifications, and 200 contract-term extractions. GLM 5.3 Flash ran with streaming enabled, reasoning low, temperature 0, top-p 1, a 256-token output ceiling, one attempt, a 20-second timeout, and per-provider concurrency one.
APIs used. The runner called each provider's OpenAI-compatible chat-completions surface with the model identifiers below; the links open the provider documentation used to verify each adapter.
- Baseten:
zai-org/GLM-5.3-Flash· API docs ↗ - Modal:
zai-org/GLM-5.3-Flash· API docs ↗ - Telnyx:
zai-org/GLM-5.3-Flash· API docs ↗ - Novita AI:
zai-org/glm-5.3-flash· API docs ↗ - Z.AI:
glm-5.3-flash· API docs ↗ - Fireworks AI:
accounts/fireworks/models/glm-5p3-flash· API docs ↗ - Parasail:
zai-org/GLM-5.3-Flash· API docs ↗ - Together AI:
zai-org/GLM-5.3-Flash· API docs ↗ - Nebius:
zai-org/GLM-5.3-Flash· API docs ↗ - DeepInfra:
zai-org/GLM-5.3-Flash· API docs ↗
Which LLM inference provider is best for tool-calling agents on latency and exact results?
The full table is retained: a tool-calling step can be blocked by complete-result latency, wrong structured output or an API failure, and no synthetic score hides those.
| Provider | End-to-end latency | Time to first output token | Time to first visible answer token | Token generation speed | Task success | Failure rate | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P50 | P95 | P99 | P50 | P95 | P99 | P50 | P95 | P99 | P50 | P95 | P99 | |||
| Basetenzai-org/GLM-5.3-Flash ↗ | 331 ms | 2.29 s | 3.23 s | 228 ms | 499 ms | 946 ms | 287 ms | 1.16 s | 2.65 s | 240.5 tok/s | 61.6 tok/s | 22.6 tok/s | 95.8% | 0.0%no operational failures |
| DeepInfrazai-org/GLM-5.3-Flash ↗ | 1.83 s | 11.0 s | 17.1 s | 808 ms | 3.78 s | 9.08 s | 1.15 s | 6.41 s | 13.2 s | 69.4 tok/s | 9.0 tok/s | 5.0 tok/s | 95.5% | 4.0%429 0.3% · transport 3.7% |
| Fireworks AIaccounts/fireworks/models/glm-5p3-flash ↗ | 1.95 s | 6.72 s | 10.5 s | 586 ms | 5.67 s | 8.74 s | 1.06 s | 5.83 s | 8.74 s | 58.2 tok/s | 27.9 tok/s | 14.9 tok/s | 99.5% | 0.0%no operational failures |
| Modalzai-org/GLM-5.3-Flash ↗ | 574 ms | 1.90 s | 3.28 s | 376 ms | 812 ms | 2.72 s | 470 ms | 1.09 s | 2.79 s | 222.2 tok/s | 84.7 tok/s | 34.9 tok/s | 99.3% | 0.0%no operational failures |
| Nebiuszai-org/GLM-5.3-Flash ↗ | 1.45 s | 7.45 s | 15.0 s | 925 ms | 7.11 s | 14.6 s | 1.14 s | 7.11 s | 14.6 s | 269.3 tok/s | 78.8 tok/s | 39.9 tok/s | 84.8% | 14.8%timeout 2.8% · other 4xx 11.5% · transport 0.5% |
| Novita AIzai-org/glm-5.3-flash ↗ | 1.74 s | 5.00 s | 7.25 s | 1.13 s | 1.79 s | 4.28 s | 1.40 s | 3.41 s | 5.58 s | 60.6 tok/s | 29.6 tok/s | 11.6 tok/s | 94.2% | 4.0%timeout 0.2% · 429 3.8% |
| Parasailzai-org/GLM-5.3-Flash ↗ | 1.59 s | 7.74 s | 11.5 s | 856 ms | 4.96 s | 10.9 s | 1.11 s | 5.54 s | 10.9 s | 52.4 tok/s | 22.8 tok/s | 15.6 tok/s | 98.3% | 1.3%timeout 0.8% · 429 0.2% · transport 0.3% |
| Telnyxzai-org/GLM-5.3-Flash ↗ | 669 ms | 2.33 s | 3.48 s | 497 ms | 802 ms | 2.90 s | 574 ms | 1.34 s | 3.00 s | 184.3 tok/s | 67.5 tok/s | 47.3 tok/s | 99.5% | 0.0%no operational failures |
| Together AIzai-org/GLM-5.3-Flash ↗ | 609 ms | 3.66 s | 12.4 s | 377 ms | 2.23 s | 12.2 s | 393 ms | 2.26 s | 12.2 s | 151.6 tok/s | 34.0 tok/s | 21.7 tok/s | 98.5% | 1.0%timeout 0.3% · 429 0.5% · 5xx 0.2% |
| Z.AIglm-5.3-flash ↗ | 1.75 s | 4.83 s | 8.59 s | 1.05 s | 1.89 s | 4.76 s | 1.39 s | 3.14 s | 5.00 s | 56.3 tok/s | 27.8 tok/s | 11.5 tok/s | 97.5% | 0.2%timeout 0.2% |
Rows are listed alphabetically. Generation-speed P95 and P99 are coverage-based slow-tail floors: 95% and 99% of measured responses respectively generated at least that fast.
Provider by provider, in the order above
- Baseten: 331 ms median, 3.23 seconds end-to-end p99, time to first token p95 1.16 seconds, 240.5 tokens per second, 95.8% task success, 0.0% failures.
- Modal: 574 ms median, 3.28 seconds end-to-end p99, time to first token p95 1.09 seconds, 222.2 tokens per second, 99.3% task success, 0.0% failures.
- Telnyx: 669 ms median, 3.48 seconds end-to-end p99, time to first token p95 1.34 seconds, 184.3 tokens per second, 99.5% task success, 0.0% failures.
- Novita AI: 1.74 seconds median, 7.25 seconds end-to-end p99, time to first token p95 3.41 seconds, 60.6 tokens per second, 94.2% task success, 4.0% failures.
- Z.AI: 1.75 seconds median, 8.59 seconds end-to-end p99, time to first token p95 3.14 seconds, 56.3 tokens per second, 97.5% task success, 0.2% failures.
- Fireworks AI: 1.95 seconds median, 10.49 seconds end-to-end p99, time to first token p95 5.83 seconds, 58.2 tokens per second, 99.5% task success, 0.0% failures.
- Parasail: 1.59 seconds median, 11.51 seconds end-to-end p99, time to first token p95 5.54 seconds, 52.4 tokens per second, 98.3% task success, 1.3% failures.
- Together AI: 609 ms median, 12.45 seconds end-to-end p99, time to first token p95 2.26 seconds, 151.6 tokens per second, 98.5% task success, 1.0% failures.
- Nebius: 1.45 seconds median, 14.97 seconds end-to-end p99, time to first token p95 7.11 seconds, 269.3 tokens per second, 84.8% task success, 14.8% failures.
- DeepInfra: 1.83 seconds median, 17.07 seconds end-to-end p99, time to first token p95 6.41 seconds, 69.4 tokens per second, 95.5% task success, 4.0% failures.
Which tool-calling patterns does this benchmark represent?
- Tool selection. Ticket triage approximates choosing a category and priority before dispatching the next action.
- Typed tool arguments. Contract extraction approximates producing a typed object that deterministic code must accept without repair.
- Fact before action. Meeting-note lookup approximates retrieving one grounded fact that controls the next call.
How is LLM inference provider latency and accuracy measured here?
- Same model. Every endpoint serves GLM 5.3 Flash; the provider layer is the only variable.
- Same request. One-turn, non-tool chat completion in JSON-object mode, streaming on, temperature 0, top-p 1, reasoning low, 256-token ceiling, 20-second timeout, concurrency one.
- Client clock, pinned tokenizer. Time to first token, end-to-end latency and tokens per second come from the client; visible answer tokens are counted with a pinned
tiktoken o200k_basetokenizer, never provider-reported usage. - Exact task success. An answer counts only when it is schema-valid and exactly correct under the task rubric. Timeouts, HTTP errors and malformed responses stay in the denominator.
- Serverless endpoints only. Dedicated capacity is outside the comparison. Runner and per-request records: openbenchmarks-labs/inference ↗.
Tool-calling inference: questions this page answers
Which inference provider is fastest for tool-calling agents at p99?
Baseten led end-to-end p99 latency at 3.23 seconds; Telnyx and Fireworks AI led task success at 99.5%. The scorecard keeps latency, tokens per second, accuracy, and failures separate rather than blending them into one synthetic score. For tool-calling agents the complete structured result is the finish line; time to first token matters only when text is shown to a person first.
Which provider returns the most exact tool arguments?
Set a task-success floor first, because a plausible but wrong argument triggers the wrong action. Then compare end-to-end p95 and p99 among the providers that clear it; sequential tool calls compound tail latency.
Should a tool-calling agent optimize time to first token or end-to-end latency?
The full table is retained: a tool-calling step can be blocked by complete-result latency, wrong structured output or an API failure, and no synthetic score hides those.
Same run, other questions
- LLM inference provider benchmark: complete scorecard and methodology →
- Best LLM inference provider for AI agents: tool calling latency and accuracy →
- Fastest LLM inference provider for AI agents →
- Most accurate LLM inference provider for AI agents →
- Best LLM inference API for chatbots and copilots: time to first token and tokens per second →
- Best LLM inference provider for chat applications →
- Best LLM inference provider: speed, accuracy and failures compared →
Read the complete LLM inference provider benchmark
Metric definitions, request scheduling, task rubrics, failure handling, and the complete provider scorecard live on the LLM Inference Provider Benchmark →




