Best PDF Parser APIs for Redlined Contracts, 2026
Claude Fable 5.1 leads 8 document parsing providers at 89.1% accuracy on redlined contracts. LlamaParse, Reducto, Datalab and 4 more benchmarked, 2026.
What this page compares. A redlined-contract parser must preserve not just the words on the page but which terms were deleted, retained, or replaced. This benchmark tests whether that edit history survives PDF parsing well enough for a downstream reader to answer from the negotiated agreement rather than superseded language.
How to read the result. The table ranks complete runs by downstream answer accuracy. Gap closed places each result between the markup-blind baseline and the best-case parse, so raw accuracy is not mistaken for parser capability alone.
Workload and controls. Each measured configuration processed the same 94 tagged contract PDFs. The same downstream reader then answered the same 1,500 questions. A best-case parse and a markup-blind baseline bound the score.
Ranked by downstream answer accuracy
The table ranks complete runs by downstream answer accuracy. Gap closed places each result between the markup-blind baseline and the best-case parse, so raw accuracy is not mistaken for parser capability alone.
| Rank | Provider | Type | Accuracy | Gap closed | Median / P95 latency | Parser $ / 1k pages | Parser $ / 1k correct | Total tokens / 1k correct | List price |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1claude-fable-5-1 · pdf in · bedrock | Frontier models (VLMs) | 89.1% | 95.5% | 327s / 461s | $109.39 | $187.18 | 21.3M | $11 / $55 per 1M tokens |
| 2 | LlamaParse agentic plustier=agentic_plus · 2026-09-24 | Long running agentic parsers | 83.4% | 86.5% | 91s / 148s | $56.25 | $102.79 | 24.3M | $0.05625 / page |
| 3 | LlamaParse agentictier=agentic · 2026-09-24 | Specialised document parsers | 80.0% | 81.1% | 52s / 77s | $12.50 | $23.81 | 24.8M | $0.0125 / page |
| 4 | Reductomodel=r-1 · 2026-09-25 build | Specialised document parsers | 78.7% | 79.1% | 7s / 10s | $10.00 | $19.36 | 28.3M | $0.010 / page |
| 5 | Datalab track changestrack-changes endpoint | Specialised document parsers | 78.6% | 78.9% | 34s / 45s | $10.00 | $19.39 | 25.6M | $0.006 / page |
| 6 | Extendengine=parse_performance | Specialised document parsers | 74.8% | 72.9% | 28s / 41s | $25.00 | $50.94 | 27.3M | $0.025 / page |
| 7 | GPT-6 Astragpt-6-astra · pdf in | Frontier models (VLMs) | 63.3% | 54.6% | 228s / 309s | $51.00 | $122.85 | 30.8M | $10 / $50 per 1M tokens |
| 8 | Datalab convertconvert · mode=accurate | Specialised document parsers | 62.4% | 53.2% | 23s / 55s | $10.00 | $24.42 | 30.7M | $0.010 / page |
| 9 | Pulsemodel=pulse-ultra-2 | Specialised document parsers | 52.7% | 37.8% | 43s / 98s | $15.00 | $43.41 | 81.5M | $0.015 / page |
| 10 | Mistral OCRmodel=mistral-ocr-4-1 | Specialised document parsers | 43.4% | 23.1% | 11s / 21s | $4.00 | $14.04 | 43.3M | $0.004 / page |
How document-processing accuracy, latency, and cost were measured
This is the compact protocol. Corpus construction, all nine task types, parser settings, scoring rules, and cost treatment are documented in the full benchmark methodology →
- Same corpus. 94 legal contracts and 1,500 questions were reused for every measured configuration; no vendor received an easier document set.
- One changing layer. Every arm received the complete PDF and returned Markdown. The downstream reader, question, prompt, and scoring path stayed fixed, so the parser output was the variable under test.
- Independent answer key. Correct and stale answers were derived from tracked changes in the source Word files before vendor output was produced. The residual-answer judge was not shown the parser identity.
- Bounded accuracy. The best-case reference preserves every deletion; the markup-blind baseline removes every mark. Gap closed reports where each parser landed between those controls.
- Separate operational metrics. Parse latency and measured parser spend are reported beside semantic accuracy, never blended into a synthetic score. Published list price remains separate from measured corpus spend.
- Reproducible evidence. The benchmark runner and published data are open in openbenchmarks-labs/document-processing ↗.
What this ranking establishes
Semantic accuracy
A correct answer follows the language the parties agreed. A stale answer follows language they struck, while a fused answer combines deleted and surviving text into a value that never existed.
Independent controls
The best-case reference preserves every deletion. The baseline removes every mark. Gap closed places each measured system between those controls instead of pretending 100% is always reachable.
Scope
A redlined-contract parser must preserve not just the words on the page but which terms were deleted, retained, or replaced. This benchmark tests whether that edit history survives PDF parsing well enough for a downstream reader to answer from the negotiated agreement rather than superseded language.
Questions answered by this comparison
Which provider leads best pdf parser apis for redlined contracts?
Claude Fable 5.1 leads this table at 89.1% accuracy. The result comes from 1,500 questions across 94 contracts.
What does document-parsing accuracy mean here?
Accuracy is the share of questions answered from the operative contract language. The markup-blind baseline scored 28.8%. Stale rate separately counts answers taken from deleted language.
How should latency and cost be read beside accuracy?
Claude Fable 5.1 recorded 327s median parse time and $109.39 in measured parse cost per 1,000 pages. The full board also divides parse cost and the answering agent's tokens by correct answers, which shows what a parser costs for every 1,000 right answers it leads to.
How was this document-processing comparison independently benchmarked?
Each system converted the same full PDFs to Markdown. The same reader model then answered the same hidden-ground-truth questions; vendor identity was not shown to the residual-answer judge.
What documents does this benchmark cover?
A redlined-contract parser must preserve not just the words on the page but which terms were deleted, retained, or replaced. This benchmark tests whether that edit history survives PDF parsing well enough for a downstream reader to answer from the negotiated agreement rather than superseded language.
Related document-processing pages
- Document Processing Benchmark: complete leaderboard and methodology →
- Best Document Parsing APIs for Redlined Contracts
- Best Document Parser Tools for Redlined Contracts
- Best OCR APIs for Redlined Contracts
- Best PDF Data Extraction APIs for Redlined Contracts
- Best Document Parsing APIs
- Best OCR APIs for Scanned Documents
Read the complete Document Processing Benchmark
Task design, corpus transformation, parser settings, scoring, controls, cost accounting, and both tagged-PDF and scanned-PDF leaderboards live on the Document Processing Benchmark →





