Why OCR Quietly Breaks Document AI
OCR sits at the front of every document pipeline. When it misreads a table or a total, every retrieval step and every answer downstream inherits that error.
Most teams treat OCR as solved plumbing and spend their attention on the model at the end of the chain. That ordering gets the failure mode backwards.
Before a model can answer a question about an invoice or summarise a filing, some engine has to turn the page into text. The quality of that text sets a ceiling on everything after it. A pipeline that spends its evaluation budget on the model while trusting extraction by default is optimising the part least likely to be the actual bottleneck.
The pipeline runs in one direction
document
↓
extraction / OCR
↓
text and structure
↓
indexing and retrieval
↓
context
↓
model
An error introduced at the first step does not stay contained there. It is preserved through indexing, it shapes what gets retrieved, and it gets amplified once a model reasons over it confidently. Retrieval cannot find a figure that was merged into the wrong cell. No prompt recovers a total that was never read correctly.
Document
Extraction, or OCR
An error introduced here does not stay contained
Text and structure
Indexing and retrieval
The error is preserved through indexing, and shapes what gets retrieved
Context
Model
Reasons over it confidently
The errors also land badly. They concentrate on the tokens that carry the most meaning: totals, dates, labels, and the relationships between cells in a table. Those are exactly the tokens a downstream answer leans on hardest, which is why a modest-looking accuracy score can still produce a badly wrong answer.
If your evaluation scores the final answer but never scores the extraction step on its own, you cannot tell whether a bad answer came from reasoning or from a page that was already garbled before the model saw it.
Why no single engine wins
It is tempting to pick whichever engine tops a public leaderboard and assume the ranking holds for your documents. It usually does not. Engines specialise. Some are tuned for clean, predictable layouts such as invoices and fixed-template forms, and lead comfortably there. The same engines can fall several places on messy, mixed-layout material, while an engine that trailed on the easy case wins the hard one.
What one benchmark found
Newtuple published an OCR benchmark comparing PaddleOCR, Docling, LlamaParse and Surya. It graded 1,600 outputs across four test sets: two invoice runs, a set of financial tables, and mixed documents drawn from the OHR-Bench corpus. Outputs were scored from 0 to 1 by a model acting as grader against fixed ground truth. One invoice run shipped without verdict labels, so its verdicts were derived from score bands calibrated against the labelled runs.
On invoice run A, Docling led at 0.98, while on the mixed set LlamaParse and PaddleOCR tied at 0.79 and Docling fell to 0.69. On financial tables Docling and PaddleOCR tied at 0.81. Within that benchmark no engine led every category, and the engine that led the easiest category was not the one that led the hardest.
The authors are explicit about the limits: absolute scores are most reliable within a single test set, cross-set comparison is relative ranking rather than a like-for-like measure, and the mapping from these datasets to any particular industry is indicative rather than a controlled sample. Their own recommendation is to evaluate on your own documents.
Disclosure: I work at Newtuple and contributed to the work referenced here.
That result describes four engines on those four test sets. It is not a statement about OCR accuracy in general, and it is not independent validation of the argument on this page. What it does support is narrow and still useful: the ranking moved with document type, on a real graded set, which is the thing worth designing around.
Route by document type
If the ranking depends on the material, the response is routing rather than standardising. Match each document profile to the engine that handles it best, with a fallback for when the primary engine errors, times out, or returns low-confidence output.
| Document profile | What tends to work best | Why |
|---|---|---|
| Fixed templates, invoices | Structured-document specialists | Predictable layout rewards engines tuned for table and field alignment |
| Mixed, unpredictable material | Generalist engines with a validation step | Flexibility matters more than peak accuracy on any one layout |
| Unknown or varied pipeline | The most consistent all-rounder, as a default | A flat accuracy curve across types beats a high peak with a steep drop |
This table is design reasoning, not a measured result. Treat it as a starting hypothesis to test against your own documents.
Re-benchmarking as a habit
A benchmark result is evidence with an expiry date, not a permanent ranking. Engines and models change from one release to the next, and a comparison run on someone else's documents tells you less than the same comparison run on your own. Treat re-checking as a recurring operational habit tied to your document mix, not a one-time procurement decision.
The format the extraction stage produces is itself a contract the rest of the pipeline depends on, and contracts need re-verifying as their sources change. That argument is made properly in structured output and why it matters.
What this changes
Stop scoring only the final answer. Score the step that decides what the model is allowed to see.
Keep a small set of your own documents with known-correct values, and re-run extraction against them when anything upstream changes. That set is worth more than any public ranking, because it is made of the material you actually process.
For what happens to this text once it is extracted, see retrieval-augmented generation in plain terms.
Where this helps, and where it stops
- What it is
- An explainer of why document extraction is a distinct reliability boundary upstream of retrieval, and what follows from that for pipeline design.
- What it does not guarantee
- A ranking of OCR engines, a benchmark of its own, or a claim about industry-wide OCR accuracy.
- When the distinction matters
- A document question returns a wrong number and the model is the first thing blamed.