← Insights

Do document agents still need an OCR pipeline?

New research suggests multimodal models extract business documents about as well from the image alone as with OCR. What that changes for appeals, claims and onboarding packs.

  • Documents
  • Agents

For a decade, document automation meant a pipeline: scan, OCR, layout analysis, field detection, rules, and a person to fix the rest. Each stage had its own vendor and its own failure mode.

A paper published this month, OCR or Not? Rethinking Document Information Extraction in the MLLMs Era, tested that assumption on large real-world business document sets. Its finding: given image-only input, capable multimodal language models perform about as well as OCR-enhanced pipelines, suggesting OCR may not be necessary. Careful schema design, examples and instructions improved results further.

OCR got better too

This is not only a story about bigger models. Dedicated document models also moved fast. A year ago, Mistral released Mistral OCR with throughput of up to 2,000 pages per minute on a single node and pricing of 1,000 pages per dollar.

For us the practical conclusion is not "OCR is dead." It is that the expensive part of document automation has moved.

Where the work is now

In our document work (denial letters, medical records, appeal packets, insurance applications), the hard part is no longer reading the page. It is:

  • The schema. What exactly should be extracted, in what format, with which definitions? The paper's emphasis on schema design matches our experience: most extraction errors are really specification errors.
  • Evidence. Every extracted field should point to where it came from on the page, so a reviewer can check it in seconds.
  • Cross-document reasoning. An appeal needs facts from several documents checked against a policy. That is an agent's job, not an extractor's.
  • Evaluation. A labelled set of real documents, scored field by field, run on every change.

How we decide the architecture

  • Clean, high-volume, standard forms: a dedicated OCR model plus validation rules is often cheapest.
  • Messy, varied documents: a multimodal model on the page image, with a strict schema.
  • Decisions across documents: an agent that calls extraction as a tool, then reasons over the results.

The right answer is often all three, routed by document type. The evaluation set decides.

Sources

  1. OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets, Shen et al., arXiv, March 3, 2026.
  2. Mistral OCR, Mistral AI, March 6, 2025.