DocOCR-Eval Selects Best OCR Tools Without Ground Truth

Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding· July 21, 2026 View original

Summary

DocOCR-Eval is an annotation-free framework for automatically assessing and selecting optimal OCR engines and MLLMs for document parsing, even without ground-truth labels. It uses a three-stage correction and ranking strategy, demonstrating that aggregating results from multiple MLLMs improves alignment with annotation-based rankings, providing practical guidance for real-world document collections.

Document parsing is a foundational step for understanding unstructured scanned documents, converting them into structured data. With a multitude of Optical Character Recognition (OCR) engines and multimodal large language models (MLLMs) available, selecting the most suitable solution for a specific document collection can be challenging, especially when manual annotations (ground truth) are scarce or non-existent. This research introduces DocOCR-Eval, an innovative framework designed to evaluate and select OCR tools without requiring ground-truth labels. The framework employs a three-staged approach involving correction and ranking strategies to approximate the performance ordering typically achieved with annotations. Experiments show that by aggregating results from multiple MLLMs, DocOCR-Eval progressively improves its alignment with ground-truth-based rankings, offering a reliable and practical method for deploying document parsing systems across diverse, label-limited real-world document collections.

Why it matters

Many organizations deal with vast amounts of scanned documents and lack the resources for extensive manual annotation. DocOCR-Eval provides a crucial, cost-effective method to select the best OCR solution, improving the accuracy of downstream document understanding tasks.

How to implement this in your domain

  1. 1Pilot DocOCR-Eval to select the most effective OCR engine for a specific document collection where ground truth is unavailable.
  2. 2Integrate the three-staged correction and ranking strategy into your document processing pipeline for automated OCR assessment.
  3. 3Leverage multiple MLLMs to enhance the accuracy of OCR tool selection in label-scarce environments.
  4. 4Apply this framework to optimize document understanding workflows, reducing manual effort and improving data extraction quality.

Who benefits

BFSILegalHealthcareGovernmentData Management

Key takeaways

  • Selecting the right OCR tool is challenging, especially without ground-truth labels.
  • DocOCR-Eval is an annotation-free framework for OCR assessment and selection.
  • It uses a three-stage correction and ranking strategy to approximate ground-truth rankings.
  • Aggregating multiple MLLMs improves the reliability of OCR tool selection.

Original post by Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding

"arXiv:2607.16203v1 Announce Type: new Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting te…"

View on X

Originally posted by Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau, Wei Liu, Sirui Li, Yifan Peng, Yihao Ding on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses