MinerU2.5-Pro: Pushing the Limit of Document Parsing via a Calibrated Evaluation-Data Flywheel
Abstract
Document parsing is now critical infrastructure for LLM systems, yet standard benchmarks suggest near saturation: on OmniDocBench v1.5, leading methods cluster around a 95-point ceiling. We argue that this reflects a stalled evaluation–training flywheel, not solved parsing: v1.5 under-samples hard, long-tail documents and uses fixed-granularity matching that can penalize semantically correct outputs. We re-engage the flywheel with two co-designed protocols. OmniDocBench v1.6 repairs evaluation with MGAM (Multi-Granularity Adaptive Matching) and a strictly held-out 296-page Hard Subset, with Base / Hard / Full reporting. The corrected benchmark exposes two data gaps—hard cases are under-represented and unreliably annotated—which the MinerU2.5-Pro Data Engine addresses through DDAS (Difficulty-Driven Adaptive Sampling) and Judge-and-Refine, a render–verify–refine annotation pipeline, yielding 65.5M auto-labeled and 192K human-annotated samples. Conversely, the same discovery and annotation mechanisms provide a scalable path for future Hard-subset refreshes. Under a fixed 1.2B MinerU2.5 backbone, these interventions raise OmniDocBench v1.6 Full performance from 92.98 to 95.69, with the clearest separation on the corrected Hard split. The result reframes apparent saturation as a repairable benchmark–data loop: document parsing progress must be easier to measure correctly, not just easier to claim.