The Missing Corpus: Infinite Ground Truth for File-Grounded LLM Evaluation
Abstract
Evaluating enterprise LLMs on file-grounded Question Answering (QA) requires large, labeled corpora of realistic enterprise PDFs, yet real documents are sensitive, scarce, and impossible to annotate at scale. Existing synthetic approaches produce visually crude renders, annotate ground truth post-hoc (introducing hallucination risk), or are non-reproducible one-off scripts with no diversity control. We present SynthDocQA, an Intermediate Representation(IR)-driven pipeline that co-generates photorealistic multi-page enterprise PDFs and provenance-grounded evaluation pairs in a single pass. A compact Document Generation IR encodes topic, content schedule, and perturbation profile before rendering begins. Templates are retrieved per page via Contrastive Language-Image Pretraining (CLIP) embedding-based cross-modal search over a 10,000-layout library; a 15-stage physical degradation pipeline models five scan quality tiers from pristine to poor. Every QA assertion is derived analytically from the same structured artifact that populated the page, making hallucination structurally impossible, and a hierarchical fork-based Random Number Generator (RNG) guarantees exact reproducibility from a single integer seed. A reference 100-document corpus spans 50 enterprise domains, 4,374 pages, 16,589 QA pairs, and 21,402 assertions at zero annotation cost.