STAGE: Source-grounded Text-to-JSON Artifact GEneration
Abstract
From financial filings to clinical records, legacy industries rely heavily on long, unstructured documents to store high-value information. Reliably extracting this information into structured, machine-readable representations is a key prerequisite to making the contents accessible to automated systems. Yet, constructing reliable and scalable text-to-JSON training data remains challenging. To address this gap, we propose STAGE (Source-grounded Text-to-JSON Artifact GEneration), a source-grounded data generation pipeline that constructs reports and JSON/schema by using LLMs for scalable synthesis while validating ground-truth values against the underlying spreadsheet. Evaluations on STAGE-Eval, our source-grounded benchmark with an 851-example test set, show that STAGE produces stronger training data than existing approaches. This improves Qwen3-4B exact match rate from 38.03% to 74.27% and value accuracy from 68.83% to 90.69%.