Opt-Arena: Evaluating, Selecting, and Generating Optimization Modeling Data via Tripartite Graphs
Abstract
Optimization modeling is the cornerstone of Operations Research (OR). While large language models (LLMs) show promise in autoformulation, existing datasets suffer from unquantified redundancies and domain imbalances. Consequently, models overfit to frequently occurring patterns and fail on underrepresented constraints, inducing a structural capability bias. To address this, we propose Opt-Arena, a data-centric framework that formally represents optimization datasets as tripartite graphs. To map this topology, we introduce Atomic Modeling Information (AMI)—the minimal semantic-mathematical rules of optimization. This representation enables the quantification of structural coverage, difficulty, and overlap. Guided by these metrics, our core-set selection mechanism distills high-density, low-redundancy training data. To further resolve intrinsic domain absences, we extend this framework with a kernel-driven generation pipeline, leveraging foundational problem backbones to synthesize structurally feasible instances for underrepresented areas. Empirically, this approach exhibits exceptional sample efficiency: utilizing only 5k curated instances, it surpasses the 20k-budget performance plateau of conventional distance- and uncertainty-based baselines. Relying entirely on standard SFT driven by our data, lightweight empirical probes outperform both massive proprietary models and complex reasoning architectures enhanced by reinforcement learning. These results reveal that LLM formulation performance is fundamentally constrained by structural data diversity, positioning topology-aware curation as an important driver for advancing OR modeling capabilities.