Verifier-Native Closed-Loop Repair for Operations Research Agents
Abstract
Generative-AI systems for operations research (OR) are often evaluated as one-shot formulation generators, although real modeling is iterative: solve, diagnose, repair, and verify. We study OR repair as a closed-loop decision process with exact solver and domain feedback. Our system compiles feedback into a typed Repair-Sufficient State (RSS) and uses verifier-state changes to rank interventions through a Counterfactual Verifier Loop (CVL). We evaluate fixed-trace diagnosis and live repair on a benchmark of 360 hardened cases. Paired ablations show that typed state is the stable contribution, while controller value is model-dependent: a probe-lite controller improves Qwen3-30B by 10.8 percentage points and GPT-5.4-nano by 17.2 points over RSS-only, but gains compress near ceiling and can vanish on other models. Focused controls isolate ranking, retrieval, output-interface, and interaction-budget effects. These findings motivate a deployment rule: expose typed verifier state by default and activate additional controller logic only after model-specific paired validation.