ASTOR: Benchmarking and Improving Autoformulation for Soft-Constraint Optimization with a Formal Grammar
Abstract
Decision-making problems in real-world settings are often presented in natural language. Operations Researchers have found useful ways to address these problems by formulating them as mathematical programs, and increasingly, by using Large Language Models to automate and assist their modeling. However, existing methods and benchmarks in automatic optimization formulation often neglect softer preferences proposed by consumers, planners, and managers, despite the fact that many practical problems require modeling them explicitly. We introduce ASTOR, an inference-time framework that maps a natural-language decision problem to a grammar-guided abstract syntax tree (AST) before deterministically compiling the resulting representation into executable optimization code. The AST explicitly separates hard constraints from soft constraints and is provably adequate. We also introduce SoftOR, a unified benchmark of natural language problems which focuses on soft constraint formulations. A range of experiments show that ASTOR achieves best aggregate performance among comparable LLM methods. Code and data are available at {https://github.com/amerlikesummer/astor-softor}.