Meaning-Preserving Rephrasings Can Break LLM-Generated Optimization Models
Jessica Udomsrirungruang ⋅ Anirudh Subramanyam
Abstract
Large language models (LLMs) are increasingly being used to solve optimization problems. Agents can now generate formulations from natural language (NL) problem descriptions, call solvers, and return a solution. Because the formulation step is implicit, the NL phrasing can produce unnoticeable errors. Existing LLM evaluation frameworks only score one fixed phrasing per instance and do not test whether semantically equivalent statements are answered identically. To address this gap, we introduce a notion called answer stability, together with two families of meaning-preserving perturbation: local phrase substitutions and global whole-statement paraphrases. On the OptiBench benchmark, we find that at lease one perturbation causes $\\texttt{gemma4:26b}$ to fail on $55.5\\%$ of the instances it originally solved. To understand these failures, we inspect the LLM-generated code and classify each failure by: the stage where the failure happens, the component that is wrong, the type of failure, and its consequence. We find that the single most common cause of failure is a dropped integrality constraint. We also check if correct answers always come from correct formulations. Across both perturbations, only $82\\%$ of value-correct phrasings produce formulations isomorphic to that generated from the original wording, whereas several others silently relax integrality yet still return the reference optimum.
Chat is not available.
Successful Page Load