Instruction Following with Composable Constraints
Abstract
Modern instruction-following benchmarks score language models by whether they satisfy long, flat lists of verifiable constraints, such as length, formatting, and lexical requirements. However, in real-world scenarios, instruction complexity often arises not only from the number of constraints, but from how constraints are layered and composed. We study this complexity with composable constraint trees, a depth- and width-controlled representation whose nodes are verified by deterministic code. We introduce a framework that generates such trees, renders them as natural-language instructions, and supports new constraint families through a typed slot signature and verifier. Using this framework, we build ComposeIF, a benchmark of 400 structurally unique, source-grounded, human-validated depth-3/4 instances. ComposeIF remains challenging for frontier and open-weight models, with the strongest evaluated model reaching 55.25%. The same pipeline produces 105,161 verifier-filtered training examples without additional annotation; a single LoRA finetune of Qwen3-8B improves ComposeIF by +12.75 points and transfers to COLLIE and IFBench by +25.16 and +11.09 points, respectively.