VERLA: Verified Lattice Decoding for Reliable Small Language Model Agents
Abstract
Small language models (SLMs) are attractive for repetitive agentic workloads, but a single malformed or semantically wrong tool call can terminate a multi-step workflow. We introduce VERLA (Verified Lattice Decoding for Agents), a training-free runtime that converts typed tool schemas and current state into a finite action lattice, ranks state-valid calls by an SLM’s conditional likelihood, and transactionally tests the top-K candidates against executable postconditions with rollback. We also introduce QUEUETOOL, a controlled five-domain benchmark with 16 tools and 2–5 step workflows designed to isolate action selection, state validity, composition, paraphrase shift, and distractor-state robustness. Across two independently trained 258K-parameter conditional action language models, greedy decoding averages 39.50%, 22.00%, 4.63%, and 31.38% workflow success across four splits. VERLA with K = 5 reaches 99.88%, 99.50%, 54.25%, and 99.88%, respectively, while requiring 1.32–2.37 attempted calls per reached step and 2.5–3.0 ms of batched candidate scoring on CPU. The result is deliberately scoped: exact executable postconditions, two controlled micro-SLM checkpoints, and a synthetic benchmark are strong assumptions. Within that regime, the study shows that allocating a small verification budget to the harness can recover reliability that the model alone does not possess.