Small Language Models as Autonomous Operational Agents
Ahmed Elmokashfi ⋅ Usama Ahmed
Abstract
Small language models are usually cast as executors inside agentic systems: they call tools, follow procedures, and defer anything hard to a frontier model. We ask whether an 8B model can instead hold the \emph{planning} role in a production incident-response pipeline, for a domain absent from public pretraining corpora. Our method combines continued pretraining over roughly 100{,}000 internal domain documents with Reinforcement Learning with Verifiable Rewards (RLVR) via GRPO against a reference-grounded judge, preceded by a diagnostic establishing that RL is the right instrument at all: across 337 incident reports the domain-adapted model's deficit tracks how much it must infer rather than what it knows, and its shortfalls are failures of precision rather than direction. Trained on 9{,}400 expert trajectories and evaluated as the planner in a live workflow over 160 paired incidents, the model raises mean end-to-end quality from 5.72 to 6.53 out of 10 ($p = 1.1\times10^{-5}$) and lifts the share of incidents clearing the actionable bar from 37\% to 60\%. A planner ablation over 561 incidents postdating the training cutoff shows the 8B model, invoked once with retrieval withheld, is statistically indistinguishable from a retrieval-augmented frontier planner on cause identification and next-action accuracy; stratification localizes its residual deficit to one operation --- resolving an aggregate descriptor into the elements it currently contains --- which is dynamic state that does not belong in model weights. Our run settles at a KL divergence of ${\sim}0.10$ without degenerating.
Chat is not available.
Successful Page Load