Host-Materialized State for Small Language Models: A Synthetic Paired Study
Abstract
Answering state queries from an event stream requires a model to replay updates and bind each value to the correct entity. We test a narrow host-side intervention for small language models: replace the full event transcript with query-sufficient current-state facts computed before inference. On 96 paired width-3 queries, Qwen3-8B improves from 85/96 to 96/96 exact answers, an absolute gain of 11.46 percentage points. A supporting Ministral-3-8B run improves from 88/96 to 95/96 (+7.29 points). Secondary empirical-task bootstrap 95% confidence intervals are [5.21, 17.71] and [2.08, 13.54] percentage points, respectively. Because materialization both reconstructs state and compacts the prompt, these results estimate the bundled intervention rather than either component alone. They apply only to the retained tasks and model revisions; they do not establish a universal memory limit, an internal binding mechanism, or a tool-observation-boundary effect.