An Error in the Context Is a Demonstration
Abstract
Small language models fail long agentic trajectories more often than large ones, and the usual explanation is per step competence: failures multiply independently over the horizon. We measure a different mechanism. An error that enters the context does not merely fail its own step. The model reads it as a demonstration and reapplies it, most strongly at the steps that play the same structural role. We inject controlled errors into teacher forced ledger audit trajectories across nine instruction tuned models from five families between 0.36B and 3.2B parameters, and score only later steps that are logically independent of the corrupted one, so that any degradation is contagion rather than task dependence. Three findings follow. First, contagion is flat across that ladder while per step competence spans a wide range, so scaling inside the small model regime buys competence but not immunity, and a model that executes the clean task perfectly is as susceptible as one that fails a third of its steps. Extending the ladder to twenty models, up to 9B and two newer generations, bends contagion down but does not end it: preregistered out of sample, the decline replicates in two families and reverses in a third, and of three families near 9B one is indistinguishable from zero and two are not, so the ceiling belongs to the model and not to the size. Second, the effect follows structural role rather than distance, sharpens with scale, and grows with repeated demonstrations of one wrong rule but not with an equal number of different wrong rules, which is rule induction rather than a corrupted context. Third, guardrails that detect but do not remove the error do nothing, and even redacting it in place leaves damage. At equal escalation budget no placement policy transfers across model pairs, but one rule does: the step kind carrying a driver’s reproduction number, measured under teacher forcing, predicts which single kind is worth protecting in free running rollouts, and protecting the other three is worth nothing. Error handling in the context, not only model choice, is a first class design axis for small model agents.