Harness Recession: When Small Models Outgrow Their Scaffolding
Abstract
Small language models (SLMs) often need extra support, such as examples, plans, tool instructions, or controllers, to work reliably as agents. However, the same support does not help every model. In this work, we study when such guidance is useful and when it can get in the way. We introduce a receiver-relative decision model in which guidance changes the actions favored by a base policy. The model yields a simple criterion: guidance helps when it corrects errors the receiver is likely to make. This motivates harness recession: the realized value of a persistently imperfect, receiver-conditioned guidance channel contracts as the receiving model improves and can eventually become negative. We test the account at three resolutions. Controlled policies verify the decision criterion; a three receiver open-model intervention measures bare, weak-, and full-guidance action distributions and recovers the predicted sign on 51 of 59 local identified shifts; paired task experiments vary receiver capability at fixed protocol and guidance quality at fixed receiver. Across arithmetic reasoning and interactive tool use, scaffolding gives large gains to weaker models, smaller gains to stronger models, and negative effects for some specialized models; improving the demonstrator raises protocol value for the same receiver. A harness therefore has no fixed value on its own. Its usefulness depends jointly on the receiving model and the guidance source, and should be re-evaluated whenever either changes