Harness Engineering Can Make Small Language Models Explore Like Large Language Models
Abstract
On data science tasks, most AI agents follow one analytical direction unless a user prompts them. This paper introduces Axum, a harness that gives one language model separate roles. An orchestrator plans and delegates, workers execute the analysis and report back, and a goal-holding agent outside the execution context alone can declare the analysis complete and stop the run. The number of concurrent workers is the main control. The paper also introduces ExploreBench, which measures how much of the human analytical move space a system covers, that is, the analytical directions a run examines against those that human analysts examined. With one small model, Gemma 4 31B, Axum covers 51.9\% of that move space against 30.2\% for the best single-agent harness from Claude Code, Codex, OpenCode and Pi, and the same small model through Axum matches the best harness on a near-frontier model with about 14 times the parameters, at 5.1 times lower estimated cost. That comparison changes the model and the harness together, so we report it as descriptive. Coverage rises from 33.2\% with one worker to 64.7\% with fifteen, while it starts to flatten after three. Because coverage is a union over moves, repeated sampling is a demanding control: five combined runs of the best single-agent harness reach 41.9\% at the token cost of Axum at five workers, where Axum reaches 56.1\%. Harness engineering can thus be a primary control on exploration breadth, and part of its gain is not only more compute.