Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality
Abstract
For CodeAct-style agents fine-tuned on execution traces, an evaluation score reflects task skill and compatibility with the runtime demonstrated in training. Interactive-agent evaluations usually hold the harness fixed, which leaves those two hard to separate. That compatibility is partly learned: some runtimes preserve the agent's interpreter variables across turns, while others reset them after every action. We study runtime transfer: whether an agent trained under one runtime still solves the task under the other. On three tasks that require a working record built from tool feedback and reused in later actions, we generate instance-paired training traces under persistent and stateless runtimes, fine-tune a separate Qwen3-8B agent for each task and runtime, and evaluate all four combinations of training and deployment runtime. We then vary an evaluation setting that is commonly fixed: the per-turn tool-call cap. The cap determines whether mismatch appears as reconstruction overhead or as quality failure. On the Opaque Knapsack task, where the agent reveals hidden item attributes before selecting a capacity-constrained subset, both appear. When the persistent-trained agent runs on a mismatched stateless runtime under a slack cap (80 calls per turn, which its action blocks rarely reach), it retains much of the quality of matched deployment (0.61 against 0.77), but it uses 4.4x as many tokens rebuilding lost state. Under a binding cap (25, which routinely truncates them) it collapses to 0.07, while matched deployment remains at 0.66. Across the tasks and controls studied, quality loss caused specifically by the runtime mismatch appears only when the cap truncates an intended action and matched execution can resume from surviving state. That dependence on the cap replicates across training seeds, an independent second rollout and two additional base models. Evaluations should therefore report and vary the runtime contract, the per-turn cap and the episode horizon together. Otherwise the same mismatch may appear as an efficiency cost under one configuration and as a catastrophic failure under another.