When Does Monitoring Long-Horizon Agents Lead to Better Performance?
Abstract
Long-horizon language-model agents can experience performance degradation as their interaction histories grow and context accumulates in their agent state [1]. We study how well different monitoring methods predict agent degradation, and whether compacting or reconstructing agent state improves downstream task performance. Across conversational reasoning tasks (Evolving-Intent GSM8K [2]) and multi-turn tool-use tasks (BFCL Multi-Turn [3]), we compare “active” observational monitors, “passive” monitors, and baseline context- and time-based approaches. We find that monitoring is not a free observation: active probes can alter the trajectories they are intended to measure, producing an “observer effect” where task success tends to decrease. We further find that the value of monitoring can depend on how agent state is recovered. In a controlled study, compaction and deterministic reconstruction provide similar gains without a carried probe (+5.9 and +6.3 percentage points). Carrying the probe reduces the gain under compaction to +2.8 points, while the gain under reconstruction remains +6.5 points. These results suggest a taxonomy in which observation method, signal quality, and recovery mechanism are evaluated jointly, rather than treating monitoring as a standalone prediction problem.