When Are Recursive Language Models Useful? A Cost-Aware, Task-Conditional Evaluation
Abstract
Recursive Language Models (RLMs) handle long inputs through controller-generated programs and additional model calls. Published RLM evaluations typically compare these scaffolds with generic long-context or coding baselines. Deployment requires a different comparison: whether adaptivity improves on routes suited to the task under a common information, scoring and budget contract. We evaluate the released standard RLM against deterministic operators, retrieval, semantic workers and direct inference in an operation-matched audit. On controlled operations and known-family Oolong aggregation, matched routes generally match or outperform standard RLM at lower cost. A source-transfer replication preserves this direction although its clustered exact-rate difference includes zero. Open multi-document search is less stable. An initial untouched-shard replication favors RLM over BM25 but neither a prospectively locked equal-cap comparison nor a post-hoc same-row repeat yields an RLM contrast that survives multiplicity correction. Route ordering changes across these rollouts while RLM remains less reliable and 5.8–15.9× costlier than the stronger alternatives. Mechanism probes localize errors to operation specificity, output-schema fit, evidence selection and scaffold reliability. These results support reporting recursive inference as a task-conditional accuracy–cost frontier rather than as a single scaffold ranking.