KARL: Kairotic Training for Efficient Agentic Reinforcement Learning System
Abstract
Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents with reinforcement learning requires long, heterogeneous trajectories, leaving learner engines idle in synchronous systems until the slowest trajectory finishes. To shrink these pipeline bubbles, asynchronous training overlaps rollout and learning across updates, but comes with the pitfall of policy staleness. We introduce KARL, a system that overlaps these stages with zero policy staleness through kairotic execution: gradient computation starts as soon as its required inputs are fixed. For group relative policy optimization (GRPO), KARL computes each trajectory's score gradient as soon as its reward arrives, without waiting for the group. For on-policy distillation (OPD), it computes gradients for each completed agentic turn's teacher-scored actions while tool calls run in the sandbox. KARL leaves the original objectives intact, and we prove mathematical equivalence to their synchronous updates for the formulations studied here. On SWE-bench Verified and Terminal Bench 4.0, KARL trains models to the same performance up to 1.9 \times faster than synchronous training. With zero policy staleness, KARL also outperforms asynchronous training in final performance under the same GPU-hours.