Hardware-Aware Resource Management for Fully Asynchronous Agentic RL with Small Language Models
Abstract
Fully asynchronous reinforcement learning (RL) overlaps rollout and policy updates, yet substantial producer–consumer imbalance can remain in disaggregated execution. Large-scale systems address related imbalance through heterogeneous pools, elastic scheduling, or resource-ratio tuning. We study a complementary setting—Qwen3-4B agentic RL on a fixed two-GPU allocation within a homogeneous AMD MI350X node—and exploit large HBM, multi-process device residency, and AMD GPU's DPX hardware partitioning. Temporal packing crosses the rollout and trainer roles of two concurrent, independent RL jobs, improving post-warm-up sequence throughput per physical GPU by 15.1% relative to a token-normalized serial baseline. DPX right-sizing keeps rollout on a full GPU and assigns training to one half-GPU partition; over a matched update prefix, job throughput does not decrease, while logical capacity-normalized throughput improves by 36.8% when the released half is reusable. Taken together, these results suggest that hardware-aware placement and partitioning can increase throughput per unit of GPU capacity, offering a path toward lower infrastructure costs for fully asynchronous agentic RL with small language models.