Agent-Conditioned Routing: Diverse Reasoning Paths Within One Frozen Model
Ananth Kalyanasundaram ⋅ Ankush Kadu ⋅ Aswanth Krishnan
Abstract
Generating several solutions is useful only when the additional trajectories explore meaningful alternatives rather than paraphrase one dominant approach. We investigate whether the conditional capacity of a frozen sparse mixture-of-experts (MoE) model can provide such diversity without method prompts, fine-tuning, or model copies. Agent-Conditioned Routing (ACR) assigns four parallel trajectories persistent, untrained shared/private expert-eligibility masks while preserving the learned, input-dependent token router. Because full-trajectory masking often causes loops or premature stops, our practical configuration applies ACR only to a 2,000-token method-selection prefix and then restores unrestricted base-model routing for execution. We evaluate the resulting system at three scales and read every reported derivation in full. To evaluate method diversity directly, we also construct a 126-question multi-mode mathematics and physics benchmark. On an 11-problem development set, ACR reaches ${\approx}0.93$ soft correctness with 2.00 distinct methods per four trajectories, versus 0.95 and 1.27 for the strongest tested temperature setting. The larger 126-problem study yields a more qualified conclusion. At the two high-correctness settings, a full method-diversity re-read of the ACR outputs finds 2.29 and 2.10 distinct methods among correct trajectories per four-trajectory pool, at 0.95 and 0.97 soft correctness. At the same correctness levels, temperature sampling reaches only 1.03 and 1.05 methods, respectively. The union over three pools contains 2.85 methods per problem, with 90\% of problems admitting at least two. Thus, on this screened benchmark and at matched generation budget, ACR roughly doubles UCM@4 relative to temperature sampling at the two matched high-correctness operating points. ACR stores one frozen backbone and about 160~KiB of mask state for four identities. In our measured implementation, latency and VRAM are within 3\% of batched temperature sampling.
Chat is not available.
Successful Page Load