Beyond Pessimism and Diversity: What Critic Ensembles Regularize in Offline-to-Online RL
Abstract
RLPD obtains strong offline-to-online performance with a large N-head critic ensemble and min-over-M target subsetting, yet why the ensemble helps is unresolved. The standard pessimism and head-diversity accounts leave a puzzle that neither resolves: dropout substitutes for ensemble members at small N but adds nothing at large N. We propose that the operative property is the action-space smoothness of the ensemble-mean critic, the function the actor's gradient actually climbs. We measure it as a Q-scale-normalized sharpness, the variance of Qbar(s, a+eps) under small action perturbations, normalized by |Qbar|^2. Across a multi-seed grid on sparse Adroit manipulation (pen, door; N in {2,4,6,10}, dropout p in {0,0.01}, with min-free M=1 ablations), dropout reduces sharpness and improves score only at small N, and the benefit decays monotonically in N along a saturation curve, reconciling DroQ's small-ensemble gains with the absence of a dropout benefit at large N, while the min-free failure is an amplitude event: |Qbar| inflates by ~1009x with no accompanying rise in normalized sharpness, which separates pessimism from geometry. We then add a pre-registered TD3-style target-policy-smoothing arm, an action-space knob that adds no ensemble diversity, to test the central prediction of dose-dependent gains at small N. It did not reproduce dropout's effect, and because it moved measured sharpness only marginally we report it as inconclusive. Finally, across non-divergent configurations sharpness is associated with final score (config-level rho_s = -0.74, permutation p = 0.045, n = 8 configs), a weak early signal that is not yet a practical selection rule.