Adaptive Action Chunking via Multi-Chunk Q Value Estimation
Abstract
Action chunking emerged as a pivotal technique in imitation learning, enabling policies to predict cohesive action sequences rather than single actions. Recently, this approach has expanded to reinforcement learning (RL), enhancing behavioral consistency and reducing bootstrapping errors in value function estimation. However, existing methods rely on a fixed chunk length, creating a performance bottleneck as the optimal length varies across states and tasks. In this paper, we propose Adaptive Action CHunking (ACH), a novel offline-to-online RL algorithm that dynamically modulates chunk length during both training and inference. For each state, ACH takes the action-value of each candidate chunk length as its optimality criterion. We estimate the action-values of all candidate chunk lengths simultaneously in a single forward pass by employing a causal Transformer-based critic. Based on these values, the agent adaptively selects the most effective chunk length for the current state. Evaluated on 34 challenging tasks, ACH consistently outperforms fixed-length baselines, demonstrating superior generalization and learning efficiency in complex environments.