StylePlan: Style-Conditioned Intent Planning for Zero-Shot Coordination
Abstract
Zero-shot coordination requires agents to collaborate with previously unseen partners without test-time fine-tuning or explicit communication. A persistent difficulty is that partner behavior contains information at multiple timescales: stable conventions such as role preference or spatial habit, and fast local intent such as the next object to collect or deliver. Most partner-conditioned policies compress these signals into a single latent embedding, which can blur long-term convention with short-term action evidence. We propose \textbf{StylePlan}, a partner modeling framework that separates slow partner style from fast partner intent and combines them through conflict-aware arbitration. StylePlan builds an online behavioral fingerprint, predicts short-horizon partner intent, forms a style-conditioned intent prior, and uses the disagreement between prior and immediate evidence to modulate a recurrent actor-critic policy. On the verified Overcooked V2 artifacts, StylePlan obtains the best reported average cross-play among methods with available entries under the legacy evaluator. We additionally provide a scan-based JAX evaluation with unseen-partner XP, within-episode partner switches, unified SP/FCP reruns, and wired module ablations. The results show that StylePlan is most reliable as an online adaptation mechanism: it improves recovery and drop under partner switches, while the current implementation still needs broader unified reruns to establish a stronger XP claim.