From Skill Representations to Coaching Decisions: A Skill-Conditioned Policy for Real-Time Instruction
Abstract
Pretrained representations of people can support personalization by encoding stable individual characteristics. Once such a representation conditions an agent’s actions, however, its consequences extend beyond prediction: the represented person becomes the recipient of decisions shaped by the representation, including its errors. We therefore ask whether a pretrained human representation is actionable: which decisions it influences, whether those influences align with human behavior, and how misalignment between the representation and the person propagates through interaction. Our setting is real-time verbal coaching for high-performance driving, where an agent must combine a driver’s immediate behavior with persistent skill to decide when and what to instruct. We develop a hierarchical coaching policy conditioned on a skill embedding learned independently of coaching data and frozen during policy training. In participant-held-out evaluation, skill conditioning provided little improvement in moment-level instruction-onset discrimination, but reduced participant-level coaching-rate error by 46% and improved instruction-category prediction over both a no-skill ablation and a scalar lap-time baseline. To identify which decisions depend on the representation, we substitute embeddings from other drivers while holding driving context fixed. These interventions change coaching frequency and three of six instruction categories, while other categories remain unchanged despite exhibiting skill-related patterns in held-out predictions, showing that observational agreement alone does not establish that a policy uses a representation. Finally, in a 36-participant within-subject study, participants experienced coaching conditioned on a matched representation, a deliberately mismatched representation, and no skill representation. Relative to mismatched conditioning, matched conditioning produced greater above-chance compliance, greater end-of-block lap-time improvement, and better perceived skill calibration. Relative to no skill representation, its clearest advantage was perceived calibration; compliance and lap-time improvement favored the matched condition but did not reach significance. Pretrained human representations should therefore be evaluated not only by what they encode, but by which agent decisions they influence and how representational errors affect the people who experience those decisions.