Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era
Abstract
Over the past two years, agentic AI has moved from research demos to broad deployment in coding, clinical decision support, and conversational settings, yet the community's evaluation toolkit, built largely for static benchmark performance, has not kept pace. A growing body of peer-reviewed evidence documents a persistent gap between benchmark results and deployed reality, with evaluations dominated by technical metrics while human-centered, safety, and economic dimensions remain peripheral. This workshop, the second edition of the Human-AI Coevolution (HAIC) series, builds a methodological foundation for the empirical evaluation of human-agent teams as they coevolve with the people who use them. We organize the program around three interlocking bottlenecks, each anchored in recent peer-reviewed work: the validity of evaluation in deployed contexts, where measurement targets move and core constructs are contested; expert disagreement and the limits of human feedback, where annotator disagreement may signal genuine domain pluralism rather than noise; and adaptive testing and continual evaluation for systems that drift as humans and agents coevolve. Grounded in high-stakes domains including healthcare, mental health, aviation, and finance, the workshop solicits new methods, benchmarks, datasets, critiques, and case studies.