Can We Trust AI Evaluation? Robustness, Causality, and Risk in Modern AI Assessment
Abstract
AI capabilities are advancing rapidly, yet our ability to evaluate these systems has not kept pace. Benchmarks, leaderboards, and aggregate metrics increasingly influence decisions about model selection, deployment, regulation, and investment, but growing evidence shows that evaluation conclusions can be fragile, misleading, or insufficient for real-world use. Performance may vary under small evaluation changes, benchmark reuse can induce overfitting, contamination can distort comparisons, and offline metrics often provide limited evidence about safety, reliability, and deployment outcomes. As foundation models, AI agents, and autonomous systems are increasingly deployed in high-stakes settings, evaluation is becoming a critical bottleneck for trustworthy AI adoption. This workshop is motivated by a central question: When is AI evaluation evidence strong enough to guide deployment decisions? We argue that the next decade of AI research requires a science of AI evaluation – a research agenda that studies evaluation protocols themselves, not only the models being evaluated. The goal is to develop principled foundations for determining what an evaluation measures, what assumptions it relies on, what uncertainty remains, and when its conclusions can be trusted. The workshop will bring together researchers and practitioners from machine learning, statistics, causal inference, robustness, AI safety, and high-risk application domains to address three interconnected challenges: (1) uncertainty and robustness of evaluation evidence, (2) benchmark and leaderboard auditing, and (3) deployment risk and decision relevance. Through invited talks, contributed papers, posters, and panel discussions, the workshop aims to catalyze a research community around trustworthy AI evaluation and help establish the methodological foundations needed for reliable, transparent, and accountable deployment of increasingly capable AI systems.