Surfacing Jagged Intelligence via Agentic Search for Weak-Model Wins
Nicholas Di ⋅ Brian Bartoldson ⋅ Chloe Georgiou ⋅ Hongjun Choi ⋅ Jaywon Koo ⋅ Christine Klymko ⋅ Shusen Liu ⋅ Bhavya Kailkhura ⋅ Ruben Glatt
Abstract
Benchmarks coarsely depict model capabilities and can hide a model's jagged intelligence -- systematic failures on tasks similar to the benchmark's. We aim to automate the creation of datasets that characterize this jaggedness by inverting the roles of the strong and weak models in an AutoData loop: a challenger LLM modifies verifiable questions, and we accept its proposals if a weaker model succeeds and a stronger model fails. A task the weaker model solves is not inherently hard, suggesting the generated dataset reveals jaggedness in the stronger model's intelligence rather than lack of general capability. We demonstrate our approach's effectiveness on synthetic instruction-following tasks involving retrieval and math, a synthetic code tracing task, and via modifications to problems from a real coding benchmark (CRUXEval-O). Analysis of the generated datasets reveals triggers that cause the performance drops, including susceptibility to misleading context, index operations, and empty loops. Introducing these triggers to the original tasks causes the stronger model's performance to drop by $46-93$ percentage points, while the weaker model's performance fluctuates by less than 5 points.
Chat is not available.
Successful Page Load