Passers, Partial Solvers, and Early Quitters: Evaluating Skill-Augmented Agents Beyond Pass Rate
Abstract
Agent skills package reusable procedural knowledge for LLM agents. Their evaluation, however, still rests almost entirely on terminal pass rate, which collapses very different failure trajectories into one binary outcome and leaves execution cost unmeasured. We complement pass rate with two metrics computed on a controlled SkillsBench subset. Partial Credit comes from deterministic verifier groups or task-native rewards rather than an LLM judge; Token Consumption normalizes provider-specific logs onto a common basis. We evaluate three harness–model configurations on eight verifier-decomposable Data Analysis tasks under No Skills and Curated Skills, for 240 trajectories. The three metrics disagree. Claude Code / Claude Sonnet 4.5 leads on Pass Rate (37.5%), Codex CLI / GPT-5.2 on Partial Credit (0.705), and Gemini CLI / Gemini 3 Flash on token efficiency. Partial Credit also separates trajectories that Pass Rate scores identically, and curated skills can reduce intermediate progress while terminal success stays flat. Skill benchmarks should report terminal success, deterministic partial progress, and execution cost together.