Finding Blind Spots in AppWorld and WorkArena Task Verifiers
Abstract
Execution-based task verifiers decide whether an agent succeeded. Leaderboards, model comparisons, and verifier-reward training inherit their errors. We audit the shipped AppWorld and WorkArena verifiers with source-informed, mutation-inspired tests. The main audit never modified a shipped checker. AppWorld has a concrete blind spot: duplicating a non-idempotent write preserves every value the evaluator checks while creating an extra record. The verifier accepted all three task variants from each of two affected scenario generators, or 2/5 eligible generators and 6/15 constructed effects. On copies specified after that Extra-NI census, a documented len(added_*)==N patch flips all six extra-row cells to FAIL. Table 2 still reports the unmodified shipped tests. WorkArena has a second blind spot. We prospectively fixed and reran the 23 earlier extra-field candidates with independent Table API readback before checker evaluation. The Table API confirmed a nondefault persisted value in 21 of the 23 selected reruns, and the shipped checker returned PASS in all 23. Those 23 cells were chosen because they had already passed. Two requested strings were display-label aliases of their stored defaults. The 21 confirmed wrong effects span three form templates. This confirms a mechanism; it does not estimate a rate. No other construction produced an independently confirmed false accept. The remaining checker-PASS cells were effect-correct degeneracies. Retained evidence is not uniform across the zero-PASS families, so we report their verdict vectors separately and do not pool them. In fixed intent-swap grids, the checkers returned no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells used a session-scoped evidence path, and the other 2,632 remain unclassified by rejection cause and independent target ground truth. Each increment was specified before its own cells were scored. An anonymous supplement accompanies the submission with the construction grammar, primary evidence, content-bound stage lineage, and count reproducer.