MORPH-DA: A Mutation-Grounded Benchmark for Metamorphic Verification of Data-Analysis Agents
Prateek Kohli ⋅ Karun Thankachan
Abstract
Data-analysis agents increasingly generate and execute Python programs, but successful execution does not establish that the program implemented the requested filters, aggregation, grouping, denominator, or date window. We introduce MORPH-DA, a mutation-grounded benchmark for evaluating runtime verifiers of data-analysis agents, together with a specification-conditioned metamorphic reference verifier for wrong-but-executable programs. The reference verifier keeps each generated program $p$ fixed, applies operator-conditioned transformations to the underlying tabular environment, and checks necessary output relations $R(p(D), p(T(D)))$ without using a precomputed numeric gold answer or an additional model call. The benchmark contains 101 compositional tasks and 563 validated, non-equivalent semantic mutants. MORPH-DA detects 364/563 faults (64.7%, 95% task-clustered CI [58.4, 70.5]) versus 9/563 for universal robustness checks ($p=9.4\times10^{-79}$). Re-executing each seed-42 agent program unchanged on held-out data seeds reveals that 13--16 programs per model (20--26% of apparently correct programs) are accidental corrects: correct on the source seed but wrong on at least one held-out seed. With cross-seed evaluation labels, MORPH-DA obtains 79--89% precision and 67--79% recall on executable programs, with 8--21% false-positive rate. These results position metamorphic execution as a complementary, zero-additional-LLM-call verifier rather than a correctness certificate.
Chat is not available.
Successful Page Load