SeismoAgentBench: Environment-Grounded Verification of Long-Horizon Scientific Agents
Abstract
Interactive scientific agents transform natural-language goals into multi-stage tool trajectories, simulator runs, artifacts, and final reports. Evaluating only the final answer obscures whether an agent selected the correct environment, satisfied intermediate obligations, recovered from execution failures, or merely produced a plausible completion claim. We present SeismoAgentBench, a runnable benchmark and evaluator for trajectory-level analysis of seismic-simulation agents. The release contains 316 tasks across three open environments—OpenSeesPy, OpenQuake, and ANUGA—organized into capability, controlled out-of-distribution robustness, and reliability/recovery layers. Each task is specified by a scenario contract defining required trajectory evidence, scientific acceptance checks, artifacts, budgets, and permitted recovery actions. The evaluator combines deterministic schema and artifact checks, simulator-specific scientific checks, recovery semantics, and a 24-task expert-adjudicated judgment slice. Across five architecture-family controls, the strongest reference system achieves task-success rates of 0.941, 0.781, and 0.722 across the three layers. All controls degrade on private hidden evaluation, while limited real-mode reruns reveal additional environment discrepancies. SeismoAgentBench is intended to diagnose how tool-using agents succeed and fail over extended trajectories, rather than to establish deployment readiness or a single measure of general agent capability.