SIGA: Scientific Simulation Coding Agent Adapter- A Geophysics Case Study
Abstract
Frontier LLMs are increasingly capable of expert-level scientific reasoning, but using them as reliable scientific agents requires more than general reasoning ability. Scientific simulation is a central test case: modern simulators are configured through executable interfaces such as XML decks, input scripts, and namelists that function as domain-specific languages tied to the simulator's internal API. Translating a researcher's natural-language intent into a runnable configuration remains a recurring expert bottleneck. We study how much of this bottleneck a general-purpose LLM coding harness, Claude Code, can absorb when wrapped with a Simulator-Interface Grounding Adapter (SIGA): a package of skills, tools, and workflow control flow that grounds the agent's outputs in simulator documentation, schema, and example libraries. We instantiate SIGA for GEOS, an open-source multiphysics simulator used in CO₂-storage and induced-seismicity research. A Resolution-IV factorial surfaces three benefits over vanilla Claude Code. First, SIGA improves reliability, reducing across-seed variance by roughly 40× by preventing unparseable or empty decks on a hard tail of compound multi-physics tasks. Second, it improves quality, raising mean structural similarity by about +7 percentage points on the same hard tail. Third, a self-evolved variant matches the best hand-designed cell with roughly 16% fewer tool calls, suggesting that automatic SIGA discovery is tractable. A preliminary human baseline finds that geoscience-domain-expert volunteers new to GEOS take between 8 and 36 times as long as the agent on a representative task. An explicit human-consultation tool is used in only about 3% of under-specified trials, with the agent instead relying on the on-disk example library as a cheaper retrieval substitute. A small OpenFOAM transfer study indicates that the recipe is not specific to GEOS XML: the same stop-hook component again dominates reliability, and the best SIGA cell outperforms both vanilla Claude Code and a constrained Foam-Agent lint-only baseline. We close with SIGA-design recommendations grounded in failure modes that our best configuration still does not fix.