SimCAE-Bench: Evaluating Graph Representations for Language-Model Understanding of FEA Results
Abstract
Finite-element simulations often enter engineering review through post-processed stress and displacement summaries. Because these summaries remain tied to modeling assumptions, discretization choices, boundary conditions, and missing validation information, a language model must read both the computed relations and the limits of the available evidence. We study this problem for completed static FEA outputs: how should they be represented when the model is asked to answer review questions after the solve? To make the task measurable, SimCAE-Bench pairs 381 SimJEB-derived samples with 32 post-processing review questions selected through simulation-engineer screening. We then introduce graph-based relation-readout, which converts each completed FEA record into a question-agnostic CAE review graph over load cases, response quantities, rankings, margins, stress-displacement correspondence, and evidence boundaries. With DeepSeek Flash, this representation reaches 72.2% answer accuracy, outperforming the evaluated text and graph baselines while using fewer tokens than field-topology and Mapper-style graph serializations. Code and data are available in an anonymous repository: https://anonymous.4open.science/r/simcae-bench-anonymous-artifact-7452/