MINEGRID: A Multi-Modal Benchmark for LLMs Evaluation on Dynamic Modeling of Power Systems
Abstract
Dynamic modeling and simulation of modern power systems are inefficient and error-prone due to the manual translation required for complex differential-algebraic equations (DAEs) and the reliance on information in unstructured, multimodal documents. Vision-language model (VLM) capable of synthesizing multimodal information from diagrams has the potential to automate this modeling process; however, it lacks specialized datasets to ensure physical consistency and structural accuracy and systematic frameworks to evaluate the performance. This paper presents \textbf{MINEGRID}, a first-of-its-kind multimodal benchmark for automating the conversion of power systems diagrams into DAE systems. The benchmark includes a comprehensive dataset of control diagrams paired with expert-verified LaTeX equations, JSON metadata, and grid diagrams paired with topology annotations, both stratified by physical and structural complexity. This benchmark also evaluates the generation progress from multimodal structures to precise DAE terms using metrics covering structural correctness, algebraic accuracy, and physical consistency. In the experiments, the proposed benchmark provides a comprehensive performance analysis across multiple state-of-the-art models to establish a foundational baseline for automated power system modeling, bridging the gap between multimodal document understanding and automated modeling of power systems. The dataset and the code are available at https://anonymous.4open.science/r/Minegrid-Benchmark-4280.