FactoryBench: Evaluating Industrial Machine Understanding
Abstract
We introduce FactoryBench, a benchmark for evaluating whether time-series models and LLMs hold an implicit world model of a physical system: an internal representation of its state and dynamics accurate enough to predict how the system evolves under an intervention and to reason about it counterfactually. Q&A pairs are organized along four levels, state estimation, action-conditioned forward prediction, counterfactual rollout, and decision-making, instantiating Pearl's ladder of causation, and span five answer formats: four are scored deterministically and free-form answers are scored by an LLM-as-judge voting protocol. We propose a scalable Q&A generation framework built around structured templates, release FactoryWave (a dense, multitask, multivariate sensor dataset from a UR3 cobot and a KUKA KR10 industrial arm), and construct FactoryBench as a benchmark of over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Zero-shot evaluation of six frontier LLMs shows that no model exceeds 50% (chance-corrected) on state, intervention, or counterfactual reasoning, and that the ranking reshuffles entirely on decision-making, with the leader on the first three levels falling to the bottom of the fourth: current models' implicit world models are far from adequate for physically grounded, closed-loop decision-making.