FactoryBench: Evaluating Industrial Machine Understanding
Abstract
Industrial robots emit dense multivariate telemetry that determines their operating state, and acting on a fault requires both reading that state and knowing the procedure that resolves it. These are separate abilities, and they are usually measured separately: perception benchmarks stop at the reading, while control benchmarks score the resulting action without isolating which of the two failed. We introduce FactoryBench, a benchmark that measures both on the same episodes from real industrial robots, so that physical perception and physically grounded decision-making can be told apart. Q&A pairs are organized along four levels, physical state estimation, intervention, counterfactual reasoning, and engineering decision-making, and span five answer formats: four are scored deterministically and free-form answers by an LLM-as-judge voting protocol. We release FactoryWave (a dense, multitask, multivariate sensor dataset from a UR3 cobot and a KUKA KR10 industrial arm, with counterfactual episodes recorded on hardware by re-executing a task under held-fixed conditions to approximate the do-operation) and build FactoryBench as over 70k Q&A items grounded in roughly 15k normalized episodes from FactoryWave, AURSAD, and voraus-AD. Evaluating six frontier models, none exceeds 50% (chance-corrected) on the perception-side levels, and the ranking reshuffles entirely at the decision level: the leader on the first three levels falls to the bottom of the fourth. Linear probes show this is not one deficit but two, recovering fault presence from a frozen model's activations at 0.83 where its own answer is below chance (a read-out failure), while fault type is no more decodable than from an untrained network (a representational one). Physical understanding and physically grounded decision-making are dissociable, and current models fail at both in different ways.