What Does an Assembly Monitor Need to Remember?
Abstract
An assembly monitor may remember which actions occurred, their frequencies, or their full order. More detailed memory preserves more distinctions, but does the monitor use them? We audit these three interfaces using original human Assembly101 action labels and human instruction graphs. Counts separate most observed label conflicts hidden by a set of actions. On 944 held-out targets, however, retaining order gives no clear gain for a learned feature model and has opposing effects on two fixed language-model readouts. A third readout largely collapses to one label. Unrestricted outputs and a confidence-threshold check expose further limits that aggregate scores miss. The finding is practical: a richer symbolic trace does not by itself make a better monitor. Evaluation should check both what the input distinguishes and how the classifier reaches its verdict. This is an audit of annotated execution traces, not a test of physical-state estimation or robot control.