Origami as a Real-Image Benchmark for Procedural State-Transition Reasoning
Abstract
General-purpose embodied AI systems are increasingly expected to interpret multimodal instructions and perform fine-grained physical tasks. However, many manipulation benchmarks provide limited support for jointly evaluating procedural correctness, local geometric sensitivity, and human-guided correction. We propose origami as a benchmark based on real images for procedural state-transition reasoning in embodied AI. Origami requires precise operations such as tucking, inserting, reversing, and folding overlapping layers, making it a focused domain for studying physical state transitions. We introduce a dataset of step-by-step folding images, textual entries, and difficulty annotations, and define two tasks: autonomous state transitions from multimodal instructions and human-guided collaborative transitions. Unlike crease-pattern- or diagram-based origami benchmarks, our setting exposes agents to actual intermediate folding states. For an initial simulator-based evaluation, we use a multimodal LLM controller to evaluate procedural understanding, action-interface grounding, and retry-based correction.