PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. However, a fundamental question remains unresolved: can LLMs reason about physical constraints—the rules governing how abstract computations must map onto real-world hardware? This capability is distinct from, and arguably more demanding than, software generation alone, as it requires grounding language understanding in spatial, electrical, and topological constraints. To address this gap, we introduce PCEval (Physical Computing Evaluation), a benchmark and evaluation protocol for physical computing education that enables fully automatic, execution-based assessment of LLMs across both logical and physical implementation artifacts, without human grading. Our evaluation framework assesses LLMs in generating circuits and producing compatible code across varying levels of project complexity. Through comprehensive testing of 14 leading models, PCEval provides the first reproducible and automatically validated empirical assessment of LLMs' ability to reason about fundamental hardware implementation constraints within a simulation environment. Our findings reveal that while LLMs perform well in code generation and logical circuit design, they struggle significantly with physical breadboard layout creation, particularly in managing proper pin connections and avoiding circuit errors. PCEval advances our understanding of AI assistance in hardware-dependent computing environments and establishes a foundation for developing more effective tools to support physical computing education.