PCFBench: How Far Can Large Vision-Language Models Go in Physics-aware Photonic Inverse Design?
Abstract
Photonic crystal fibres (PCFs) underpin critical applications in telecommunications, sensing, and high-power laser delivery, yet their design remains a manual, expert-intensive process that couples geometric intuition with deep waveguide physics. Vision-language models (VLMs) offer a promising path toward automating this pipeline, but can they truly reason about photonic physics from microscopy images, or do they merely recognise visual patterns? We introduce PCFBench, the first comprehensive benchmark for evaluating VLMs on the complete photonic inverse-design workflow. PCFBench spans 26 tasks organised into six categories of increasing cognitive demand (Geometry Perception, Physics Understanding, Text Generation, Multi-modal Reasoning, Inverse Design, and Code Generation), built from over 220K FDTD simulation samples across 18 structurally diverse fibre families. We evaluate both individual VLMs on isolated tasks and multi-agent systems (MAS) on the full sequential pipeline from perception to design. Our results reveal a striking perception-design gap: current models can identify fibre geometry with moderate success, yet consistently fail to translate that visual understanding into physics-aware design actions. Moreover, cascading agents through the pipeline exposes significant error propagation, where upstream perception mistakes compound into downstream design failures. These findings suggest that neither model scaling nor naive agent chaining suffices; closing the loop demands explicit physical grounding. PCFBench provides the community with a rigorous, multi-dimensional testbed to measure progress toward AI-assisted photonic engineering.