Talk Like a Molecular Graph: Chemical Representations for Large Language Models
Abstract
Large Language Models (LLMs) are increasingly being used for chemistry tasks such as reaction prediction and structure elucidation. To complete these tasks, LLMs must reliably reason about molecular graph structures. A molecule can be written in several standard text notations, which differ in syntactic form but encode the same structure. Most previous studies of LLMs in chemistry have used SMILES strings or IUPAC names as molecular representations; however, the suitability of these formats has not been systematically assessed. In this work, we introduce MolJSON, a novel JSON schema for representing molecular graphs, and compare it with five standard chemical formats. We evaluated each representation with GPT-5-nano, GPT-5-mini, GPT-5, and Claude Haiku 4.5 using 78,045 questions spanning translation, shortest path, and constrained generation reasoning tasks. These tasks assess graph-reasoning performance without requiring specialised chemistry knowledge, allowing the impact of different molecular representations to be isolated. On translation tasks, GPT-5 achieved 71.0% accuracy when converting IUPAC names to MolJSON, compared with 43.7% when converting the same inputs to SMILES. Similar improvements were observed for constrained generation (95.3% versus 64.0% for SMILES) and shortest-path reasoning (98.5% versus 92.2%), where MolJSON also used fewer reasoning tokens. MolJSON outperformed SMILES and IUPAC, despite these formats being ubiquitous in training corpora, and despite the LLMs receiving no MolJSON-specific training. We observed systematic errors for SMILES and IUPAC associated with molecular size and the nesting of fused ring systems, whereas MolJSON was more robust to these failure modes. Our results show that the choice of molecular representation has a material impact on LLM performance, and suggest that explicit graph schemas, such as MolJSON, are intrinsically better suited to current LLMs than traversal- or nomenclature-based encodings.