ReaXpert: From Explanation to Intervention with Auditable Chemical Reasoning
Abstract
Scientific foundation models need more than plausible explanations: experts need to see what the system assumed, what evidence it used, and what changed when it was adapted. We study this requirement in reasoning about reaction yield. A frozen large language model (LLM) asked directly to classify Suzuki-Miyaura reactions as low, mid, or high yielding recovers only 7.9% of reactions in the low yield class, concealing a severe optimism bias behind a single verdict. ReaXpert replaces that verdict with six factors specified by chemists and grounded in retrieved industrial precedent, then adapts the instruction with Genetic-Pareto (GEPA), an evolutionary prompt optimizer that reflects on scored execution traces written in natural language. The resulting factor claims, evolved instruction, and calibration rules remain readable, providing surfaces on which a scientist can locate and contest an error without inspecting or updating model weights. Calibration based on attributed failures raises recall for the low yield class to 48.3% and reduces the spread across class recalls from 52.5 to 5.3 percentage points. The complete prompt pipeline comes within 0.4 percentage points of an open model adapted with low-rank adaptation (LoRA) on exact accuracy and exceeds it on macro-F1. Under a production shift with structures as the only inputs, it retains the best macro-F1 and class balance among the evaluated systems. Its factor reasoning also reaches useful yield thresholds in fewer Bayesian optimization evaluations than standard optimization or random search. A proxy audit finds 90% factor traceability but only 75.5% structure grounding, exposing fluent reasoning built on parsing errors. We therefore claim interpretability through readable scientific claims and adaptation artifacts, not faithfulness to the model's hidden computation or chemical explanations validated by experts.