Recovering Hidden Chemistry: Improving LLM Fine-Tuning on Metabolic Pathways
Yimeng Liu
Abstract
Tabular data encode much of biochemical knowledge, yet converting reaction tables into text for large language models (LLMs) often preserves only opaque enzyme and reaction IDs while discarding the chemistry underlying pathway reasoning: the intermediate metabolites connecting successive reactions. We study whether recovering this hidden chemistry improves LLM transfer learning on metabolic pathways. From raw stoichiometry, we reconstruct missing intermediates using a recursive backtracking algorithm that resolves reaction directionality under reversible reactions and shared cofactors, yielding directed metabolic graphs $G=(V,E)$ for 275 validated pathways. We linearize these graphs into Chain-of-Thought (CoT) traces and fine-tune Gemma-3 (1B to 27B) and DeepSeek-R1 (8B) with parameter-efficient LoRA, comparing against HTML-table, plain-text, and SMILES-based representations. Graph-derived CoT substantially improves Gemma-3 performance, increasing BLEU from $<0.2$ to $\sim0.5$ and ROUGE-1 to $>0.6$, with gains also observed for DeepSeek-R1. We further find that standard early stopping can underfit these repetitive, compositional data, highlighting the importance of training dynamics for biochemical fine-tuning. Overall, our results show that exposing implicit chemical relationships, rather than simply reformatting tabular data, improves LLM learning of metabolic pathways, suggesting that AI for science may benefit from representations that explicitly expose the structures underlying expert reasoning.
Chat is not available.
Successful Page Load