Can a Knowledge-Graph Agent Close the Gap Between a Low-Cost and a Frontier-Class LLM? An Audited Upstream Petroleum Case Study
Abstract
Small, low-cost language models are attractive agent backbones but are assumed to trail frontier models on tasks requiring exact identifiers, deterministic aggregation, and multi-hop reasoning. We test how close a knowledge-graph (KG) agent brings them. One ReAct pipeline – a structured graph-query tool, a deterministic aggregation tool, and a citation-traceability guardrail – runs a 50-question upstream-petroleum benchmark with a low-cost backbone (Gemini 3.1 Flash-Lite) and a frontier-class one (Claude Opus 4.8), three repeats each. Normalized exact match (EM) gives 0.898 vs. 0.889; a symmetric hand audit of all 216 EM answers gives 0.926 vs. 0.954, and the backbones differ on only 3 of 36 items, so this benchmark cannot separate them. The audit also shows that most strict-EM failures of both backbones are correct answers (82% and 81%), that grading rules derived from one backbone missed a format of the other, and that 7 of the 50 items are defective. An ablation ladder places the gain in the graph plus deterministic tools: graph retrieval without agency scores 0.278, below Text-RAG at 0.583. Since gold answers are computed from the graph the tools query, we read this as both backbones driving one toolchain to similar accuracy, not as a claim about model size.