Reliability for Deployed Analytics Agents: One-Hop Retrieval, Deterministic Computation, and Single-Source Truth Maintenance
Abstract
An analytics agent deployed on real business metrics fails in two ways that matter operationally: it returns a confidently wrong number, and it drifts silently out of correctness as the definitions underneath it change. We treat both as reliability problems and address them in the agent's knowledge substrate rather than its prompt. We present a metric-centered knowledge graph in which every question lands on one typed Metric node; the agent reads only that node's one-hop context — guardrails, sliceable dimensions, the matched slice — and delegates all arithmetic to a deterministic engine that walks the derivation cone to full depth. Governance is a first-class node collected at every node the engine grounds through, so scope refusals and mandated formulas fire before an answer is produced rather than being hoped for in top-k retrieval. Treating vocabulary as a node property and computation as graph structure confines every fact to exactly one node. We evaluate on a deployed enterprise cost and operations reporting domain against a prose-skill basstore ablation. Acrossthree language models andons (1,999 runs), the graphanswers 83–89% within a 0% for the baseline, using1–2 data-tool calls per q3.6–7.3× fewer tokens, and4.4–5.5× lower latency; ion pauses and zero withheldanswers, and on the 25 ruthe data it declined andsaid what was missing rate wrong source. A second,pre-registered study measse each fact has one home,a semantic change is a siprose baseline requires amedian of 16 and up to 96nge requests — leaving upto 95 latent contradictioagainst zero. We formalizethe type system and travehere the design does notwin.