Navigate, Don’t Generate: Making AI Mediated Access to Public Agricultural Data Citizen-Auditable
Abstract
Agricultural decisions such as variety selection, fertilizer dosing, and scheme eligibility directly affect farmers' yields and livelihoods. Farmers increasingly turn to LLM-based assistants for these \textbf{high-stakes decisions}. However, these systems are stochastic, prone to \textbf{hallucinations}, and provide no \textbf{claim-level provenance} for their outputs. Retrieval augmentation grounds answers in documents but does not enforce structured constraints on eligibility, location, or source validity. We present \textbf{India-Agri-KG}, a nationwide knowledge-graph construction system based on the \textbf{Navigate, Don't Generate} paradigm. Ontology-guided LLM agents navigate heterogeneous government portals, APIs, and PDFs and propose candidate triples. A \textbf{deterministic verification layer} then checks each candidate for source registration, schema conformance, referential integrity, location anchoring, and provenance before permitting graph mutation. Applying the system across \textbf{784 Indian districts}, the resulting graph contains \textbf{14,478 entities and 38,590 source-traceable relations} spanning \textbf{26 agricultural subdomains}, with provenance enforced at write time. The resulting knowledge substrate provides a reusable, auditable foundation for downstream AI systems, while maintaining an explicit authority boundary in which LLMs propose knowledge but deterministic verification controls what becomes persistent graph state.