Ontoscope: Ontology-Guided Retrieval Scoping for Efficient Retrieval-Augmented Generation
Dhananjay Khulbe ⋅ Siddarth Tumu ⋅ Siddharth Mahesh ⋅ Jonathan Wang
Abstract
Retrieval-augmented generation uses large document collections as external memory, but searching global indexes is costly when a query concerns a narrow semantic region. We isolate predicate execution from predicate inference: for each benchmark question, we construct an oracle Wikidata predicate from its gold passages and evaluate retrieval after this scope is known. Ontoscope partitions a corpus by an existing Wikidata ontology and reorders storage so each category occupies contiguous memory; under these predicates, the scan returns exactly the predicate extension, so recall equals an exact prefilter by construction. At matched recall, ontology-ordered scanning is 5.8-68x faster on HotpotQA and 1.8-43x on MuSiQue at measured grid points, against post-filtered baselines widened until they reach Ontoscope's recall. Against FAISS $\texttt{IDSelector}$, recall is empirically identical and Ontoscope is faster in 7 of 10 cells. The mechanism is cost per comparison: 10,056 comparisons in 0.583 ms against HNSW ef=256's 5,030 in 0.922 ms, so contiguous memory costs roughly one-third at this operating point.
Chat is not available.
Successful Page Load