Quantifying the Latency-Accuracy Trade-off in Local Agentic RAG Workflows
Abstract
Agentic workflows improve cloud-based RAG, but their practicality on consumer hardware is still unclear. We evaluate the latency-accuracy trade-off of five local RAG pipelines, from a single-pass baseline to stateful self-correction, across three Small Language Models (Llama 3.1 8B, Qwen 3.5 9B, and Qwen 3.5 4B) on a single consumer GPU. On 500 HotpotQA distractor questions with retrieval restricted to ten paragraphs per question and a two-document generation context, agentic pipelines achieve significant accuracy gains over their baselines in 10 of 12 comparisons. Surprisingly, single-pass pre-retrieval decomposition is consistently the most efficient approach, while its corrective variant achieves the highest accuracy. Qwen 3.5 4B paired with single-pass pre-retrieval decomposition also achieves the highest marginal accuracy gain per added second. The same rankings are observed across all three tested models, suggesting a "start simple, scale up" heuristic for similar local deployments.