LatentAudit: Resource-Efficient White-Box Auditing for Trustworthy RAG in High-Stakes Domains
Abstract
Retrieval-augmented generation can retrieve the right source and still produce an answer that departs from it. Checking every answer with a second LLM adds latency, cost, and possible data egress. We ask whether the generator's own activations contain a cheaper warning signal. LatentAudit aligns frozen evidence embeddings with pooled answer states on faithful calibration pairs, then scores covariance-normalized alignment residuals. The rule is fit in closed form and adds 0.24–0.77 ms of post-generation computation in our measurements. Across three QA domains and five open-weight model families, this signal approaches GPT-4o judging and transfers to HaluEval. RAGTruth exposes its boundary: zero-shot AUROC is 0.6947 and rises to 0.7816 after 200-example recalibration because whole-answer pooling can dilute localized errors. LatentAudit is therefore a first-pass routing monitor, not a truth oracle or claim localizer. It requires open weights, trustworthy retrieval, and validation after deployment shifts.