Fast but Forgetful: Cross-Agent KV Reuse Breaks Decision Fidelity in LLM Security Review
Abstract
Multi-agent software engineering increasingly relies on specialized agents to generate, review, and validate code, creating redundancy when the same program is processed across different contexts. Cross-context key-value (KV) cache reuse can reduce this cost, but whether it preserves security-critical decisions remains unclear. We study this as an inference-fidelity problem by holding the model, code, prompt, and decoding configuration fixed and comparing approximate reuse against identical-input exact dense inference. Across 533 paired gate-hit executions spanning three open-weight models, KVCOMM admits 65.8% of validator calls to reuse and achieves a mean 2.52× TTFT speedup. Yet approximate execution changes 40.6–85.7% of security reports and silences 21.7–89.5% of dense vulnerability findings while creating at most 5.3% new findings. Most suppressed findings transition from Fail to Not Applicable, revealing a strongly directional fidelity failure. A structurally distinct KV-reuse method, CacheBlend, reproduces the same pattern of decision deviation. PrimeVul experiments further show that both mechanisms can suppress dense detections corresponding to externally labeled vulnerabilities. Cross-context KV reuse can therefore improve inference efficiency while silently altering downstream security behavior.