Interpretable Semantic Resonance for Steerable Dense Retrieval
Abstract
Dense retrieval models based on Large Language Model (LLM) embeddings deliver strong semantic matching performance but introduce operational challenges in production search systems, including limited diagnostic traceability and absence of determin- istic control. When retrieval errors occur, maintainers often lack mechanisms to isolate and adjust specific semantic drivers without retraining the model. We present Interpretable Semantic Resonance (ISR), a light- weight sparse bottleneck framework that decomposes dense em- beddings into steerable semantic facets while preserving retrieval topology. ISR is trained using a reconstruction, sparsity, and inner product preservation objective and supports query time semantic weighting via a modifier vector. Across five BEIR subsets, ISR maintains statistically compara- ble NDCG@10 to its dense baseline while improving diagnostic traceability by 2–5× and enabling controlled ranking adjustments with bounded latency overhead ( 9ms on NVIDIA T4). We analyze hardware, sparsity, and storage tradeoffs and discuss deployment guardrails. ISR demonstrates that controllable and auditable neural retrieval can be integrated into existing dense pipelines without retraining encoders or sacrificing ANN compatibility.