A Coherence-Gated SLM for Verifiable Security Triage
Abstract
On-device agents often rely on small language models to perform repetitive, high-volume tasks under tight compute and memory budgets. In security triage, however, generating an explanation for every prediction creates a distinct safety problem: a model can cite evidence that is not present in the flow or produce a plausible-sounding rationale for an incorrect decision. Rather than improving explanation generation itself, we ask a prior question: when should an on-device model be allowed to explain, and how can the explanation be restricted to claims grounded in observed evidence? We introduce a safety-first selective-generation protocol that separates when an explanation is generated from what claims are allowed to reach the analyst. A coherence gate invokes the rationale generator only when its prediction agrees with the upstream ensemble decision. The resulting rationale is then passed to a deterministic structured verifier, which checks each feature/value claim against the observed flow and blocks unsupported claims, replacing rejected rationales with a grounded summary. The full pipeline runs entirely on a single consumer GPU. We evaluate the proposed protocol on two network-flow security triage benchmarks. The coherence gate selects the ensemble's more trustworthy decisions to explain, in contrast to disagreement-based routers, while the verifier, validated through controlled fault injection, catches every injected violation with a 0\% false-positive rate. End-to-end, the protocol reduces unsupported claims in delivered rationales by 55.6\%, ensuring that every flow receives either a verified rationale or a grounded summary.