A Lightweight Safety Prefilter for LLM Agents: Mitigating Context-Dependent Shortfalls in SAEs
Abstract
Autonomous LLM agents act through tool-calls, some of which are irreversible, such as transferring money or deleting a file. Such agents must check every irreversible tool-call for safety, but running a heavy guard at every step is too costly. A lightweight prefilter that reads the agent's own hidden state through a sparse autoencoder (SAE) is an attractive candidate, but it remains unclear which kinds of harm such a prefilter can reliably catch. We analyze where the prefilter falls short. The prefilter catches concept-level harm, visible on the surface of the call itself, reasonably well, but misses context-dependent harm, where the same call is harmful only in relation to its surrounding context: the relevant signal is present, but not concentrated into the few features a practical prefilter can afford to read. This shortfall depends on how the feature dictionary was trained; it is not inherent to SAEs. We show that escalating irreversible actions makes the guard deployable, while installing the missing feature directly into the dictionary demonstrates that the underlying limitation is surmountable in principle rather than fundamental.