Self-Other Simulated Introspection
Abstract
Self-adapting language models can improve on new tasks through model-generated training signals and parameter updates, but these adaptations may also degrade safety behaviors that are not represented in the adaptation objective. We propose Self-Other Simulated Introspection (SOSI), a safety-repair framework that uses mechanistic evidence from a model's own safe and unsafe trajectories to identify candidate intervention sites and test bounded repairs. SOSI ranks internal components using activation contrasts and gradient-based evidence, presents the resulting evidence to the same checkpoint under an Agent-B framing, and tests bounded interventions before distilling accepted repairs into targeted parameter updates. We evaluate SOSI on six Qwen2.5-7B and Mistral-7B-Instruct-v0.3 SEAL checkpoint arms using generated-response safety, hazardous-knowledge, and benign-preservation evaluations. In the paired temporary comparison available for Qwen Iter2, a single behavioral judge finds no reliable generated-response improvement; across six arms, SOSI's aggregate temporary outcomes vary, while persistent gradient-distilled updates reduce held-out harmful generations on some checkpoints. Our results support a verification architecture in which an LLM may participate in proposing its own repair, while trusted host-side mechanisms retain localization, causal testing, and authorization authority before any persistent model change is accepted.