Self-Other Simulated Introspection
Abstract
Self-adapting language models can improve on new tasks through model-generated training signals and parameter updates, but these adaptations may also degrade safety behaviors that are not represented in the adaptation objective. We propose Self-Other Simulated Introspection (SOSI), a safety-repair framework that uses mechanistic evidence from a model's own safe and unsafe trajectories to localize and repair adaptation-induced failures. SOSI ranks internal components using activation contrasts and gradient-based evidence, presents the resulting evidence to the same checkpoint under an Agent-B framing, and tests bounded interventions before distilling accepted repairs into targeted parameter updates. We evaluate SOSI on six Qwen2.5-7B and Mistral-7B-Instruct-v0.3 SEAL checkpoint arms using generated-response safety, hazardous-knowledge, and benign-preservation evaluations. Across the reported evaluations, SOSI improves generated-response safety and shows a favorable observed safety--preservation tradeoff relative to adaptive activation steering, with performance in a similar range to an external automated interpretability baseline. Our results support a verification architecture in which an LLM may participate in proposing its own repair, while trusted host-side mechanisms retain localization, causal testing, and authorization authority before any persistent model change is accepted.