SIREN: Evidence-Coupled Rule Editing for Selective Memory Use
Abstract
A persistent assistant must learn both what behavior a user prefers and when that behavior is appropriate. We propose SIREN, an evidence-coupled rule-memory framework that separates rule installation, applicability refinement, activation eligibility, and execution on a frozen language-model backbone. On a controlled collection of 82 rules, mean condition-recognition AUC improves from 0.743 to 0.922 as evidence accumulates. In a full-library comparison with matched supervision, SIREN achieves a balanced exact-routing score of 63.80%, versus 33.07% for a supervised MiniLM selector, through substantially fewer wrong selections and false activations. Evidence eligibility reduces accumulated wrong-rule decisions by 63.1%. A trainable attention-memory implementation, SIREN-KV, completes 430/464 requested actions when given the correct rule; actual selective execution completes 221/464, versus 136/464 for the base model. Sequential installation preserves each rule's post-installation behavior. These results connect evidence-based applicability to effective persistent execution. The experiments use a known rule collection, and correct-rule recall on new contexts remains 30.03%: broader selective coverage is the main remaining challenge.