When Typed Vocabulary Helps and When It Hurts: A Controlled Audit of Memory-Policy Routing
Abstract
Long-term-memory agents increasingly rely on structured intermediate representations before deciding how retrieved user memories should affect a current request. We study this issue in memory-policy classification, where an agent must choose Use, Ignore, Update, or Ask. Exposing the four state definitions significantly improves accuracy on both endpoints (+9.17 pp Llama, Holm p=0.0009; +5.00 pp GPT-OSS, Holm p=0.0135), while requiring the model to emit a state field does not. Because the four states map one-to-one onto the four policies, this may convey label structure rather than useful typed knowledge; in a separate exploratory follow-up on a different controlled set, a non-isomorphic six-field evidence taxonomy hurts Llama (-8.8 pp, Holm p=0.0003), while the GPT-OSS arm had provider failures. On a frozen controlled set with 40 matched four-way families and rule-derived labels, an isolated state-output field changes Llama-3.3-70B by 0.6 pp with 95% CI [-2.9, 4.0] pp and GPT-OSS-120B by 3.3 pp with 95% CI [0.0, 6.7] pp. Per-policy reanalysis shows that the same state-output intervention redistributes policy bias: it shifts Llama toward Use and away from Ask, while shifting GPT-OSS toward Ignore and away from Update. A normalized-input probe replacing natural-language history with structured [subject|attribute|value|tag] tuples hurts accuracy. These results support an audit framing for typed intermediate representations: intermediate fields can expose failure modes, but they do not by themselves solve the decision mapping.