RecMem: Recurrent Memory Compression for Long-Sequence Recommendation
Abstract
Sequential recommenders need long histories, but full-prefix attention scales with user lifetime. We argue that an effective memory mechanism should provide space invariance, semantic fidelity, and lifelong evolvability, which existing fixed-window or capped-memory methods do not jointly satisfy. We propose RecMem, a recurrent compression framework that folds history segments into M forget-gated memory slots. RecMem has a history-length-independent rollout error bound, and its fixed-capacity memory is trained by next-item prediction to preserve useful signals. On MerRec, RecMem uses 5.6x fewer decoder-side tokens than full attention while retaining 88–97% of Recall@10–200. Compared with a compressed truncation baseline at similar decoder cost, RecMem improves Recall@50+ by 4–6 percentage points by compressing the full history rather than discarding old interactions. It also remains stable over 1,800 test steps.