A Schema-Grounded Benchmark for Long-Term Memory in Workplace Assistant
Abstract
Long-term memory benchmarks for conversational agents now span personal as- sistants, television narratives, multi-party dialogue, and in-vehicle settings, but workplace conversation remains uncovered. Enterprise assistants must retain a different class of fact – incidents, change requests, approvals, asset ownership, and reporting relationships – for which correctness depends on who owns a record, who may act on it, and which state it currently occupies. We build WFMemBench, a schema-grounded enterprise memory benchmark whose ground truth derives from an explicit entity–relation schema, by reconfiguring the BEAM long-conversation pipeline [48] at 100K and 500K tokens. It exposed eleven reproducible defects – six in generation, five in the verification tooling we built to catch them – which we report alongside the per-stage validation layer that detects them, its judges calibrated against human annotation at both scales. Two findings generalise be- yond our domain: the unmodified pipeline scores identically on conversation-flow realism in our domain and in its original ones, locating that limit in the pipeline rather than the domain; and generation quality and verification reliability move in opposite directions as scale increases. We then evaluate five deployed memory architectures under matched conditions – the same conversations replayed through every framework. Published rankings do not transfer – two frameworks lose 10–14 points, reordering the ranking. Architecture dominates retrieval depth: quadrupling retrieval breadth changes accuracy by 0.01, while restructuring memory into ab- straction tiers gains 19–23 points. Temporal reasoning is weak in every system we test. A reimplementation of the strongest design attains 94.6% on LongMemEval and 73.6% on the final, validated WFMemBench.