DossierBench: Scoring a Recommender’s Model of You on Calibration, Not Recall
Abstract
Recommendation apps increasingly keep a standing preference profile of the user, built from conversation. The memory systems that build such profiles are evaluated on recall, which rewards committing and cannot see the failure that matters once a profile drives recommendations: asserting a preference the user never expressed. We introduce DossierBench, which scores a profile on calibration: does the system commit exactly where the conversation has given it grounds to? Each synthetic subject is a persona whose dining and going-out preferences are fixed on a 24-field card and surface gradually over a scripted conversation of many sessions. An evidence map records, for every field at every checkpoint, what the conversation so far has shown: STRONG, WEAK, or UNEVIDENCED. A system is served the conversation one checkpoint at a time and emits the whole card, abstaining where it will not commit. Committing to an unevidenced field is a fabrication even when the value is right; abstaining there is free. Fidelity over the timeline is normalized to a skill score with always-abstain at 0 and an oracle at 1. On 50 subjects (612 sessions) and four frontier backbones, single-pass extractors that re-read the transcript calibrate well, while on the one backbone where we scored it a stateful memory layer buys nothing at twice the call budget. We release the corpus, scorer, and reference systems; we score the profile, not the recommendations.