The Occupancy Curve: Why Note-Placement Accuracy Depends on How Full the Folder Already Is
Abstract
An agent with long-term memory has to put each new note somewhere in the user's own folder tree. That tree is personal and it grows: some folders already hold a dozen related notes, others were made this morning and hold nothing. Recent work scores placement with a single accuracy number that averages over both. We call the number of notes already in the target folder its occupancy, and we report accuracy separately at each occupancy level, including a zero-occupancy split in which whole folders are held out of training so that every evaluation note lands in a folder the system has never seen used. Four methods (prompting with the whole folder list, lexical and embedding nearest-neighbour retrieval, a retrieve-then-pick cascade, and a fine-tuned 8B model) were run on three corpora: one private working directory, 27 public personal note vaults, and 16 public software repositories. No method wins everywhere, and the ordering itself moves. On 27 public vaults the cascade ranks second of five when the gold folder holds one or two notes and fourth when it holds ten or more, while embedding kNN moves in the opposite direction, from fourth to first. The cascade that wins outright on our private tree loses on both public corpora. We tested the obvious explanation for that reversal (tree size) at the correct unit of analysis and it did not hold. For personal memory this matters twice over: the method that looks best on a mature tree is not the method that handles a folder the user made this morning, and the ranking measured on one person's tree did not transfer to other people's. We therefore recommend three reporting changes for this task: break accuracy out by occupancy, always include a retrieval-only baseline, and always include an evaluation on folders that have never been used.