How Should a Personal Agent Store Its Facts? Granularity, Compute, and Robustness in Parametric Memory
Mahdi Ghodsi ⋅ Megha Sri Satya Sai Devineni ⋅ Chandra Shekhar Pandey
Abstract
A personal language-model agent must keep a user's facts somewhere: as prompt context at query time or encoded in adapter weights. When memory is partitioned, both approaches require selection: one retrieves text into context and the other activates a memory partition. The open design question is then not whether to select but how finely to partition what is selected. We study this on a 4B instruction-tuned model over a fixed 200-fact user memory, holding per-adapter rank constant at 16 and varying the number of partitions across three adapter-construction regimes. Recall from adapters compiled by a hypernetwork falls steeply as partitions coarsen, from 0.602 at 20 facts per adapter to 0.116 at 200, against a 0.050 base rate. Gradient-trained adapters given exactly 249 optimizer updates each decline similarly (0.746 to 0.354, $p=8.4\times10^{-17}$). When instead every partition receives equal per-fact exposure at 83 epochs, original- question recall remains 0.740–0.773 without a monotonic decline, while recall on paraphrased questions still deteriorates with coarser partitions (0.564 to 0.387, $p=2.1\times10^{-5}$). Granularity therefore strongly affects parametric memory, but the relationship depends on the construction regime and training exposure rather than on fixed adapter capacity alone, and robustness and recall respond differently. We measure two in-context points on the same items: among the five matched operating points, placing all 200 facts in every prompt reaches 0.862 original and 0.834 paraphrased recall at 131.4× the prompt tokens of the parametric arms per query, with top-5 retrieval intermediate on both accuracy (0.729) and cost (4.67×). No mechanism dominates every measured quality and cost axis; we report what each one costs.
Chat is not available.
Successful Page Load