MemEdge: Dual-Temperature Edge-Cloud Orchestration for On-Device Persistent Memory
Jifu Li ⋅ Xueshu Chen ⋅ Shiming Wang ⋅ Zhen Bi ⋅ Jungang Lou
Abstract
Persistent memory benefits long-horizon agents but strains edge devices, where storage, retrieval, and generation share limited memory and compute. We present MemEdge, a dual-temperature edge--cloud controller that makes the LightMem pipeline deployable under these constraints. MemEdge keeps persistent memory and vector retrieval on the device. Before retrieval, a resource temperature $T_r$ summarizes device pressure to set the active-memory ratio and initial retrieval budget; after retrieval, a query temperature $T_q$ evaluates evidence sufficiency to select local answering, cold-tier rescue, or cloud escalation. On the 1,000-question MobileMem held-out split, MemEdge improves QA accuracy over static length routing from 7.86\% to 16.20\% with DeepSeek-V4 Flash and from 8.02\% to 16.50\% with DeepSeek-V4 Pro, while reducing peak incremental RAM, retrieval latency, and cloud cost. Controlled ablations quantify the effects of the two signals on resource allocation and evidence-aware execution.
Chat is not available.
Successful Page Load