Edge-Aware Dynamic Adaptation for Personalized On-Device Assistants
Ouiame Marnissi
Abstract
Modern on-device AI assistants increasingly rely on multiple personalization mechanisms, including user profiles, long-term memory, retrieval-augmented generation (RAG), and parameter-efficient adaptation (PEFT). Each mechanism is effective in different situations, yet they differ substantially in latency, memory footprint, energy consumption, and persistence behavior. Existing work primarily studies these mechanisms in isolation or combines them using fixed pipelines. For example, RAG systems and PEFT methods such as LoRA are typically triggered by a static schedule rather than a learned decision. Prior work on sequential and bandit-based personalization has explored adaptive policies for individual mechanisms, but has not addressed orchestration across heterogeneous mechanisms with distinct resource profiles. This leaves open a fundamental question: how should an on-device assistant decide, at each interaction, which personalization mechanism to invoke and whether any information should persist, given that user preferences evolve, feedback is sparse and imperfect, and device resources are limited? We argue this decision is inherently sequential: today's choice of mechanism and update shapes the evidence, memory, and adapter state available for every future interaction, so mechanism selection cannot be optimized as a per-turn, single-shot choice. Furthermore, edge devices operate under strict latency, memory, and energy budgets, requiring personalization policies that balance task quality with computational efficiency. We propose EADOA, a hierarchical decision framework that formulates on-device personalization as a constrained partially observable Markov decision process. At each interaction, EADOA jointly reasons over (i) the generation strategy, selecting how the current request should be personalized (e.g., the base model, profile injection, personal RAG, or adapter routing), and (ii) the persistent-update strategy, deciding whether the interaction should trigger a memory revision, a deferred parameter-efficient adaptation, or no change at all. The controller maintains a belief over the user's latent, potentially evolving preferences and optimizes long-term personalization utility under latency and energy constraints. Our target formulation is a unified sequential policy that dynamically orchestrates heterogeneous personalization mechanisms across interactions, rather than optimizing a single mechanism or relying on a fixed pipeline. As an initial proof of concept, we construct a lightweight simulator consisting of twenty representative personalization scenarios covering profile-based preferences, personal memory retrieval, domain adaptation, and general requests. For each interaction, we compare four fixed personalization strategies (Base, Profile, Personal RAG, and Adapter) against an oracle adaptive controller that always selects the most appropriate mechanism. We evaluate a composite utility $U = \alpha\times Success - \beta \times energy_{norm} - latency_{norm}$, where success is a per-scenario correctness score (e.g., exact-match against the scenario's ground-truth personalization target) and $energy_{norm}$, $latency_{norm}$ are min-max normalized per-interaction costs and $\alpha, \beta, \gamma$ are tunable weights balancing task quality against resource costs. The oracle controller in Figure 1 achieves substantially higher average utility than any fixed strategy, indicating meaningful headroom for adaptive, per-interaction mechanism selection. This gap motivates our next step: replacing the oracle with a learned sequential controller and evaluating it under simulated preference drift against myopic and fixed-pipeline baselines. This work reframes on-device personalization not as a choice of mechanism, but as a sequential control problem over personalization mechanisms evolving over time. We believe this perspective will become increasingly important as AI assistants integrate a growing set of heterogeneous, resource-constrained personalization capabilities.
Chat is not available.
Successful Page Load