On-Device FoRA: Native Device Fine-Tuning with Selective Riemannian Adaptation
Abstract
On-device adaptation can personalize a foundation model without transmitting local data, but it must fit within a mobile training budget and produce the same learning behavior as a reference implementation. We extend Fisher-orthogonal Rank Adaptation (FoRA), which places full-rank low-rank adapters in a selected subset of layers and optimizes their down-projections on the Stiefel manifold, to native Swift/MLX. We train a 4-bit Qwen2.5-0.5B model directly on a physical iPhone 17 Pro. Over 200 updates, FoRA reduces held-out WikiText-2 perplexity from 35.75 to 33.95 (5.04\%). A matched Apple-silicon GPU run reaches the same final perplexity; the nine-point evaluation trajectories have mean absolute error 0.000705 and Pearson correlation 0.99942. At a 50\% layer budget, FoRA uses 156.9 MB peak process memory, 9.4\% below all-layer LoRA and 16.5\% below all-layer Stiefel-LoRA. A four-batch on-device projected-Fisher calibration recovers the same selected layers as a 16-batch reference in 6.6 seconds. Together with the server-scale cross-architecture evaluation, these results show that selective Riemannian adaptation is not only parameter-efficient in the data center but also executes, learns, and numerically tracks a GPU reference on a commodity phone.