LFM2.5 Encoders: On-Device Retrieval, Reranking, and Routing for Agentic Workflows
Abstract
Specialized local \emph{Small Language Models} (SLMs) are becoming a viable alternative to large cloud-based models, driven by advances in hybrid architectures, training recipes, weight quantization, and speculative decoding. However, efficient on-device agents require more than generation alone, relying on capabilities such as retrieval, reranking, routing, and classification. We extend LFM2 with three compact bidirectional models for dense retrieval, late-interaction reranking, and general-purpose encoding, obtained by adapting a causally pretrained backbone rather than training encoders from scratch. Together, these models expand the range of agentic workloads that can be handled efficiently and fully on device, while matching or outperforming larger sub-1B baselines across retrieval, reranking, and general-purpose encoding, with sub-10\,ms p50 retrieval and reranking latency on Metal.