Slim Foundation Models for Latency-Critical Recommendation Serving
DANYAL AFTAB ⋅ Anas Zafar ⋅ Assadig Abakr ⋅ Steven Davy
Abstract
Click-through rate (CTR) prediction is a core component of recommendation and advertising systems. Large language models (LLMs) offer rich semantic priors that improve prediction quality, particularly in cold-start and long-tail settings. Injecting a full-scale LLM into a CTR pipeline is, however, prohibitively expensive at inference time: existing LLM-enhanced CTR methods avoid this cost by caching LLM outputs offline rather than serving the LLM online. We propose a resource-efficient pipeline, **SlimFM**, that makes online LLM-based semantic modeling practical: (1) dependency-aware structured pruning compresses LLaMA3-8B by 44\% to a 4.5B-parameter backbone; (2) LoRA-based field-aware fine-tuning trains the pruned model to produce discriminative field-level semantic embeddings from lightweight textual prompts; (3) a semantic-guided fusion module injects these embeddings into the feature interactions of a conventional CTR model. Across four datasets and six CTR backbones, SlimFM retains over 99\% of uncompressed-model accuracy while reducing inference latency by up to $3.4\times$ relative to existing LLM-enhanced baselines.
Chat is not available.
Successful Page Load