OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework
Abstract
Generative Retrieval (GR) has emerged as a promising paradigm for modern search systems, offering end-to-end joint optimization and lower serving overhead compared to multi-stage cascaded architectures. OneSearch, as a representative industrially deployed generative framework, has delivered substantial commercial benefits. However, three key limitations constrain further improvement: shallow understanding of complex queries, insufficient personalized intent reasoning over user context, and a separately trained reward model that adapts slowly to emerging queries and is prone to sampling bias and reward hacking. To address these challenges, we propose OneSearch-V2 with three key innovations: (1) a thought-augmented query understanding module that generates compact keyword-based chains-of-thought (CoTs), overcoming the shallow-matching limitation of single-pass SID generation; (2) a reasoning-internalized self-distillation pipeline that encodes keyword-guided reasoning into the existing model weights via information-asymmetric supervision, without any extra parameters, special tokens, or inference-time CoT generation; (3) a behavior-feedback preference alignment system that replaces the separate reward model with composite rewards built from real user interactions, and introduces a token-position marginal advantage (TPMA) mechanism for position-aware credit assignment over hierarchical SID sequences. Extensive offline evaluations demonstrate OneSearch-V2's strong query understanding and personalized intent modeling capabilities. Online A/B tests further validate its business effectiveness, yielding +3.98\% item CTR, +1.17\% PV CTR, +2.90\% PV CVR, +2.07\% buyer volume, and +2.11\% order volume. Manual evaluation further confirms gains in search experience quality, with +1.37\% in page good rate and +1.65\% in query-item relevance. Importantly, OneSearch-V2 achieves these gains without any additional inference cost or serving latency, while also mitigating information bubbles and long-tail sparsity.