OmniSimulator: Aligning Small Language Models for Authentic Heterogeneous Behavior Modeling
Abstract
Large language models (LLMs) have shown great potential as general-purpose user simulators for interactive systems. Prior studies have demonstrated that training on high-quality data can substantially improve model reasoning and downstream task performance. However, the upper bound of user simulation is still constrained by the model's reasoning paradigm: standard chain-of-thought often degenerates into shallow pattern matching and fails to capture the latent mental process underlying real user behavior. In this paper, we propose OmniSimulator, a novel training framework designed to improve the predictive ability of small LLMs for realistic user simulation on short-video platform environments. Specifically, OmniSimulator models human internal deliberation through a structured reasoning schema with five dimensions, carefully selects data from real user-behavior scenarios, and introduces a teacher model to generate high-quality reasoning traces under the proposed schema, boosting small models to learn the latent logic of human decision-making during the CoT process, rather than relying on superficial behavioral correlations. Based on this process, we construct 9,881 high-quality training instances from three months of historical interaction trajectories of 200 real users, averaging about 50 distilled examples per user. Experimental results show that OmniSimulator improves performance over the baseline by 38.93\%, surpasses strong frontier models such as Claude-Sonnet-4.5, and significantly reduces inference cost. These results suggest that learning structured human-like internal reasoning is a key step toward scalable and high-fidelity LLM-based user simulation.