Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
Abstract
Agentic financial systems increasingly rely on cloud-scale large language models (LLMs), whose latency, cost, and privacy barriers are poorly matched to the repetitive, narrowly scoped sub-tasks that dominate real pipelines. We argue that alpha screening—deciding which quantitative factors are economically relevant under the current market state—is precisely such a sub-task, and present Alpha-R1, a small language model (SLM) with 8B parameters trained via reinforcement learning to perform it. Its core mechanism, semantic gating, evaluates each candidate factor's semantic profile against a dynamically constructed market state description, selecting a sparse subset of factors whose economic rationale aligns with current market conditions. Alpha-R1 is trained via group relative policy optimization (GRPO), using realized portfolio returns as the primary reward signal. Under a 12-month out-of-sample evaluation, Alpha-R1 achieves annualized returns of 47.87% on S&P 500 and 40.57% on CSI 300 with Sharpe ratios of 1.62 and 2.23, outperforming both traditional quantitative strategies and zero-shot prompting of far larger proprietary LLMs. These results, obtained under a bounded candidate-pool evaluation protocol, provide evidence that a small, domain-aligned language model can carry out the context-conditioned reasoning required for second-stage factor reranking in non-stationary markets.