PyPilot: Variance-Aware Reward Shaping for Reinforcement Learning in Code Generation
Abstract
Applied to program synthesis, Large Language Models (LLMs) can solve programming tasks but often fail due to syntax errors, runtime errors, or incorrect logic. Finetuning strategies aim to mitigate these shortcomings by specializing models to tasks. PyPilot fine-tunes an LLM (Qwen3.5-9B) using supervised parameter efficient finetuning (LoRA) followed by reinforcement learning (RL) from unit-test feedback via PPO across three reward schemes. RL theory models autoregressive code generation as a finite-horizon Markov Decision Process. We set up a measure theoretic path-space formalism, verify that the objective function is well defined and continuous, prove the policy gradient theorem, and prove a token-level KL decomposition. The main theory contributions are proofs of critic consistency for terminal execution rewards under idealizations, a sufficient condition for variance reduction under reward shaping, and a gradient variance minimizer reward. Experiments evaluate models on a cleaned LeetCode dataset using compile rate, pass rate, pass@k, and paired statistical tests. The results indicate that supervised finetuning eliminates syntax errors, while PPO-based RL improves the pass rate, with reward shaping providing a measurable variance reduction and pass-rate gain.