Don’t Repeat Yourself: Reward-Free Diversity Supervised Fine-Tuning
Abstract
In verifiable domains such as math and coding where multiple solutions can be generated in parallel and scored by a verifier, coverage at high pass@k can be more important than pass@1, especially for harder tasks where one might be looking for a needle-in-the-haystack solution. Post-training in LLMs has been shown to concentrate their outputs to a few modes. Increasing temperature is a commonly used method to improve LLM output diversity, but it hasn’t been effective. We introduce Diversity Supervised Fine-Tuning (DivSFT): a post-training method for increasing coverage and output diversity. For each problem, sequentially generate K solutions, where, for each generation, the LLM is conditioned on the prior solutions and is asked to produce a different solution. Then fine-tune the LLM as if the samples were generated independently. This process involves no reward, verifier, or correctness filter. The result is a model that can generate many diverse solutions in parallel. DivSFT raises pass@100+ by 10.8, 12.5 and 12.4 points across HumanEval+, MBPP+, and DS-1000 at a small cost to pass@1+. The structural sample diversity (abstract syntax tree (AST) edit distance) of passing solutions rises from 0.072 to 0.264 on HumanEval+. DivSFT solves 244 of 600 problems the base model was not capable of solving in 200 attempts. Across nine open-weight models, lower baseline structural diversity significantly predicts higher DivSFT performance gains, indicating that our method is especially effective on more mode-collapsed LLMs.