Beyond Independent Generation: Set-Level Optimization for Generative Models
Abstract
Many objectives for LLM generation are defined over sets of outputs rather than individual responses, yet existing reinforcement learning methods largely optimize additive per-response rewards. We propose a set-level reinforcement learning framework that jointly optimizes the collective behavior of multiple generations through a lightweight latent policy while keeping the language model frozen. We instantiate the framework on fiction generation, where the objective is to match the aggregate genre distribution of a generated set to a desired target. Experiments show that the learned policy closely matches the target distribution on held-out prompts while inducing distinct latent generation modes, despite receiving only shared set-level feedback. These results demonstrate that aggregate objectives can directly shape the collective behavior of LLM generations without full-model fine-tuning.