Escaping the Obvious: Reinforcement Learning for Conceptually Diverse Generation
Abstract
Large language models often converge on a narrow set of high-probability ideas during open-ended generation, repeatedly following familiar conceptual associations rather than exploring the broader space of concepts available. We investigate whether a simple reinforcement-learning objective can counter this tendency by directly rewarding further conceptual exploration. We post-train Qwen3-14B using Group Relative Policy Optimization (GRPO), rewarding outputs that connect the prompted topic to a semantically distant secondary topic. We evaluate the method on three creative writing tasks: analogy, poetry, and flash fiction. These tasks were chosen because they impose different structural constraints while making conceptual movement observable. We evaluate within-prompt and global diversity, semantic geometry, baseline-calibrated occupancy, and coherence. Results show that GRPO training increases the semantic distance between concepts across all three tasks. This semantic reward also shifts how the model navigates conceptual space. Specifically, the average position of generated outputs is in the least occupied regions of the embedding space. The secondary topics in generated outputs move from literal associations or near-neighbor clusters, such as volcano–eruption, toward more distant but coherent connections, such as volcano–argument. These results show that GRPO can reshape and expand the conceptual regions a model is likely to explore. Our method enables language models to explore more diverse and semantically distant topics, influencing the range of possibilities available to users. This is useful for tasks that benefit from novelty and conceptual breadth, such as creative writing, ideation, design, brainstorming, and hypothesis generation.