DREAMPOLICE: Safer Decision-Making through Language-Constrained Planning
Abstract
Society relies on rules, such as laws, norms, or safety regulations, expressed in natural language, to ensure safe coexistence. Artificial agents deployed among people must follow the same rules. Reactive agents, however, can only learn what a rule means by breaking it. We propose DREAMPOLICE, a framework for constrained model-based reinforcement learning in which language rules are enforced through latent planning. An agent learns a world model from offline data, while a pre-trained vision language model (VLM) labels that data for rule violations, which the agent learns to predict inside its world model. This allows the agent to roll out future states, detect potential rule violations in imagination, and shield a trained policy online from breaking any rules. We evaluate DREAMPOLICE on a synthetic testbed for rule-abiding agents built on the game Pokémon LeafGreen, where simple language rules such as “do not enter private residences” can be verified exactly against the game’s memory. Our experiments measure how reliably an open-weight VLM labels rule violations, how label noise propagates into cost heads, and how effective DREAMPOLICE’s safety mechanisms are: DREAMPOLICE plays the opening section of the game under two such rules without a single violation.