Position: Alignment has a Fantasia Problem
Abstract
In accomplishing complex tasks, human cognition typically progresses from abstract to concrete, for example from brainstorming ideas to producing a final artifact. With the advent of highly capable AI assistants, people now offload various parts of their task to the system. However, instruction-tuned AI systems are optimized to complete tasks as written, often producing final outputs without requiring users to traverse the intermediate decisions that structure the task. When a user approaches AI while their goals and intentions are abstract, AI systems often short circuit their cognitive process through which those goals would be refined by jumping toward a final output (e.g., writing the essay entirely). Doing so takes away the user's agency in achieving the task: they may need to spend more time revising or, worse, settle on a suboptimal outcome. We call these failures Fantasia interactions after the famous Disney scene. We argue that Fantasia interactions demand a rethinking of alignment research, where AI systems optimize how cognitive responsibility is allocated within an interaction. We highlight gaps in state-of-the-art alignment methods, and outline a research agenda for training and evaluating models to achieve this vision.