Alignability: Steering Components Without Aligning Their Objectives
Abstract
AI alignment is framed as a problem of specifying or stabilizing objectives in individual agents. This paper develops an alternative alignment architecture: alignability, the extent to which a component can be predictably redirected through externally controlled changes to its environment without requiring its objectives to match system-level goals. Profit-seeking firms provide the motivating case: prices, taxes, and changes to the payoff landscape redirect competent behavior while firm objectives remain fixed, and the control signals can themselves be shaped by the participants they steer. Externalities show how system-level alignment can improve or deteriorate without corresponding changes in component objectives. A parallel architecture in multicellular biology suggests that alignability is a recurring multiscale organization principle rather than a peculiarity of economic incentives. For AI, this reframes part of alignment as a design problem: constructing and preserving live control handles through which capable systems remain dynamically steerable as human goals and circumstances change.