Neuro-Inspired Inverse Learning for Planning and Control
Maryna Kapitonova ⋅ Tonio Ball
Abstract
We present a neuro-inspired framework for embodied planning and control. Building on three principles that enable fast and highly effective goal-directed behavior in the mammalian brain — paired forward/inverse internal models, open-loop multi-step motor commands, and sequential organization of action — our *Inverter* framework combines analytic and trained components, the latter trained end-to-end through *Inverse Learning* (IL), a paradigm we formalize and delineate from supervised, reinforcement, and imitation learning. IL bridges RL-style amortization, which runs in a single forward pass but emits one action at a time, and optimal-control-style sequence reasoning, which plans whole trajectories but with iterative test-time computation. Single Inverters or hierarchical n=2 Inverter stacks match or improve over comparable offline-RL and diffusion-planner baselines on all 3 `maze2d-v1` and 6 `antmaze-v2` D4RL variants by an average of $+24.2\%$ (range $-1.9\%$ to $+78.2\%$), at one-to-two orders of magnitude less inference compute. As one mechanism behind this effectiveness, we show that Inverters can learn control policies closer to the analytic optimum than the data-generating policy itself. We also identify a failure mode of IL: forward-model hacking under narrow training-data coverage, which we mitigate by using *random* training data with broader coverage. As an application example, a Pulse Inverter synthesizes arbitrary single-qubit quantum gates with fidelity matching the standard iterative numerical baseline (GRAPE), at more than $1000\times$ lower per-gate compute. We propose deeper hierarchical, probabilistic, and latent extensions of the Inverter framework as a differentiable world-interface, especially for latency- and compute-critical embodied AI.
Chat is not available.
Successful Page Load