History-Augmented Temporal Attention for Waypoint Handoff in NMPC Policy Distillation
Abstract
We study offline distillation of an NMPC controller for waypoint following on a resource-constrained holonomic robot. Temporal Attention combines a current-state query over two waypoints with an eight-step, 64-unit GRU history summary and training-time history dropout. We compare this complete design with Hard Selection, Joint MLP, Scalar Gate, and Lightweight Attention on four physical routes and three independently trained seeds per architecture. Relative to Lightweight Attention, Temporal Attention reduces the observed median 100-ms handoff command difference from 0.339 to 0.281 m/s (17.1\%) and mean cross-track error RMSE from 0.0481 to 0.0468 m (2.7\%), with 12/12 completion. Its RK3328 CPU p95 inference latency is 1.1451 ms, below the declared 5 ms ceiling. We create a smoothness metric to measure deviation from the pre-switch command. These physical comparisons support the feasibility of the complete policy on the deployment CPU. Current component contributions and broader generality remain unresolved.