ViAR: A Visual Autoregressive Router for the Capacitated Vehicle Routing Problem
Abstract
We present ViAR, a neural constructive solver for the Capacitated Vehicle Routing Problem (CVRP) that reformulates instance encoding as a visual understanding problem. Instead of representing an instance only as a set of coordinate-demand tokens, ViAR converts the depot and customers into a six-channel image representation, which is processed by a convolutional neural network to extract spatial context. For each node, the resulting visual features are fused with exact coordinates and demand to form a hybrid visual-coordinate embedding. A lightweight Transformer then propagates global context across nodes, and an autoregressive pointer decoder constructs routes while enforcing capacity feasibility through hard masking at every decoding step. ViAR is trained in two stages: supervised imitation learning from high-quality expert solutions, followed by self-critical sequence training to directly optimise routing quality. Controlled routing-quality ablations show that the visual architecture improves over a coordinate-token Transformer and that the CVRP-specific spatial landscapes further improve the complete policy. ViAR also provides useful warm starts for Guided Local Search under matched end-to-end budgets and shows promising transfer to instances well beyond its training scale. These results position ViAR as a practical learned representation for hybrid routing pipelines rather than as a replacement for specialised routing solvers.