Intervention-Screened Visual Perception for Small On-Device Agents
Nakul Vyas
Abstract
An on-device agent may need a small visual decision before it can plan or act, yet replacing its language model with a large vision-language model is costly for a compact local stack. We study a narrower interface: a learned continuous bridge from frozen MobileNetV3-Small features to the input space of a frozen BitNet b1.58 2B-4T decoder. The bridge contributes 34.29M trainable parameters and supplies 1-16 visual embeddings to the decoder's existing causal stream. Candidate labels are scored as text, without a classifier head or language-model fine-tuning. On a locked, balanced Oxford-IIIT Pet10 test set, the four-token system reaches 84.36 $\pm$ 1.55% balanced accuracy over five seeds. Replacing each image with an image from a different class reduces accuracy to 4.52 $\pm$ 1.33%; four further controls remain near chance, and the paired real-minus-swap effect is [77.89, 81.77] pp (95% bootstrap interval). A matched MobileNet specialist reaches 83.33 $\pm$ 0.15%. On four Arm Neoverse-N1 cores, a one-process image-to-answer implementation takes 2.104 $\pm$ 0.031 seconds for a ten-label query. The result is a compact local perception primitive whose textual output can enter a small agent's language stream.
Chat is not available.
Successful Page Load