Structure over Pixels: Learning Variable-Length Visual Programs
Piotr Wyrwiński ⋅ Kacper Dobek ⋅ Krzysztof Krawiec
Abstract
Discrete visual tokenizers map images to ordered sequences of codes, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate--distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a lower nominal rate than a fixed $K=32$ baseline, the adaptive model improves alignment with frozen DINOv3 features. At roughly $250$ nominal bits per crop, STROP programs also yield higher segmentation mIoU than FlexTok and One-D-Piece under the same readout architecture and training protocol. STROP therefore learns useful per-image sequence lengths without post-hoc search or a predefined set of compression rates.
Chat is not available.
Successful Page Load