Which Transformer Blocks Should You Keep? An Exhaustive Map of Layer-Subset Selection in a Pretrained ViT
Steffen Jung ⋅ Margret Keuper
Abstract
Removing whole blocks reduces the parameter count and FLOPs of a pretrained transformer without requiring sparse kernels. We exhaustively evaluate all $2^{12}-1=4095$ non-empty block subsets of a 12-block vision foundation model (DINOv3 ViT-S/16) after three-epoch end-to-end finetuning on ImageNet-1k. Under this protocol, top-1 accuracy varies by as much as $34.8$ percentage points among subsets of the same block budget (parameters). A common baseline, prefix truncation, simply retains the first $k$ blocks. However, truncation accuracy falls below the mean of all subsets at every pruned depth $k \in \\{1,\ldots,11\\}$. A uniformly selected subset therefore has higher expected accuracy than truncation at every pruned depth. The value of a block is also budget-dependent: early and late blocks exchange importance as the budget increases, while the final block remains beneficial throughout. Modest local search on the complete landscape approaches the best observed subsets and outperforms random search. We provide the complete measured landscape as a benchmark for budget-conditional structured on-device compression.
Chat is not available.
Successful Page Load