Delta Attention Residuals for Cross-Layer Information Flow in Indic Models
Abstract
The streams left by Transformers fetch information over the course of all network layers, creating difficulty in accessing previously found info. Our work includes Delta Attention Residuals– the new type of Attention Residuals that retains the tra ditional structure but provides weighted aggregation of information obtained before in accord with tokens. Moreover, Delta Block is introduced, which accumulates the data in residuals via grouping layers so as to decrease the number of routing and memory costs. The results of our experimentation based on tuning Qwen3-0.6B model on Hindi sub-corpus of FineWeb-2 dataset and benchmarking our method with standard tuning and Attention Residual tuning involving 10K training steps revealed that Delta Block exhibits the least validation loss and perplexity while pro viding advancement from 29.3% to 30.0% in downstream tests as well as certain results on ARC-Challenge, Winogrande and IFEval-Hindi tasks. In addition to this, routing interventions demonstrated that the increased sharpness of routing is not responsible for the enhanced performance since it seems to be based exclusively on the ability to retrieve historical information. Based on the above findings, we can conclude that the preservation of conventional residual carrier in combination with the utilization of tastefully retrieved residual differences could serve as a valuable alternative to replacement method-based deep routing in language-model adaptation.