SAVE: Sparsity-Aware Influence Estimation for Vocabulary-Expanded LLMs
Abstract
Large language models adapted to specialized targets such as Lean theorem proving or tool calling often benefit from extending the tokenizer with target-specific tokens. However, target-only vocabulary-expanded fine-tuning can leave the newly added embeddings transfer-incomplete, since the new rows receive gradients only from examples containing the corresponding tokens, so target-local gains do not necessarily imply that the embeddings are integrated with related capabilities. We propose SAVE, a sparsity-aware influence estimator for selecting auxiliary healing data from generic corpora. SAVE computes token-conditional curvature over the effective update set of each new vocabulary row and combines this vocabulary-aware signal with dense LoRA-level influence. Across Lean autoformalization and tool calling, SAVE improves target performance and related transfer while remaining competitive on broad retention benchmarks.