LDSS: Efficient Long-Context Training Data Synthesis via Self-Study
Abstract
Training long-context language models, especially with recurrent memory blocks for efficiency, remains challenging when natural training corpora and signal offer few targets requiring long-horizon dependency. We propose Long-context Data Synthesis via Self-Study (LDSS), which interleaves teacher-generated question--answer (QA) pairs from growing document prefixes with the source text to encourage use of earlier content. In correspondence to the data, LDSS is further equipped with separately normalized document and answer training losses to enhance the supervision from the augmented long-horizon dependency. We evaluate LDSS in two hybrid-model training settings: adapting full attention to sliding-window attention and context length extension. Under matched student training token budgets, LDSS improves aggregate performance over training on original documents across model sizes and teacher sources. Ablations demonstrate the interleaving and separate answer supervision designs respectively. Finally, we show evidence that LDSS gains benefits in a purely self-improvement loop: a model trained on original data generates supervision for a new student initialized from the same base checkpoint.