Romanian Small Language Models Trained on Native Child-Targeted Media
Abstract
Developmentally plausible pretraining data remain scarce outside high-resource languages. Romanian occupies BabyBabelLM's 1M-word tier, which relies heavily on fallback text such as movie subtitles. To expand the available Romanian pretraining data, we assemble 10M words of native captions through a documented procedure beginning on YouTube Kids under profiles for ages 0–4 and 5–8, then expanding across eligible channels. We train two model architectures under four protocols and evaluate agreement accuracy on Romanian MultiBLiMP. With GPT-BERT, accuracy rises from 68.2% at 1M words to 74.7% at 10M; a descriptive log-linear fit to the ten budget means explains 97% of their variance. At a matched 1M-word budget, however, models trained on the BabyBabelLM Romanian corpus outperform models trained on separately drawn 1M-word caption subsamples. This difference is consistent with greater lexical overlap between the BabyBabelLM corpus and the evaluation sentences. Models trained on the full 0–4 discovery subset outperform those trained on the full 5–8 subset under all four protocols. These results show that platform-seeded child-targeted media can provide useful Romanian pretraining data, with grammatical performance varying measurably across corpus scale, text source, and discovery profile.