Better Boundary Prediction Does Not Guarantee Better Structural Representation
Michael Zhou ⋅ Markus Frohmann ⋅ Junghyun Min
Abstract
Punctuation marks often mark phrase boundaries, carrying linguistic information. It has been suggested that as abundant signals of linguistic structure, optimizing on phrasal structure via punctuation marks as its proxy may help language models (LMs) improve their representations of argument, semantic, and syntactic structures. In this paper, we verify this claim in English across three primary LM architectures: decoder-only, encoder-only, and encoder-decoder LMs, by complementing traditional token prediction with punctuation restoration (PR). Our evidence, collected at the $<1$B parameter scale, suggests that optimizing for boundary prediction does not necessarily guarantee higher-quality structural representation in decoder-only, encoder-decoder, and encoder-only LMs, failing to improve performance on downstream tasks that require understanding structure in linguistic input.
Chat is not available.
Successful Page Load