Complementary by Coverage, Correlated by Absence: Output Representation and Complementarity in Multilingual Recognition
Abstract
How a model represents a language can determine whether model diversity pro- duces useful complementary predictions. We study this in multilingual medieval handwriting recognition, a sequence-prediction setting where output inventories and tokenizer coverage can be measured exactly. Three heterogeneous recognizers span three output representations, a general-purpose 50,265-token subword vocab- ulary, a corpus-fitted 1,000-token vocabulary, and direct character emission with no subword vocabulary at all. The corpus-fitted vocabulary provides near-complete atomic character coverage in every evaluation language despite being fifty times smaller, so vocabulary size and vocabulary suitability come apart. Using a per- line selection oracle to measure complementary correctness, we find it substantial within the training family and under zero-shot transfer to an unseen language of the same family, where oracle selection would cut error by 27.9% and 28.8%, then collapsing to 9.2% across language families as the rate of joint failure reaches 99.6%. The direct-character system supplies the best available hypothesis on 16.2% of in-family lines despite being the weakest overall. The cross-family collapse coincides with regions of the training character distribution unsupported across all three systems, with characters absent from training accounting for 0.31% of Czech text against 0.0063% of Occitan. A standard voting combiner recovers only part of the remaining headroom, its gains concentrated where the systems already agree. Output representation and training-distribution support are therefore useful axes for evaluating multilingual sequence models alongside architecture and scale.