Monoplex: Parallel Speech and Text from a Single Latent
Abstract
A person walking down a street can point, describe the scene aloud and take a note, all at once and from one thought. Machines built on language models do these tasks one at a time, so their modalities render one content stream rather than working together. We build the parallel alternative, monoplex. In monoplex a frozen vision-language model emits one latent, read at once by its own text head and by a speech head we train. No text passes between them; speech carries a summary, text the detail. With a 41M completion stage monoplex's speech head reaches 86% of a captioning cascade's content in one pass, first audio at 0.12 s against a streaming cascade's 0.32 s. The speech arrives on topic and not yet fluent: a judge reads a third of its transcripts as coherent English against the cascade's every one. The target makes the speech stream trainable, not our objective: a codec whose first codebook is distilled from a speech model learns the alignment, an acoustic codec does not. What remains is coordination between heads that never exchange a token, and it needs no training. Ranking eight draws by the sibling text stream, with their own agreement, reaches 98% of the cascade. Streams read from one source rather than chained through one another can be built today, and this paper prices what they cost.