Accommodation or Drift? Nested Controls for Probing Conversational Adaptation in Foundation-Model Embeddings
Abstract
People in conversation are thought to adapt to each other. Probes for this adaptation in speech and text embeddings are usually validated with one control, that the effect disappears when the turns are shuffled. We show that a probe can pass this test for the wrong reason. On the CANDOR corpus (1,656 conversations) we measure how the distance between two speakers' turn embeddings changes from the start of a conversation to its end, for five representations including two foundation models (wav2vec 2.0, all-MiniLM). The change is significant, differs in direction by linguistic level in the way accommodation accounts allow, and disappears under shuffling. We then pair each speaker with a stranger who never heard them. Instead of vanishing, the change persists. Strangers reproduce all of the vocal-state shift and a quarter to six tenths of the others (partner-averaged, anti-conservative tests), and outside vocal state the shift is undetectable once the stranger's turns run backwards in time. Much of it is speakers drifting the same way over a recording, not adaptation. One drift is measurable. Turns grow longer, and regressing turn duration out removes the vocal-state convergence entirely. Real pairs do carry structure that strangers lack. For the adjacent-turn metrics of prior work the strongest audio signal is again turn length, here matched between partners, and the published classifiers remain untested against it. What survives every control is an acoustic and semantic divergence for two of the four dense representations and, for text, local coherence beyond topic. A probe needs a stated direction, a nested sequence of controls that each remove one alternative cause, and an account of what survives.