Evaluating Behavioural Decomposition for LLM-Based Motivational Interviewing Fidelity Coding
Hiba Khurshid ⋅ Trang Vu ⋅ Delvin Varghese ⋅ Levin Kuhlmann
Abstract
Motivational Interviewing Treatment Integrity (MITI) coding assigns each therapist utterance a fine-grained behaviour code, and automating it is challenging because distinctions such as simple versus complex reflection depend on how an utterance functions in context rather than on surface linguistic cues. We investigate whether explicitly representing behavioural evidence before label assignment improves classification. We compare four inference-time strategies (direct prompting, chain-of-thought, self-consistency, and self-refinement) with behavioural decomposition, in which a frozen LLM answers 48 code-agnostic yes/no questions derived from MITI 4.2.1 and logistic regression maps these judgements to a code. Six open-weight models were evaluated on 735 therapist utterances from 20 MITI-coded sessions using paired transcript-level bootstrap comparisons. Decomposition did not significantly outperform direct prompting for any model and reduced macro-F1 for five of six models, with some reductions surviving Holm correction under paired bootstrap analyses based on fixed out-of-fold predictions; it was also lower than each model's best observed prompting strategy. None of 18 prompting-strategy comparisons against direct prompting survived Holm correction. Performance varied substantially more across models than across prompting strategies (mean macro-F1 SD .060 vs. .012). Decomposition changed which behaviours were recognised, including improved complex-reflection F1 for models that performed poorly on this code under direct prompting. Behavioural representations also varied across models (mean feature-level Krippendorff's $\alpha = .189$; $.318$ after excluding one outlier). Overall, this decomposition formulation did not provide a consistent aggregate advantage, while revealing substantial model-dependent variation in behavioural evidence extraction.
Chat is not available.
Successful Page Load