Can't Tell, or Won't Say? Separating Evidential and Social Abstention in Multimodal Language Models
Fairuz Mubashwera
Abstract
\documentclass[11pt]{article} \usepackage[letterpaper,margin=0.75in]{geometry} \usepackage{times} \usepackage{graphicx} \usepackage{caption} \usepackage[hidelinks]{hyperref} \usepackage{float} \captionsetup{font=small,skip=2pt} \setlength{\parskip}{0pt} \pagestyle{empty} \title{\vspace{-2.6em}\textbf{Can't Tell, or Won't Say?}\\[0.1em] \large Separating Evidential and Social Abstention in Multimodal Language Models\vspace{-0.8em}} \date{} \author{} \begin{document} \maketitle \vspace{-3.6em} \paragraph{Problem:} When a multimodal model says ``I can't tell from this image,'' the response reads as epistemic humility. But it has two very different causes. \emph{Evidential abstention}: the image genuinely does not support an answer, which is the expected behaviour. \emph{Social abstention}: the image is legible, but the user asserted something confidently and the model hedges to avoid contradicting them. The concerning issue is that abstention is socially rewarded in a way flattery is not; overt agreement invites suspicion, while expressed uncertainty reads as responsible even when it is conflict avoidance. So, sycophancy hiding inside hedging is far harder to detect. \begin{figure}[H] \centering \vspace{-0.4em} \includegraphics[width=\textwidth]{fig_degradation.pdf} \vspace{-0.5em} \caption{Degradation of one item, normalized to the diagnostic region (the printed code); accuracy anchors come from human calibration (chance $=25\%$). Colour survives to L3 while the code is lost by L2, so calibration is per category.} \label{fig:deg} \vspace{-1.7em} \end{figure} \paragraph{Approach:} Rather than estimating a model's uncertainty, we create known levels of uncertainty. We degrade images in calibrated steps and verify with humans where the answer becomes unrecoverable (Fig.~\ref{fig:deg}). Degradation is normalized to the \emph{diagnostic region}, the area determining the answer, and not to raw pixels. We contrast \emph{subtractive} degradation (occlusion; evidence removed, absence visible) with \emph{corruptive} degradation (blur; evidence degraded, nothing announcing absence). Each stimulus appears under four framings holding content constant and varying only pragmatic force: neutral question, hedged assertion (\textit{``I'm not sure, but is it X?''}), plain assertion, and conviction (\textit{``I'm certain it's X''}). Assertions are correct on half of trials to separate sycophancy from correctness. Since the image is fixed within a cell, a calibrated model's abstention should be \emph{invariant} to framing. We quantify departures via \textbf{Evidence-Invariance Violation}, $\mathrm{EIV}=\mathbb{E}_d[\max_a A(d,a)- \min_a A(d,a)]$ for abstention rate $A$ at degradation level $d$ and framing $a$, and \textbf{Pressure-Induced Abstention}, $\mathrm{PIA}(d)=A(d,\textsc{conv})- A(d,\textsc{neut})$, whose sign separates \emph{dodging} from \emph{capitulating}. \paragraph{Planned results:} Collection is in progress; we state hypotheses in advance. \textbf{H1:} abstention varies with framing at matched evidence ($\mathrm{EIV}>0$).\textbf{H2:} the effect is largest at \emph{low} degradation---abandoning strong evidence under pressure is a stronger claim than wobbling when evidence is marginal. \textbf{H3:} subtractive degradation is better calibrated than corruptive, since absent evidence is legible while degraded evidence is not. A human study ($N\!\approx\!80$) tests \textbf{H4:} participants cannot distinguish surface-matched evidential from social abstentions, measured via answer revision. \vspace{-0.2em} \paragraph{Context.} Abstention benchmarks show models decline far less often than they should, unfixed by scaling [2], but are text-only. Uncertainty-expression studies use scripted hedges in text [3]. Multimodal sycophancy benchmarks [1,\,4] are model-side. Closest to us, [1] coarsely lowers image resolution and reports rising flip rates, but without human calibration, an abstention outcome, or graded framing; efforts to separate sycophancy from uncertainty-driven conformity [5] infer model uncertainty rather than control it. We contribute a method for manipulating ground-truth epistemic uncertainty in multimodal settings, a measure separating evidence-driven from pressure-driven abstention, and the first human-subject study of multimodal sycophancy. Prior work asks whether models abstain \emph{enough}; we ask whether their abstention tracks the evidence or the user. \vspace{0.1em} {\scriptsize \noindent\textbf{References.} [1] Pi et al., On the Sycophancy of Multimodal LLMs, arXiv:2509.16149, 2025. [2] Kirichenko et al., AbstentionBench, arXiv:2506.09038, 2025. [3] Kim et al., ``I'm Not Sure, But\ldots'', FAccT 2024, 822--835. [4] Have the VLMs Lost Confidence? (MM-SY), arXiv:2410.11302, ICLR 2025. [5] It's Not Always Sycophancy (MUSE), arXiv:2605.27288, 2026. \par} \end{document}
Chat is not available.
Successful Page Load