Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
Shashwat Pandey ⋅ Satwik Pandey ⋅ Suresh Raghu
Abstract
On-device language models now ship to hundreds of millions of devices with no server-side moderation, yet the configuration developers can actually deploy is rarely audited independently. We present a reproducible, black-box reliability audit of a developer-accessible on-device foundation model ($\sim$3B parameters), framed as an oversight question: can a user or a resource-constrained developer tell when the model is wrong? We find a \emph{task-asymmetric miscalibration}, its guardrails fail in opposite directions across tasks (confabulating on 69\% of false-premise questions while refusing 18\% of benign inputs), atop a self-reported confidence that is saturated and non-discriminative (AUROC 0.47; ECE 70, worst among comparable small models). Crucially, confident-correct and confident-wrong outputs are \emph{surface-indistinguishable}: a classifier over 15 user-visible features separates them at AUROC only 0.55 (equivalence-confirmed), so there is no signal for oversight at inference time. No cheap single-generation signal flags these failures ($\le$0.68 AUROC), whereas a black-box consistency wrapper requiring no model access recovers reliability (confident confabulation 75\%$\to$3\%; selective accuracy 43\%$\to$83\%) at a tunable cost. We contribute a model-agnostic audit protocol, a surface-indistinguishability test, and released code and frozen items as reusable infrastructure for auditing deployed on-device models.
Chat is not available.
Successful Page Load