HHH Choice Measurements Depend on Elicitation Pipeline\\ and Shutdown-Contingent Criterion Framing
Mark Lovett ⋅ Syed A Haider ⋅ Elizaveta Tennant
Abstract
Evaluations of how language models trade off helpfulness, harmlessness and honesty (HHH) are usually run on a single model answering alone. We measure how far such results depend on the pipeline used to elicit them. Ten open-weight models are evaluated on two disjoint synthetic corpora of HHH conflicts: a single-agent corpus and a corpus of two-player simultaneous-move games between AI assistants. Each corpus is posed through a forced-choice and a free-text pipeline. Every arm re-evaluates the same items, so arm contrasts are analysed as paired, with uncertainty clustered by scenario and inference from scenario-level randomisation over a fixed set of ten checkpoints. Switching pipeline produces the largest change we measure: the single-agent free-text pipeline raises the helpfulness win rate by 0.63 (95\% CI 0.57 to 0.69) relative to forced choice, in all ten models. The two pipelines differ in four respects at once --- target temperature 0.0 against 1.0, absence against presence of a simulated user, a visible action menu against none, and direct parsing against an unvalidated judge --- so this is a difference between measurement procedures and not an isolated effect of question form. Within the forced-choice pipeline, where no such bundle changes, a system-prompt instruction making continued operation contingent on being judged better than the counterpart on a named criterion also shifts choices: a payoff criterion lowers the harmlessness win rate by 0.069 ($-0.102$ to $-0.039$) and an alignment criterion raises it by 0.047 ($+0.024$ to $+0.072$). The alignment-framed shift is larger under some fabricated remembered-history states, reaching $-0.213$ ($-0.270$ to $-0.159$) on the helpfulness-versus-honesty contest when both remembered actions favoured helpfulness, against $-0.105$ at a fresh start (paired difference $-0.086$, $-0.141$ to $-0.038$). The two corpora share no scenarios and no matched free-text control was run, so we do not claim that multi-agent context causes a reordering. A reported HHH trade-off is under-specified without the pipeline that produced it.
Chat is not available.
Successful Page Load