When Are Semivalue-Based Decisions Identifiable? Robust Data Selection under Utility Ambiguity
Hannah Diehl ⋅ Justin Steil ⋅ Ashia Wilson
Abstract
Semivalues are widely used to assign credit and guide data curation decisions, yet their outputs depend on a utility function that is not uniquely determined by the task. We study partition-based decisions, which directly model data curation tasks such as selecting high-quality subsets or flagging noisy examples, under two structurally unavoidable sources of utility underspecification: monotone transformations of performance scores and unconstrained small-sample behavior. We formalize outcome robustness in this setting as invariance of the induced partition across admissible utility respecifications. We establish that partition robustness is necessary for guaranteed success of semivalue-based data selection. To assess this condition in practice, we provide algorithms that certify partition robustness or produce a concrete witness utility demonstrating instability, requiring no utility evaluations beyond those already computed for the semivalues themselves, with formal correctness guarantees under both exact and approximate computation. We show that utilities admitting a coalitionally dominant top-$k$ subset, where high-quality points contribute more than low-quality points across every possible coalition, guarantee robust partition decisions, subsuming the modular utility case previously identified as sufficient. Together, these results provide a practical framework for determining when semivalue-based selection is well-defined under utility ambiguity.
Chat is not available.
Successful Page Load