What Classical MIAs Actually Consider a Member? Definition–Attack Inconsistency in Membership Inference for LLMs
Abstract
Membership inference attacks (MIAs) are evaluated against a binary ground truth: whether a sequence occurred in training. Yet their scores measure model behavior rather than training inclusion itself. We ask whether classical score-based MIAs operationalize the same notion of membership against which they are evaluated. We reverse-engineer four widely used MIAs: LOSS, Min-K\%, reference-based, and Zlib through controlled interventions that independently vary exact inclusion, target-specific familiarity, exposure frequency, and semantic consistency. Across three random seeds and all four attacks, repeated non-exact exposure to target-specific identifiers produces stronger member-like scores than a single exact inclusion, revealing a systematic mismatch between exact provenance and attack evidence. Further interventions show that this signal arises even from individual identifiers, accumulates without their co-occurrence, and is weakened but not eliminated by semantic contradiction. We call this mismatch definition-attack inconsistency: classical MIA scores behave as graded measures of exposure-sensitive familiarity rather than direct evidence of exact sequence inclusion. Our results motivate evaluating not only whether an MIA predicts membership, but what training evidence actually drives its prediction.