Neural Artifacts as Multi-View Data: Competence, Dependence, and Stability in Backdoor Evaluation
Abstract
A neural artifact can support many analytical readouts, but heterogeneous readouts need not provide independent evidence or remain stable across distributions. Neural representations therefore form a multi-view data modality in which two properties are central: view competence, the predictive value of each readout, and error dependence, the extent to which readouts fail on the same examples. In a controlled Qwen3-4B backdoor setting, we evaluate nine established activation analyses across five aligned artifact conditions that preserve at least 90% backdoor success while changing the training stage or internal representations. Pairwise joint errors exceed an independence baseline in every condition, by as much as 4.63 times, and the dependence structure itself changes across artifact conditions. Stability is also limited: when the measurement system fitted on the final model is transferred without target-condition recalibration, 30 of 36 view applications to shifted conditions fall to approximately chance accuracy, compared with none of the nine views on the source distribution. A downstream stress test over all 511 non-empty view subsets shows that these properties materially affect behavioral inference. The results motivate neural-artifact evaluation protocols that report per-view competence, inter-view error dependence, and cross-distribution stability together.