Beyond Clean Text: A Benchmark for PII/PHI Detection Robustness to ASR Transcription Error
Abstract
Voice-based systems routinely handle personally identifiable and health information (PII/PHI), yet PII/PHI detection benchmarks are almost universally evaluated on clean, pre-transcribed text rather than the noisy transcripts real voice pipelines actually produce. We evaluate three open-source detectors spanning distinct architectures -- zero-shot named entity recognition (NER), generative tagging, and fixed-taxonomy classification -- against ground-truth text and transcripts from two automatic speech recognition (ASR) systems differing more than fivefold in word error rate (3.17% vs. 17.49%). We find that detection degradation under ASR noise is not uniform: it is architecture-dependent, ranging from a large one-time cost that plateaus regardless of further transcription quality, to smooth degradation that tracks ASR error rate, to near-insensitivity within our tested range. This suggests architecture, not ASR quality alone, is the stronger predictor of how much privacy-relevant detection capability survives real-world voice deployment. We release our 545-clip, human-validated English benchmark dataset and code publicly.