Deduplication Deletes the Normal Findings: What Radiology Models Lose to Embedding Curation
Abstract
Near-duplicate removal is standard curation for clinical foundation models, and radiology is where it is most likely to go wrong. Its sentences are templated, and the closest pairs differ only by a negation, a side or a number: what an embedding cannot see, and what carries the finding. We ask what the filter deletes, and whether a model trained on what remains notices. We audit sentence-level cosine deduplication on radiology and case reports, judging every deletion with an entailment cross-encoder, and train small language models from scratch on each filter's output against random deletion of the same number of sentences. The filter deletes 52.4% of radiology sentences where only 16.4% are exact duplicates, and half of what it deletes is clinically distinct from what it keeps. Its deletions fall disproportionately on negated normal findings, and the model learns what the corpus lost. Read by sentence type, the filter-trained model falls +0.046 nats further behind the control on new negated sentences than on other new ones. The gap holds in every seed, appears for no other sentence type, and all but disappears on case reports, which lack negated templates. Clinical-text curation should measure what leaves, and read held-out loss by sentence type.