One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
Abstract
Ambient AI scribes draft clinical notes at scale, and the reassurance offered is that a clinician signs every note. We audited three commercial scribe products on the same 142 consultations - 565 notes across recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two AI models from different families, each told to refute where defensible, and 618 survived. One note in three (31.3% [27.0, 35.6]) contains at least one verified failure, and the failures concentrate in allergy and medication information, invented patient identity, and - on telephone consultations that could contain none - history written up as physical examination. No product here was given a patient record; set aside the two classes one would have prefilled, invented identities and dates, and the rate is 24.8% [20.8, 29.0]. The findings organise into a three-tier taxonomy whose classes come from the published scribe-error taxonomies, plus one failure mode it could not place: a treatment the clinician explicitly retracts recorded as delivered care. Two clinicians adjudicated blind, non-overlapping samples. A physician author upheld 20 of 21 verified findings (95.2% [77.3, 99.2]), and an independent clinician, not an author and with no involvement in the study, upheld 12 of 12 ([75.8, 100]). Both judged every sampled refusal genuine. A failure rate, however, depends on the instrument that counted it as much as on the scribes, and we measured its share. With the model, the evidence and every setting held fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the model family moves the headline: run alone at the same strict instruction, the gentler family flags 54.8% of the sampled notes against the harsher one's 27.8%, roughly double. Published audits disagree by a margin instrument differences alone can produce: omission is 54-86% of errors across them and 23.1% of ours, and the one study counting over notes as we do reports 18% where we find 15.4%. The census shows a signing clinician where attention matters most, and gives a buyer the sharper question: an error rate, under what instrument? We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline, so the census can be repeated on other products.