Verification Turns Capability Gaps into Routing Decisions: Small Language Models for Clinical Documentation
Abstract
Clinical note generation from doctor–patient dialogue is high-volume and hallucination-intolerant: a fabricated laboratory value is a patient-safety event, not a quality regression. We ask whether an agentic verification architecture, rather than generator scale, can make a small language model viable here. Holding a four-evaluator verification stack and refinement policy fixed and swapping only the generator between a 4B-parameter model with a QLoRA adapter and a frontier baseline, the small model reaches 0.984 claim-level factual entailment against 0.992 and comes within 3.3 points on completeness, generating entirely on-device. Its deficit is confined to inter-section structural fidelity — arrangement rather than clinical comprehension, and the axis on which it also falls furthest below the corpus's reference notes. Because that deficit is an explicit verifier decision rather than a silent quality difference, it becomes a routing policy: run the small model by default, escalate what the verifier rejects. Measured end to end, that policy matches frontier-only acceptance (99.6% vs. 98.4% over 250 encounters) while routing 62, documenting three quarters without a frontier generation call. Ablating the loop separates architecture from scale: unrefined, the small model is accepted on 26.8% of encounters and the frontier model on 74.8%, so verification does the most work where the generator is weakest. We also contribute section-scoped refinement, which preserves unflagged sections by construction rather than instruction, and find the baseline is not uniformly stronger — its provider refused two trauma-related transcripts the on-device model handled and certified, one facet of a provider dependency a served adapter does not carry.