A Policy-Driven Privacy Router for Clinical Text
Abstract
Cloud-based large language models (LLMs) are increasingly useful for clinical documentation and administrative workflows, but clinical text often contains highly sensitive patient disclosures that make uncontrolled cloud processing unsafe. We present a policy-driven privacy router that can decide locally whether a text input may be safely sent to a cloud LLM unchanged, must be obfuscated in some way before being sent to a cloud LLM, or must remain strictly within a local execution environment. The router utilizes deterministic Personally Identifiable Information (PII) detection and quasi-identifier (QI) detection. A lightweight local LLM-driven contextual gate is used only as an exception handler for uncertain cases and analyzes already-masked text for residual contextual risks for re-identification, including references to small communities and rare medical conditions disclosures. We evaluate the pipeline on a modified 600-sample benchmark with balanced critical and non-critical privacy examples and analyze precision-recall, false positive rates, over-censoring, and threshold-sweep sensitivity. The results show how local, policy-aware routing can reduce unnecessary local-only processing while retaining conservative controls for unmasked identifiers, low empirical anonymity, and uncertain contextual disclosure, providing a different modality for LLMs in medicine over a separate healthcare pipeline.