CausalOMOP: An Auditable, Human-Governed System for Causal Study Design and OMOP Operationalization
Abstract
Generative models can draft a clinical causal study quickly, but fluent prose is not a valid causal analysis: a target trial, a directed acyclic graph (DAG), an OMOP operationalization, and a statistical estimand each encode assumptions that cannot be judged from writing quality. Demonstration paper. We present CausalOMOP, a working system that keeps these scientific objects as separate, versioned, reviewable artifacts and routes every consequential transition through a deterministic check and a human approval gate. A language model may classify a question, draft a target trial and candidate DAG, critique the draft, or write a plain-language results narrative. It can never approve a gate, its only write actions are human-confirmed cards behind a code-level allow-list, and it never receives patient-level data. We demonstrate the system on a running example, proton-pump-inhibitor versus histamine-2-receptor-antagonist stress-ulcer prophylaxis anchored to the PEPTIC trial, showing the seven gates, the advisory-to-data boundary, and the point at which it fails closed rather than produce an unsupported estimate. A separate appendix then carries one exploratory run of the same question the full length of the pipeline, through a reviewer-corrected DAG and the statistical-estimand gate to a prespecified estimate on a materialized cohort, reported as an honest end-to-end result and not a validated finding: a new-user restriction leaves 35 analytic patients, far too few for a clinical conclusion. It is implemented as a command-line workflow, a Streamlit reviewer interface, and a governed OMOP transport, with more than 1,400 automated tests, and ships with three scaffolded reference-anchored study packages and a dated feasibility audit of the OMOP conversion it reads. It makes no clinical-effect claim; the contribution is the system and its demonstrated governance behavior.