Policy Is Not a Workflow: Can LLMs Reliably Operationalize University AI Governance?
Abstract
Universities are rapidly publishing guidance on the use of generative AI in teaching, research, authorship, and review. However, policy documents alone may not be sufficient for academics to determine what is acceptable in concrete scholarly workflows. This paper assesses whether large language models (LLMs) can support the operationalization of university AI governance policies in realistic academic scenarios, and whether a structured governance checklist improves their decision-making. We introduce a small benchmark of academic AI-use scenarios spanning authorship, citation, peer review, confidential research data, and provenance. We compare varied conditions: scenario-only reasoning, policy-based reasoning, and policy-based reasoning augmented with a workflow checklist covering role, information sensitivity, AI contribution type, human accountability, and disclosure. Model outputs are scored against reference labels derived from publicly available institutional guidance for policy compliance, disclosure correctness, confidentiality recognition, human oversight, unsupported policy claims, and false-permissive recommendations. The study frames AI-native academic governance as an implementation problem: institutions must not only state AI principles, but translate them into auditable workflows that support transparent, accountable, and context-sensitive human-AI scholarship.