Clinical and Regulatory Workflows as Stress Tests for Interactive-Agent Evaluation
Abstract
Long-horizon agents in expert workflows are difficult to evaluate: expert review is costly, plausible deliverables can hide serious errors, and useful progress depends on the task. We use two high-value life-sciences workflows—regulatory drafting and clinical statistical programming—as stress tests for interactive-agent evaluation. In Investigational New Drug (IND) drafting, leaving a field blank can be correct when the source provides no evidence; in clinical tables, listings, and figures (TLF) programming, executable code can still contain a wrong derivation. Across both tasks, activity counts are weak measures of progress. Termination state, intermediate regressions, repeated-run variation, and harness choice reveal failures that endpoint scores miss. We also find that evaluator design depends on placement: an evaluator that is appropriate for post-hoc scoring can leak privileged information or change behavior when its feedback is returned during refinement. IND judges calibrate the same 0–10 scale differently, while a TLF reviewer panel tracks exact reference-based accuracy when explicitly validated against it. We therefore recommend reporting four layers—outcome, trajectory, reliability, and system configuration—together with evaluator position, inputs, and feedback.