Demo: Blocking Silent Safety Regressions in Prompt Optimization for Health AI
Abstract
Reflective prompt optimizers can improve aggregate accuracy while silently losing rare safety behavior, such as correctly escalating a crisis message. We present gepa-guard, which keeps GEPA’s scalar objective but rejects candidates that regress a named check passed by the seed prompt. The guard checks both the sampled rows and a fixed protected floor, rejects unverifiable proposals, and records the failed check for each regression veto. On synthetic health triage data, all three unguarded seeds improved aggregate accuracy while losing 3 to 17 of the 59 protected checks that the seed passed on test (5–29%). The guard vetoed 49 to 58% of proposals, 65 to 132 per seed. One guarded seed beat its unguarded counterpart on all four reported rates, with no emergency losses and two encoding losses on test. The other two returned the seed prompt as nulls. As a result, gepa-guard turns a silent regression into a blocked proposal with a named cause and an auditable record, before any revised prompt reaches deployment. A captioned demonstration video and the run records behind every number accompany the paper.