Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI Evaluation
Abstract
Auditing AI deployed in the public interest requires unbiased estimates with valid uncertainty, under affordable budgets for auditing institutions. Expert annotation is costly; LLM-as-judge proxies are biased. Prediction-powered inference (PPI) combines both into debiased estimates whose intervals stay valid however poor the proxy, but its methods are scattered across papers under partial implementations. We introduce GLIDE, an open-source Python library unifying PPI estimators and samplers under a scipy-style API specialized to mean estimation, covering stratified, clustered, non-uniform and multi-proxy designs under both central-limit and bootstrap confidence intervals. The same primitives extend to post-deployment safety monitoring through anytime-valid confidence sequences, so a metric can be re-checked after every batch without inflating false-alarm probability. GLIDE ships with a Monte Carlo validation suite and an agent-safety case study showing substantial annotation savings at equivalent precision.