SRE-World: How ready are agents for On-Call?
Andre Fu ⋅ Malik Drabla ⋅ Meji Abidoye ⋅ Marek Suppa ⋅ Lata Mishra ⋅ Adnan E Assadi
Abstract
We present SRE-World, a benchmark of live incident response on real systems. Each task deploys a production application to an ephemeral Kubernetes cluster, injects a fault that emerges only under sustained load, and gives an agent an operator shell and a fixed time budget. Grading is dual gate: the repair must hold service-level objectives through a graded soak, and safely repair the incident. Across 44 tasks on three substrates, GPT 5.6 solves 29\% of episodes at pass@1 and GLM 5.3 solves 24\%, with performance collapsing on compound faults that require distinguishing two independent causes from one symptom.
Chat is not available.
Successful Page Load