Towards Measuring Secure Agent Deployment In The Indian Institutional Context
Abstract
Agentic benchmarks are based on multi-turn QA framework for assessing model capabilities under agentic workload. Current benchmarks are primarily designed with tool-call-user interaction as the primary modality while neglecting subtle contextual, task-specific requirements, and linguistic differences for workflows. These limitations translate to an underestimation of the downside risk of large-scale deployment of agentic systems in digital public infrastructure. We present three proof-of-concept Inspect environments covering a range of specific Indian institutional use cases (Govt. welfare eligibility, UPI dispute triage, and IRCTC rail booking) and present validation baseline metrics on a set of five LLMs. We built a deterministic oracle-based evaluator, reporting per-step success in the agentic loop instead of aggregate success/failure metrics across more than 300 instances. We also introduced a probe for measuring model response to new rules in presence of prior seen stale rules. On our validation set, we could pinpoint the exact failure step in the pipeline due to the detailed environment design. We believe our environments will serve as a framework for future development of high-fidelity agentic environments focused on the needs of the user and provide a more trustworthy evaluation proxy to inform governance policies and vendors for the large-scale deployment of such systems while mitigating harm. The code repo can be found \href{https://anonymous.4open.science/r/evals_neurips2026-DAF1/}{here}.