Tail-Risk-Aware Stackelberg Learning
Yujing Chen
Abstract
We study Stackelberg learning in which followers use lower-tail CVaR as a reward utility. In cooperative bandits, where both agents share the CVaR objective, we prove a spectral-risk lower bound that gives $\Omega(\sqrt{K(AB-1)/\tau})$ for CVaR and give a Stackelberg-factorized CVaR-UCB algorithm with matching $\widetilde O\sqrt{ABK/\tau})$ regret up to logarithmic factors. Thus CVaR learning admits a minimax-optimal realizable response certificate in bandits. In Markov games, a risk-neutral leader faces a CVaR-sensitive follower. We introduce a certified follower-response oracle based on shortfall planning and decompose true Stackelberg regret into certified oracle regret and benchmark response certification error. The first term is sublinear without coverage, while true regret requires comparison coverage and response compatibility.
Chat is not available.
Successful Page Load