Policy Regret Minimization in Partially Observable Markov Games
Lan Sang ⋅ Raman Arora ⋅ Thanh Nguyen-Tang
Abstract
We study policy regret minimization in partially observable Markov games (POMGs) between a learner and a strategic adaptive adversary who adapts to the learner's past strategies. We develop a model-based optimistic framework that operates on the learner-observable process using *joint* MLE confidence set and introduce an Observable Operator Model-based causal decomposition that disentangles the coupling between the world and the adversary model. Under multi-step weakly revealing observations and a bounded-memory, stationary and posterior-lipschitz adversary and planner stability, we prove an $\mathcal{O}(\sqrt{T})$ policy regret bound. This work advances regret analysis from Markov games to POMGs and provides the first policy regret guarantee under imperfect information against an adaptive opponent.
Chat is not available.
Successful Page Load