NeurIPS Poster Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems

Poster

Logarithmic Regret Bound in Partially Observable Linear Dynamical Systems

Sahin Lale · Kamyar Azizzadenesheli · Babak Hassibi · Anima Anandkumar

Poster Session 0 #173

[ Abstract ] [ Paper PDF ]

[ Paper ]

Abstract: We study the problem of system identification and adaptive control in partially observable linear dynamical systems. Adaptive and closed-loop system identification is a challenging problem due to correlations introduced in data collection. In this paper, we present the first model estimation method with finite-time guarantees in both open and closed-loop system identification. Deploying this estimation method, we propose adaptive control online learning (AdapOn), an efficient reinforcement learning algorithm that adaptively learns the system dynamics and continuously updates its controller through online learning steps. AdapOn estimates the model dynamics by occasionally solving a linear regression problem through interactions with the environment. Using policy re-parameterization and the estimated model, AdapOn constructs counterfactual loss functions to be used for updating the controller through online gradient descent. Over time, AdapOn improves its model estimates and obtains more accurate gradient updates to improve the controller. We show that AdapOn achieves a regret upper bound of

polylog (T)

$\text{polylog}\left(T\right)$ , after

T

$T$ time steps of agent-environment interaction. To the best of our knowledge, AdapOn is the first algorithm that achieves

polylog (T)

$\text{polylog}\left(T\right)$ regret in adaptive control of \textit{unknown} partially observable linear dynamical systems which includes linear quadratic Gaussian (LQG) control.

Chat is not available.