Online Linear Quadratic Control

Alon Cohen,Avinatan Hassidim,Tomer Koren,Nevena Lazic,Yishay Mansour,Kunal Talwar

Online Linear Quadratic Control

2018

Alon Cohen
Avinatan Hassidim
Tomer Koren
Nevena Lazic
Yishay Mansour
Kunal Talwar

We study the problem of controlling linear time-invariant systems with known noisy dynamics and adversarially chosen quadratic losses. We present the first efficient online learning algorithms in this setting that guarantee $O(\sqrt{T})$ regret under mild assumptions, where $T$ is the time horizon. Our algorithms rely on a novel SDP relaxation for the steady-state distribution of the system. Crucially, and in contrast to previously proposed relaxations, the feasible solutions of our SDP all correspond to "strongly stable" policies that mix exponentially fast to a steady state.

Keywords:

Time horizon
Quadratic equation
Mathematical optimization
Steady state
Exponential growth
Regret
Mathematics

Correction
Cite
Save
Machine Reading By IdeaReader

References

Citations