PAPER / ARXIV:2609.19712
Yan Li , Chengze Xie
RESUMO
We study the convergence of the vanilla stochastic policy gradient method applied to the linear quadratic regulator (LQR) problem. The method is cheap in the following sense: (1) at each iteration only $\tilde{O}(1)$ interactions with the environment are needed, therefore allowing frequent policy improvement steps, and (2) to ensure stability throughout and convergence to an $\epsilon$-optimal policy with probability $1-\delta$, only $O(\mathtt{Polylog}(1/\delta)/\epsilon)$ interactions are needed. To the best of our knowledge, this appears to be the first time that a stochastic model-free policy optimization method for LQR converges with high probability with $\tilde{O}(1)$ per-iteration computation and polylogarithmic dependence on the confidence level. The convergence analysis presented here is agnostic to LQR specifics and hence could be potentially generalized to a broader class of problems.
NO MESMO MAPA