A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model
This paper proposes novel classical and quantum online reinforcement learning algorithms for finite- and infinite-horizon Markov Decision Processes under a generative model that bypass traditional paradigms like optimism in the face of uncertainty to directly compute optimal policies, achieving improved regret bounds including a polylogarithmic dependence on time steps for quantum methods.