Example: quiz answers
DoubleQ-learning - NeurIPS

DoubleQ-learning - NeurIPS

Back to document page

1 Introduction Q-learning is a popular reinforcement learning algorithm that was proposed by Watkins [1] and can be used to optimally solve Markov Decision Processes (MDPs) [2]. We show that Q-learning’s performance can be poor in stochastic MDPs because of large overestimations of the action val-ues.

  Introduction, Processes, Learning, Stochastic, Markov, Doubleq learning, Doubleq

Download DoubleQ-learning - NeurIPS


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries