Transcription of Deep Reinforcement Learning with Double Q-learning
{{id}} {{{paragraph}}}
Deep Reinforcement Learning with Double Q-learningHado van HasseltandArthur GuezandDavid SilverGoogle DeepMindAbstractThe popular Q- Learning algorithm is known to overestimateaction values under certain conditions. It was not previouslyknown whether, in practice, such overestimations are com-mon, whether they harm performance, and whether they cangenerally be prevented. In this paper, we answer all thesequestions affirmatively. In particular, we first show that therecent DQN algorithm, which combines Q- Learning with adeep neural network, suffers from substantial overestimationsin some games in the Atari 2600 domain.
using Q-learning (Watkins, 1989), a form of temporal dif-ference learning (Sutton, 1988). Most interesting problems are too large to learn all action values in all states sepa-rately. Instead, we can learn a parameterized value function Q(s;a; t). The standard Q-learning update for the param-eters after taking action At in state St and ...
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}