Transcription of Reinforcement Learning with Deep Energy-Based …
{{id}} {{{paragraph}}}
Reinforcement Learning with deep Energy-Based PoliciesTuomas Haarnoja* 1 Haoran Tang* 2 Pieter Abbeel1 3 4 Sergey Levine1 AbstractWe propose a method for Learning expressiveenergy-based policies for continuous states andactions, which has been feasible only in tabulardomains before. We apply our method to learn-ing maximum entropy policies, resulting into anew algorithm, called soft Q- Learning , that ex-presses the optimal policy via a Boltzmann dis-tribution. We use the recently proposed amor-tized Stein variational gradient descent to learna stochastic sampling network that approximatessamples from this distribution.
Reinforcement Learning with Deep Energy-Based Policies 2007;Ziebart et al.,2008), which are covered in more de-tail inSection 4. Note that this objective differs qualita-
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}