Transcription of Soft Actor-Critic: Off-Policy Maximum Entropy Deep ...
{{id}} {{{paragraph}}}
soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor Tuomas Haarnoja 1 Aurick Zhou 1 Pieter Abbeel 1 Sergey Levine 1. Abstract networks holds the promise of automating a wide range of Model-free deep reinforcement learning (RL) al- decision making and control tasks, but widespread adoption gorithms have been demonstrated on a range of of these methods in real-world domains has been hampered challenging decision making and control tasks. by two major challenges. First, model-free deep RL meth- However, these methods typically suffer from two ods are notoriously expensive in terms of their sample com- major challenges: very high sample complexity plexity.
sensitivity (Duan et al.,2016;Henderson et al.,2017). We explore how to design an efficient and stable model-free deep RL algorithm for continuous state and action spaces. To that end, we draw on the maximum entropy framework, which augments the standard maximum reward
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}