Transcription of Soft Actor-Critic: Off-Policy Maximum Entropy Deep ...
{{id}} {{{paragraph}}}
Soft Actor-Critic: Off-Policy Maximum Entropy Deep ReinforcementLearning with a Stochastic ActorTuomas Haarnoja1 Aurick Zhou1 Pieter Abbeel1 Sergey Levine1 AbstractModel-free deep reinforcement learning (RL) al-gorithms have been demonstrated on a range ofchallenging decision making and control , these methods typically suffer from twomajor challenges: very high sample complexityand brittle convergence properties, which necessi-tate meticulous hyperparameter tuning. Both ofthese challenges severely limit the applicabilityof such methods to complex, real-world this paper, we propose soft actor-critic, an Off-Policy actor-critic deep RL algorithm based on themaximum Entropy reinforcement learning frame-work. In this framework, the actor aims to maxi-mize expected reward while also maximizing en-tropy. That is, to succeed at the task while actingas randomly as possible.
1. Introduction Model-free deep reinforcement learning (RL) algorithms have been applied in a range of challenging domains, from games (Mnih et al.,2013;Silver et al.,2016) to robotic control (Schulman et al.,2015). The combination of RL and high-capacity function approximators such as neural networks holds the promise of automating a wide range of
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}
Maximum Entropy Inverse Reinforcement Learning, Introduction, Learning, Network, Reinforcement, Reinforcement learning, Deep Reinforcement Learning Framework for News, Representation Learning, Community-reinforcement, Growing Success: Assessment, Evaluation and Reporting, INTRODUCTION MACHINE LEARNING, Machine Learning