Example: bankruptcy

Reinforcement Learning: An Introduction - Stanford University

IReinforcement Learning: An IntroductionSecond edition, in progress**Draft**Richard S. Sutton and Andrew G. Bartoc 2014, 2015, 2016A Bradford BookThe MIT PressCambridge, MassachusettsLondon, EnglandiiIn memory of A. Harry KlopfContentsPreface to the First EditionixPreface to the Second EditionxiiiSummary of Notationxvii1 The Reinforcement learning Reinforcement learning .. Examples .. Elements of Reinforcement learning .. Limitations and Scope .. An Extended Example: Tic-Tac-Toe .. Summary .. History of Reinforcement learning .. Bibliographical Remarks .. 23I Tabular Solution Methods252 Multi-arm Ak-Armed Bandit Problem .. Action-Value Methods .. Incremental Implementation .. Tracking a Nonstationary Problem .. Optimistic Initial Values .. Upper-Confidence-Bound Action Selection .. Gradient Bandit Algorithms .. Associative Search (Contextual Bandits) .. Summary.

learning system, or, as we would say now, the idea of reinforcement learning. Like others, we had a sense that reinforcement learning had been thoroughly ex- plored in the early days of cybernetics and arti cial intelligence.

Tags:

  Introduction, Learning, An introduction, Reinforcement, Reinforcement learning

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Transcription of Reinforcement Learning: An Introduction - Stanford University

1 IReinforcement Learning: An IntroductionSecond edition, in progress**Draft**Richard S. Sutton and Andrew G. Bartoc 2014, 2015, 2016A Bradford BookThe MIT PressCambridge, MassachusettsLondon, EnglandiiIn memory of A. Harry KlopfContentsPreface to the First EditionixPreface to the Second EditionxiiiSummary of Notationxvii1 The Reinforcement learning Reinforcement learning .. Examples .. Elements of Reinforcement learning .. Limitations and Scope .. An Extended Example: Tic-Tac-Toe .. Summary .. History of Reinforcement learning .. Bibliographical Remarks .. 23I Tabular Solution Methods252 Multi-arm Ak-Armed Bandit Problem .. Action-Value Methods .. Incremental Implementation .. Tracking a Nonstationary Problem .. Optimistic Initial Values .. Upper-Confidence-Bound Action Selection .. Gradient Bandit Algorithms .. Associative Search (Contextual Bandits) .. Summary.

2 42iiiivCONTENTS3 Finite Markov Decision The Agent Environment Interface .. Goals and Rewards .. Returns .. Unified Notation for Episodic and Continuing Tasks .. 54 The Markov Property .. Markov Decision Processes .. Value Functions .. Optimal Value Functions .. Optimality and Approximation .. Summary .. 734 Dynamic Policy Evaluation .. Policy Improvement .. Policy Iteration .. Value Iteration .. Asynchronous Dynamic Programming .. Generalized Policy Iteration .. Efficiency of Dynamic Programming .. Summary .. 955 Monte Carlo Monte Carlo Prediction .. Monte Carlo Estimation of Action Values .. Monte Carlo Control .. Monte Carlo Control without Exploring Starts .. Off-policy Prediction via Importance Sampling .. Incremental Implementation .. Off-Policy Monte Carlo Control .. 118 Return-Specific Importance Sampling.

3 Summary .. 1236 Temporal-Difference TD Prediction .. Advantages of TD Prediction Methods .. Optimality of TD(0) .. Sarsa: On-Policy TD Control .. Q- learning : Off-Policy TD Control .. Expected Sarsa .. Maximization Bias and Double learning .. Games, Afterstates, and Other Special Cases .. Summary .. 1467 Multi-step TD Prediction .. Sarsa .. Off-policy learning by Importance Sampling .. Off-policy learning Without Importance Sampling:Then-step Tree Backup Algorithm .. 160 A Unifying Algorithm:n-stepQ( ) .. Summary .. 1658 Planning and learning with Tabular Models and Planning .. Dyna: Integrating Planning, Acting, and learning .. When the Model Is Wrong .. Prioritized Sweeping .. Planning as Part of Action Selection .. Heuristic Search .. Monte Carlo Tree Search .. Summary .. 186II Approximate Solution Methods1899 On-policy Prediction with Value-function Approximation.

4 The Prediction Objective (MSVE) .. Stochastic-gradient and Semi-gradient Methods .. Linear Methods .. Feature Construction for Linear Methods .. Basis .. Coding .. Coding .. Basis Functions .. Nonlinear Function Approximation: Artificial Neural Networks .. Least-Squares TD .. Summary .. 22210 On-policy Control with Episodic Semi-gradient Control .. Semi-gradient Sarsa .. Average Reward: A New Problem Setting for Continuing Tasks .. Deprecating the Discounted Setting .. Differential Semi-gradient Sarsa .. Summary .. 24011 Off-policy Methods with Semi-gradient Methods .. Baird s Counterexample .. The Deadly Triad .. 24912 Eligibility The -return .. TD( ) .. An On-line Forward View .. True Online TD( ) .. Dutch Traces in Monte Carlo learning .. 26313 Policy Gradient Policy Approximation and its Advantages .. The Policy Gradient Theorem .. REINFORCE: Monte Carlo Policy Gradient.

5 REINFORCE with Baseline .. Actor-Critic Methods .. Policy Gradient for Continuing Problems (Average Reward Rate) .. Policy Parameterization for Continuous Actions .. 278 III Looking Deeper28014 Terminology .. Prediction and Control .. Classical Conditioning .. The Rescorla-Wagner Model .. The TD Model .. TD Model Simulations .. Instrumental Conditioning .. Delayed Reinforcement .. Cognitive Maps .. Habitual and Goal-Directed Behavior .. Summary .. Conclusion .. and Historical Remarks .. 31515 Neuroscience Basics .. Reward Signals, Reinforcement Signals, Values, and Prediction Errors The Reward Prediction Error Hypothesis .. Dopamine .. Experimental Support for the Reward Prediction Error Hypothesis .. TD Error/Dopamine Correspondence .. Neural Actor-Critic .. Actor and Critic learning Rules .. Hedonistic Neurons .. Reinforcement learning .. Methods in the Brain.

6 And Historical Remarks .. 35716 Applications and Case TD-Gammon .. Samuel s Checkers Player .. The Acrobot .. Watson s Daily-Double Wagering .. Optimizing Memory Control .. Human-Level Video Game Play .. Mastering the Game of Go .. Personalized Web Services .. Thermal Soaring .. 39917 The Unified View .. 403 References407 Preface to the First EditionWe first came to focus on what is now known as Reinforcement learning in late were both at the University of Massachusetts, working on one of the earliestprojects to revive the idea that networks of neuronlike adaptive elements might proveto be a promising approach to artificial adaptive intelligence. The project exploredthe heterostatic theory of adaptive systems developed by A. Harry Klopf. Harry swork was a rich source of ideas, and we were permitted to explore them criticallyand compare them with the long history of prior work in adaptive systems. Ourtask became one of teasing the ideas apart and understanding their relationshipsand relative importance.

7 This continues today, but in 1979 we came to realize thatperhaps the simplest of the ideas, which had long been taken for granted, had receivedsurprisingly little attention from a computational perspective. This was simply theidea of a learning system thatwantssomething, that adapts its behavior in order tomaximize a special signal from its environment. This was the idea of a hedonistic learning system, or, as we would say now, the idea of Reinforcement others, we had a sense that Reinforcement learning had been thoroughly ex-plored in the early days of cybernetics and artificial intelligence. On closer inspection,though, we found that it had been explored only slightly. While Reinforcement learn-ing had clearly motivated some of the earliest computational studies of learning ,most of these researchers had gone on to other things, such as pattern classifica-tion, supervised learning , and adaptive control, or they had abandoned the study oflearning altogether.

8 As a result, the special issues involved in learning how to getsomething from the environment received relatively little attention. In retrospect,focusing on this idea was the critical step that set this branch of research in progress could be made in the computational study of Reinforcement learninguntil it was recognized that such a fundamental idea had not yet been field has come a long way since then, evolving and maturing in several direc-tions. Reinforcement learning has gradually become one of the most active researchareas in machine learning , artificial intelligence, and neural network research. Thefield has developed strong mathematical foundations and impressive computational study of Reinforcement learning is now a large field, with hun-dreds of active researchers around the world in diverse disciplines such as psychology,control theory, artificial intelligence, and neuroscience. Particularly important havebeen the contributions establishing and developing the relationships to the theoryixxPreface to the First Editionof optimal control and dynamic programming.

9 The overall problem of learning frominteraction to achieve goals is still far from being solved, but our understanding ofit has improved significantly. We can now place component ideas, such as temporal-difference learning , dynamic programming, and function approximation, within acoherent perspective with respect to the overall goal in writing this book was to provide a clear and simple account of thekey ideas and algorithms of Reinforcement learning . We wanted our treatment to beaccessible to readers in all of the related disciplines, but we could not cover all ofthese perspectives in detail. For the most part, our treatment takes the point of viewof artificial intelligence and engineering. Coverage of connections to other fields weleave to others or to another time. We also chose not to produce a rigorous formaltreatment of Reinforcement learning . We did not reach for the highest possible levelof mathematical abstraction and did not rely on a theorem proof format.

10 We triedto choose a level of mathematical detail that points the mathematically inclined inthe right directions without distracting from the simplicity and potential generalityof the underlying book is largely self-contained. The only mathematical background assumed isfamiliarity with elementary concepts of probability, such as expectations of randomvariables. Chapter 9 is substantially easier to digest if the reader has some knowledgeof artificial neural networks or some other kind of supervised learning method, but itcan be read without prior background. We strongly recommend working the exercisesprovided throughout the book. Solution manuals are available to instructors. Thisand other related and timely material is available via the the end of most chapters is a section entitled Bibliographical and Histori-cal Remarks, wherein we credit the sources of the ideas presented in that chapter,provide pointers to further reading and ongoing research, and describe relevant his-torical background.


Related search queries