Reinforcement Learning: Theory and Algorithms
Reinforcement Learning: Theory and AlgorithmsAlekh AgarwalNan JiangSham M. KakadeWen SunJanuary 31, 2022WORKING DRAFT:Please any typos or errors you appreciate it!iiContents1Fundamentals31 Markov Decision (Infinite-Horizon) Markov Decision Processes . . . . . . . . . . . . . . . . . . . . . . . objective, policies, and values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Consistency Equations for Stationary Policies . . . . . . . . . . . . . . . . . . . . . Optimality Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Markov Decision Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Complexity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Iteration.
In reinforcement learning, the interactions between the agent and the environment are often described by an infinite-horizon, discounted Markov Decision Process (MDP) M= (S;A;P;r;; ), specified by: •A state space S, which may be finite or infinite. For mathematical convenience, we will assume that Sis finite or countably infinite.
Download Reinforcement Learning: Theory and Algorithms
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: