Reinforcement Learning: Theory and Algorithms
Reinforcement Learning: Theory and AlgorithmsAlekh AgarwalNan JiangSham M. KakadeWen SunJanuary 31, 2022WORKING DRAFT:Please any typos or errors you appreciate it!iiContents1Fundamentals31 Markov Decision (Infinite-Horizon) Markov Decision Processes . . . . . . . . . . . . . . . . . . . . . . . objective, policies, and values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Consistency Equations for Stationary Policies . . . . . . . . . . . . . . . . . . . . . Optimality Equations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Markov Decision Processes.
A transition function P: SA! ( S), where ( S) is the space of probability distributions over S(i.e., the probability simplex). P(s0js;a) is the probability of transitioning into state s0upon taking action ain state s. We use P s;ato denote the vector P( s;a). A reward function r: SA!
Download Reinforcement Learning: Theory and Algorithms
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: