Transcription of Reinforcement Learning: Theory and Algorithms
{{id}} {{{paragraph}}}
Reinforcement Learning: Theory and AlgorithmsAlekh AgarwalNan JiangSham M. KakadeWen SunJanuary 31, 2022 WORKING DRAFT:Please any typos or errors you appreciate it!iiContents1 Fundamentals31 Markov Decision (Infinite-Horizon) Markov Decision Processes .. objective, policies, and values .. Consistency Equations for Stationary Policies .. Optimality Equations .. Markov Decision Processes .. Complexity .. Iteration .. Iteration .. Iteration for Finite Horizon MDPs .. Linear Programming Approach .. Complexity and Sampling Models .. : Advantages and The Performance Difference Lemma.
•An action space A, which also may be discrete or infinite. For mathematical convenience, we will assume that Ais finite. •A transition function P: SA! ( S), where ( S) is the space of probability distributions over S(i.e., the probability simplex). P(s0js;a) is the probability of transitioning into state s0upon taking action ain state s ...
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}