Example: barber
Reinforcement Learning: Theory and Algorithms
In reinforcement learning, the interactions between the agent and the environment are often described by an infinite-horizon, discounted Markov Decision Process (MDP) M= (S;A;P;r;; ), specified by: •A state space S, which may be finite or infinite. For mathematical convenience, we will assume that Sis finite or countably infinite.
Download Reinforcement Learning: Theory and Algorithms
Information
Domain:
Source:
Link to this page: