Example: barber
Bellman Equations and Dynamic Programming
The value function for π is its unique solution. Backup diagrams: ... The terminal state is shaded in the figure (although it is shown in two places, it is formally one state). The expected reward function is thus r(s,a,s0)=1forallstatess,s0 and actions a. Suppose the agent follows
Download Bellman Equations and Dynamic Programming
Information
Domain:
Source:
Link to this page: