Example: barber
Bellman Equations and Dynamic Programming

Bellman Equations and Dynamic Programming

Back to document page

The value function for π is its unique solution. Backup diagrams: ... The terminal state is shaded in the figure (although it is shown in two places, it is formally one state). The expected reward function is thus r(s,a,s0)=1forallstatess,s0 and actions a. Suppose the agent follows

  Value, Terminal

Download Bellman Equations and Dynamic Programming


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Other abuse

Advertisement

Related search queries