Policy Gradient Methods for Reinforcement Learning with ...
Advances in Neural Information Processing Systems 12, pp. 1057{1063, MIT Press, 2000Policy Gradient Methods forReinforcement Learning with FunctionApproximationRichard S. Sutton, David McAllester, Satinder Singh, Yishay MansourAT&T Labs { Research, 180 Park Avenue, Florham Park, NJ 07932AbstractFunction approximation is essential to Reinforcement Learning , butthe standard approach of approximating a value function and deter-mining a Policy from it has so far proven theoretically this paper we explore an alternative approach in which the policyis explicitly represented by its own function approximator, indepen-dent of the value function, and is updated according to the gradientof expected reward with respect to the Policy parameters. Williams'sREINFORCE method and actor{critic Methods are examples of thisapproach.}}}
The second formulation we cover is that in which there is a designated start state s 0, and we care only about the long-term reward obtained from it. We will give our results only once, but they will apply to this formulation as well under the deflnitions ...
Download Policy Gradient Methods for Reinforcement Learning with ...
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: