Transcription of Maximum Entropy Inverse Reinforcement Learning
{{id}} {{{paragraph}}}
Maximum Entropy Inverse Reinforcement LearningBrianD. Ziebart, Andrew Maas, Bagnell,andAnind K. DeySchool of Computer ScienceCarnegie Mellon UniversityPittsburgh, PA research has shown the benefit of framing problemsof imitation Learning as solutions to Markov Decision Prob-lems. This approach reduces Learning to the problem of re-covering a utility function that makes the behavior inducedby a near-optimal policy closely mimic demonstrated behav-ior. In this work, we develop a probabilistic approach basedon the principle of Maximum Entropy .
normalize locally over each state’s available actions (Ra-machandran & Amir 2007; Neu & Szepesvri 2007). Background In the imitation learning setting, an agent’s behavior (i.e., its trajectory or path, ζ, of states si and actions ai) in some planning space is observed by a learner trying to model or imitate the agent.
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}