Transcription of Maximum Entropy Inverse Reinforcement Learning
{{id}} {{{paragraph}}}
Maximum Entropy Inverse Reinforcement LearningBrianD. Ziebart, Andrew Maas, Bagnell,andAnind K. DeySchool of Computer ScienceCarnegie Mellon UniversityPittsburgh, PA research has shown the benefit of framing problemsof imitation Learning as solutions to Markov Decision Prob-lems. This approach reduces Learning to the problem of re-covering a utility function that makes the behavior inducedby a near-optimal policy closely mimic demonstrated behav-ior. In this work, we develop a probabilistic approach basedon the principle of Maximum Entropy . Our approach providesa well-defined, globally normalized distribution over decisionsequences, while providing the same performance guaranteesas existing develop our technique in the context of modeling real-world navigation and driving behaviors where collected datais inherently noisy and imperfect. Our probabilistic approachenables modeling of route preferences as well as a powerfulnew approach to inferring destinations and routes based onpartial problems ofimitation learningthe goal is to learn to pre-dict the behavior and decisions an agent would choose ,the motions a person would take to grasp an object or theroute a driver would take to get from home to work.
Recovering the agent’s exact reward weights is an ill-posed problem; many reward weights, including degenera-cies (e.g., all zeroes), make demonstrated trajectories opti-mal. Ratliff, Bagnell, & Zinkevich (2006) cast this problem as one of structured maximum margin prediction (MMP). They consider a class of loss functions that directly measure
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}