PDF4PRO ⚡AMP

Modern search engine that looking for books and documents around the web

Example: air traffic controller

Abstract - arXiv

Back to document page

Offline reinforcement learning as One BigSequence Modeling ProblemMichael JannerQiyang LiSergey LevineUniversity of California at Berkeley{janner, learning (RL) is typically concerned with estimating stationarypolicies or single-step models, leveraging the Markov property to factorize prob-lems in time. However, we can also view RL as a generic sequence modelingproblem, with the goal being to produce a sequence of actions that leads to asequence of high rewards. Viewed in this way, it is tempting to consider whetherhigh-capacity sequence prediction models that work well in other domains, suchas natural-language processing, can also provide effective solutions to the RLproblem.}

quence model can be applied to reinforcement learning problems without the need for the components usually associated with RL algorithms. 3 Reinforcement Learning and Control as Sequence Modeling In this section, we describe the training procedure for our sequence model and discuss how it can be used for control.

  Control, Learning, Reinforcement, Reinforcement learning, Reinforcement learning and control

Download Abstract - arXiv


Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Spam in document Broken preview Other abuse

Related search queries