PDF4PRO ⚡AMP

Modern search engine that looking for books and documents around the web

Example: barber

Abstract - arxiv.org

On-line Active Reward Learning for Policy Optimisationin Spoken Dialogue SystemsPei-Hao Su, Milica Ga si c, Nikola Mrk si c, Lina Rojas-Barahona,Stefan Ultes, David Vandyke, Tsung-Hsien Wen and Steve YoungDepartment of Engineering, University of Cambridge, Cambridge, UK{phs26, mg436, nm480, lmr46, su259, djv27, thw28, ability to compute an accurate re-ward function is essential for optimisinga dialogue policy via reinforcement learn-ing. In real-world applications, using ex-plicit user feedback as the reward sig-nal is often unreliable and costly to col-lect.}

data is available to pre-train a task suc-cess predictor off-line. In practice neither of these apply for most real world applica-tions. Here we propose an on-line learn-ing framework whereby the dialogue pol-icy is jointly trained alongside the reward model via active learning with a Gaussian process model. This Gaussian process op-

Loading..

Tags:

  Cess, Suc cess

Information

Domain:

Source:

Link to this page:

Please notify us if you found a problem with this document:

Spam in document Broken preview Other abuse

Transcription of Abstract - arxiv.org

Related search queries