Transcription of Abstract - arxiv.org
{{id}} {{{paragraph}}}
On-line Active Reward Learning for Policy Optimisationin Spoken Dialogue SystemsPei-Hao Su, Milica Ga si c, Nikola Mrk si c, Lina Rojas-Barahona,Stefan Ultes, David Vandyke, Tsung-Hsien Wen and Steve YoungDepartment of Engineering, University of Cambridge, Cambridge, UK{phs26, mg436, nm480, lmr46, su259, djv27, thw28, ability to compute an accurate re-ward function is essential for optimisinga dialogue policy via reinforcement learn-ing. In real-world applications, using ex-plicit user feedback as the reward sig-nal is often unreliable and costly to col-lect.}
data is available to pre-train a task suc-cess predictor off-line. In practice neither of these apply for most real world applica-tions. Here we propose an on-line learn-ing framework whereby the dialogue pol-icy is jointly trained alongside the reward model via active learning with a Gaussian process model. This Gaussian process op-
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}