Transcription of Abstract - arxiv.org
{{id}} {{{paragraph}}}
On-line Active Reward Learning for Policy Optimisationin Spoken Dialogue SystemsPei-Hao Su, Milica Ga si c, Nikola Mrk si c, Lina Rojas-Barahona,Stefan Ultes, David Vandyke, Tsung-Hsien Wen and Steve YoungDepartment of Engineering, University of Cambridge, Cambridge, UK{phs26, mg436, nm480, lmr46, su259, djv27, thw28, ability to compute an accurate re-ward function is essential for optimisinga dialogue policy via reinforcement learn-ing. In real-world applications, using ex-plicit user feedback as the reward sig-nal is often unreliable and costly to col-lect.}
strapping estimates of sparse value functions from minimal numbers of samples (dialogues). The quality of each dialogue is defined by its cumu-lative reward, where each dialogue turn incurs a small negative reward (-1) and the final reward of either 0 or 20 depending on the estimate of task success are provided by the reward model.
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}