ATutorialonThompsonSampling - Stanford University
A Tutorial on Thompson SamplingDaniel J. Russo1, Benjamin Van Roy2, Abbas Kazerouni2, IanOsband3and Zheng Wen41Columbia University2Stanford University3Google DeepMind4Adobe ResearchABSTRACTThompson sampling is an algorithm for online decision prob-lems where actions are taken sequentially in a manner thatmust balance between exploiting what is known to maxi-mize immediate performance and investing to accumulatenew information that may improve future performance. Thealgorithm addresses a broad range of problems in a compu-tationally efficient manner and is therefore enjoying wideuse. This tutorial covers the algorithm and its application,illustrating concepts through a range of examples, includingBernoulli bandit problems, shortest path problems, productrecommendation, assortment, active learning with neuralnetworks, and reinforcement learning in Markov decisionprocesses.
TS, specialized to the case of a beta-Bernoulli bandit, proceeds similarly,aspresentedinAlgorithm2.Theonlydifferenceisthatthe successprobabilityestimate ...
Download ATutorialonThompsonSampling - Stanford University
Information
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document: