Start with a demo of the basic UCB algorithm. Use the slider to change the confidence parameter and click "Step" to advance the algorithm.
An asymptotically optimal way to set the confidence parameter is to use the following confidence parameter.We can demonstrate this choice by plotting its total regret over time.
This next experiment analyzes the effect of the optimality gap on the UCB algorithm. We switch to a Gaussian bandit setting to allow for larger optimality gaps. Each trial uses a Gaussian bandit with two arms that both have unit variance. The optimal arm has zero mean.
We can also examine the effect of changing the number of arms. Here are two such experiments that use Bernoulli bandits. The first contains one arm with an expected reward of . We then introduce other arms, each with an optimal expected reward of . If we run our algorithm for a fixed number of iterations, the total regret varies based on the number of arms.
Instead of varying the number of optimal arms, we can fix one optimal arm and vary the number of suboptimal arms. For our experiment, we keep one arm with an expected reward of and introduce a varying number of arms with an expected reward of .