Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 1 of 5
Prof. Dr. Utku Kose, Süleyman Demirel University
Part A is a bandit game in which you play against the strategies of the day. Part B names the ingredients of a decision problem. Part C computes a discounted return and a Bellman backup by hand. Parts D and E repeat two calculations of the lecture step by step, the upper confidence bound and value iteration, with sliders for their numbers. In these parts every intermediate value is rounded, and the next step continues with the rounded value, as in a calculation by hand.
Five slot machines pay a reward with hidden probabilities. You have fifty pulls: Click a machine to pull it. An algorithm plays the same machines with the same number of pulls. At the end the true probabilities are revealed and the game is added to the scoreboard. Can you beat the algorithm?
The three opponents are the strategies of the lecture, where their formulas are given. Epsilon-greedy plays the machine with the best record so far and tries a random one in one pull out of ten. UCB, the upper confidence bound, adds a bonus to machines that it has tried rarely. Thompson sampling draws a guess for every machine from what it has seen and plays the best guess.
| Game | Opponent | Your reward | Its reward | Your regret | Its regret |
|---|
Continue in Colab, section 1: four strategies over thirty seeds.
Name the ingredient that each description refers to.
Continue in Colab, section 2: TokenWorld as a decision process.
An episode gives the rewards 0, 0 and 1, in this order, and the discount factor is 0.9. In another state, an action leads with probability 0.8 to a state of value 10 and with probability 0.2 to a state of value 0, and every step costs 1.
Continue in Colab, section 3: value iteration on the Windy Cliff.
The worked example of the lecture: three arms, their pulls and their rewards, and the decision of greedy and of the upper confidence bound (UCB) for the next pull. The sliders set the pulls and the rewards of each arm and the constant \(c\). The figure stacks the bonus of each arm on its estimate.
Continue in Colab, section 1: the same rule over 1500 pulls and thirty seeds.
The corridor of the lecture: two cells, A and B, and a goal behind B. Each sweep computes the values of B and A from the values of the previous sweep, rounded to three decimals. The sweeps stop when no value changes by more than 0.001. The sliders set the probability that forward moves the agent, the discount factor and the reward of a step that does not reach the goal.
Continue in Colab, section 3: value iteration on the 32 cells of the Windy Cliff.
Answer each question, rate your confidence and check the answer. Results stay in this browser.
Write a short answer to each question. The text is saved in this browser and is included when you export the learning log.