Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 1 of 5

Foundations of Reinforcement Learning

Prof. Dr. Utku Kose, Süleyman Demirel University

Foundations lab

Part A is a bandit game in which you play against the strategies of the day. Part B names the ingredients of a decision problem. Part C computes a discounted return and a Bellman backup by hand. Parts D and E repeat two calculations of the lecture step by step, the upper confidence bound and value iteration, with sliders for their numbers. In these parts every intermediate value is rounded, and the next step continues with the rounded value, as in a calculation by hand.

Part A: The bandit game

Five slot machines pay a reward with hidden probabilities. You have fifty pulls: Click a machine to pull it. An algorithm plays the same machines with the same number of pulls. At the end the true probabilities are revealed and the game is added to the scoreboard. Can you beat the algorithm?

The three opponents are the strategies of the lecture, where their formulas are given. Epsilon-greedy plays the machine with the best record so far and tries a random one in one pull out of ten. UCB, the upper confidence bound, adds a bonus to machines that it has tried rarely. Thompson sampling draws a guess for every machine from what it has seen and plays the best guess.

Scoreboard

GameOpponentYour rewardIts rewardYour regret Its regret

Open in ColabContinue in Colab, section 1: four strategies over thirty seeds.

Part B: The vocabulary of reinforcement learning

Name the ingredient that each description refers to.

Open in ColabContinue in Colab, section 2: TokenWorld as a decision process.

Part C: A return and a Bellman backup by hand

An episode gives the rewards 0, 0 and 1, in this order, and the discount factor is 0.9. In another state, an action leads with probability 0.8 to a state of value 10 and with probability 0.2 to a state of value 0, and every step costs 1.

Open in ColabContinue in Colab, section 3: value iteration on the Windy Cliff.

Part D: The upper confidence bound, step by step

The worked example of the lecture: three arms, their pulls and their rewards, and the decision of greedy and of the upper confidence bound (UCB) for the next pull. The sliders set the pulls and the rewards of each arm and the constant \(c\). The figure stacks the bonus of each arm on its estimate.

Open in ColabContinue in Colab, section 1: the same rule over 1500 pulls and thirty seeds.

Part E: Value iteration on a corridor, sweep by sweep

The corridor of the lecture: two cells, A and B, and a goal behind B. Each sweep computes the values of B and A from the values of the previous sweep, rounded to three decimals. The sweeps stop when no value changes by more than 0.001. The sliders set the probability that forward moves the agent, the discount factor and the reward of a step that does not reach the goal.

Open in ColabContinue in Colab, section 3: value iteration on the 32 cells of the Windy Cliff.