Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 3 of 5
Prof. Dr. Utku Kose, Süleyman Demirel University
Part A trains a policy with and without the clip of proximal policy optimisation (PPO) and shows its mean reward and its entropy over the iterations. Part B matches the ingredients of policy gradient methods with their effects. Part C evaluates the clipped objective by hand.
A softmax policy learns to choose among four arms with close win rates, 0.70, 0.74, 0.77 and 0.80. Each iteration collects a small batch of 8 pulls and reuses it for the chosen number of gradient passes. Compare plain policy gradient with the clipped objective of PPO, raise the learning rate and the reuse, and add an entropy bonus. Every run is the mean of eight seeds.
The bench applies the formulas of the lecture to a bandit, in which every pull is an episode of one step. The advantage of a pull is its reward minus the mean reward of the batch, divided by the standard deviation of the rewards of the batch. In every pass, plain policy gradient adds the advantage times the gradient of the log-probability of the arm. The clipped objective multiplies this term by the probability ratio and leaves a pull out when its ratio has left the range from 0.8 to 1.2 in the direction that its advantage favours. The entropy bonus adds beta times the slope of the entropy to the mean gradient of the batch. A pass changes the preferences by this sum times the learning rate of the slider times 0.1, and a run has 150 iterations. The curves show the mean reward of the policy, at most 0.80, and its entropy in nats, at most ln 4 = 1.386.
| Run | Objective | Learning rate | Passes | Entropy bonus | Final reward | Worst seed | Final entropy |
|---|
Continue in Colab, section 5: PPO against the plain gradient on TokenWorld.
Match each ingredient of a policy gradient method with its effect.
Continue in Colab, section 4: actor-critic with four values of lambda.
PPO maximises the smaller of two terms: the ratio times the advantage, and the clipped ratio times the advantage, with the ratio clipped to the range from 0.8 to 1.2. This is the clipped objective of the lecture for a single step, with a clip range of 0.2. Compute it for three steps.
Continue in Colab, section 8: how much learning rate the clip absorbs.
Answer each question, rate your confidence and check the answer. Results stay in this browser.
Write a short answer to each question. The text is saved in this browser and is included when you export the learning log.