Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 3 of 5

Policy Gradient Methods

Prof. Dr. Utku Kose, Süleyman Demirel University

Policy gradient lab

Part A trains a policy with and without the clip of proximal policy optimisation (PPO) and shows its mean reward and its entropy over the iterations. Part B matches the ingredients of policy gradient methods with their effects. Part C evaluates the clipped objective by hand.

Part A: The trust-region bench

A softmax policy learns to choose among four arms with close win rates, 0.70, 0.74, 0.77 and 0.80. Each iteration collects a small batch of 8 pulls and reuses it for the chosen number of gradient passes. Compare plain policy gradient with the clipped objective of PPO, raise the learning rate and the reuse, and add an entropy bonus. Every run is the mean of eight seeds.

The bench applies the formulas of the lecture to a bandit, in which every pull is an episode of one step. The advantage of a pull is its reward minus the mean reward of the batch, divided by the standard deviation of the rewards of the batch. In every pass, plain policy gradient adds the advantage times the gradient of the log-probability of the arm. The clipped objective multiplies this term by the probability ratio and leaves a pull out when its ratio has left the range from 0.8 to 1.2 in the direction that its advantage favours. The entropy bonus adds beta times the slope of the entropy to the mean gradient of the batch. A pass changes the preferences by this sum times the learning rate of the slider times 0.1, and a run has 150 iterations. The curves show the mean reward of the policy, at most 0.80, and its entropy in nats, at most ln 4 = 1.386.

Scoreboard

RunObjectiveLearning ratePassesEntropy bonus Final rewardWorst seedFinal entropy

Open in ColabContinue in Colab, section 5: PPO against the plain gradient on TokenWorld.

Part B: What does each ingredient do?

Match each ingredient of a policy gradient method with its effect.

Open in ColabContinue in Colab, section 4: actor-critic with four values of lambda.

Part C: The clipped objective by hand

PPO maximises the smaller of two terms: the ratio times the advantage, and the clipped ratio times the advantage, with the ratio clipped to the range from 0.8 to 1.2. This is the clipped objective of the lecture for a single step, with a clip range of 0.2. Compute it for three steps.

Open in ColabContinue in Colab, section 8: how much learning rate the clip absorbs.