Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 4 of 5

Reinforcement Learning for Language Models

Prof. Dr. Utku Kose, Süleyman Demirel University

RLHF lab

Part A turns the weight of the Kullback-Leibler (KL) penalty on real outputs of the notebook and shows the optimum of the penalised objective. Part B names the stages of reinforcement learning from human feedback (RLHF). Part C computes a Bradley-Terry preference and a penalised reward by hand.

Part A: The KL leash on real outputs

The candidates below are real outputs of the notebook: The most likely sentences of the fine-tuned model and every sentence that the reinforcement-learning runs produced outside the grammar. For each one the lab knows its probability under the reference model, its reward-model score and its gold reward. For a KL weight beta, the optimal policy is proportional to the reference probability times exp(reward-model score / beta). Lower beta and watch what it buys.

This is the optimum of the KL-regularised objective from the lecture, computed over the 392 candidates only, with their reference probabilities rescaled to add up to one. The reference model is the fine-tuned model, and beta is the weight of the Kullback-Leibler (KL) penalty. The horizontal axis shows the logarithm of beta to base 10: -1 stands for 0.1 and 0 for 1. The dashed line, the share of probability on sentences of the grammar, is drawn at four times its value, so the height 4 means 100 percent.

The ten most probable outputs at this beta

ProbabilityOutputReward modelGoldIn the grammar

Open in ColabContinue in Colab, section 7: the KL weight as a dial, and reward hacking.

Part B: Which stage of the pipeline?

Name the stage of RLHF that each description refers to.

Open in ColabContinue in Colab, section 5: a Bradley-Terry reward model from noisy pairs.

Part C: A preference and a penalised reward by hand

The reward model scores answer A with 1.2 and answer B with 0.4. During RLHF, an answer receives the reward 2.0 from the reward model, the log-ratio of its probability under the new and the fine-tuned model is 3.0, and the KL weight is 0.2.

Open in ColabContinue in Colab, section 6: RLHF with a KL penalty.