Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 4 of 5
Prof. Dr. Utku Kose, Süleyman Demirel University
Part A turns the weight of the Kullback-Leibler (KL) penalty on real outputs of the notebook and shows the optimum of the penalised objective. Part B names the stages of reinforcement learning from human feedback (RLHF). Part C computes a Bradley-Terry preference and a penalised reward by hand.
The candidates below are real outputs of the notebook: The most likely sentences of the fine-tuned model and every sentence that the reinforcement-learning runs produced outside the grammar. For each one the lab knows its probability under the reference model, its reward-model score and its gold reward. For a KL weight beta, the optimal policy is proportional to the reference probability times exp(reward-model score / beta). Lower beta and watch what it buys.
This is the optimum of the KL-regularised objective from the lecture, computed over the 392 candidates only, with their reference probabilities rescaled to add up to one. The reference model is the fine-tuned model, and beta is the weight of the Kullback-Leibler (KL) penalty. The horizontal axis shows the logarithm of beta to base 10: -1 stands for 0.1 and 0 for 1. The dashed line, the share of probability on sentences of the grammar, is drawn at four times its value, so the height 4 means 100 percent.
| Probability | Output | Reward model | Gold | In the grammar |
|---|
Continue in Colab, section 7: the KL weight as a dial, and reward hacking.
Name the stage of RLHF that each description refers to.
Continue in Colab, section 5: a Bradley-Terry reward model from noisy pairs.
The reward model scores answer A with 1.2 and answer B with 0.4. During RLHF, an answer receives the reward 2.0 from the reward model, the log-ratio of its probability under the new and the fine-tuned model is 3.0, and the KL weight is 0.2.
Answer each question, rate your confidence and check the answer. Results stay in this browser.
Write a short answer to each question. The text is saved in this browser and is included when you export the learning log.