Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 5 of 5
Prof. Dr. Utku Kose, Süleyman Demirel University
Part A is an evaluation desk with 400 real outputs of four policies, judges and intervals. Part B reads the intervals of six comparisons. Part C computes a win rate and its interval by hand.
The desk holds 400 real outputs of each of four policies, with their gold reward, reward-model score and length. Three policies come from the notebook of this day: the model after supervised fine-tuning (SFT), the policy trained by reinforcement learning from human feedback (RLHF) with a weight of 0.2 on the Kullback-Leibler (KL) penalty, and the policy trained by direct preference optimisation (DPO) with beta 0.1. The fourth is the RLHF policy that Day 4 trained without the KL penalty. Choose two policies, a judge, a rule for ties and a number of comparisons. The desk pairs outputs at random, computes the win rate of A over B and a bootstrap interval, and keeps a log of every evaluation you run. An evaluation is decided at the 5 % level when its 95 % interval does not contain one half.
| A | B | Judge | Ties | n | Win rate | Interval | Decided |
|---|
Continue in Colab, section 4: three judges on the same models.
Each line reports the win rate of model A against model B with its 95 percent interval.
Continue in Colab, section 5: intervals, sample size and seeds.
Model A met model B 400 times: A won 230 comparisons, 40 were ties and A lost 130.
Answer each question, rate your confidence and check the answer. Results stay in this browser.
Write a short answer to each question. The text is saved in this browser and is included when you export the learning log.