Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 5 of 5

Direct Preference Optimization and Evaluation

Prof. Dr. Utku Kose, Süleyman Demirel University

Evaluation lab

Part A is an evaluation desk with 400 real outputs of four policies, judges and intervals. Part B reads the intervals of six comparisons. Part C computes a win rate and its interval by hand.

Part A: The evaluation desk

The desk holds 400 real outputs of each of four policies, with their gold reward, reward-model score and length. Three policies come from the notebook of this day: the model after supervised fine-tuning (SFT), the policy trained by reinforcement learning from human feedback (RLHF) with a weight of 0.2 on the Kullback-Leibler (KL) penalty, and the policy trained by direct preference optimisation (DPO) with beta 0.1. The fourth is the RLHF policy that Day 4 trained without the KL penalty. Choose two policies, a judge, a rule for ties and a number of comparisons. The desk pairs outputs at random, computes the win rate of A over B and a bootstrap interval, and keeps a log of every evaluation you run. An evaluation is decided at the 5 % level when its 95 % interval does not contain one half.

Evaluation log

ABJudgeTiesnWin rateIntervalDecided

Open in ColabContinue in Colab, section 4: three judges on the same models.

Part B: Is the comparison decided?

Each line reports the win rate of model A against model B with its 95 percent interval.

Open in ColabContinue in Colab, section 5: intervals, sample size and seeds.

Part C: A win rate and its interval by hand

Model A met model B 400 times: A won 230 comparisons, 40 were ties and A lost 130.

Open in ColabContinue in Colab, section 6: your own judge.