VTR UGE 21 value added course, Vel Tech University, Chennai

Reinforcement Learning and Language Model Alignment

Prof. Dr. Utku Kose

A five-day course from Bellman equations to direct preference optimisation on one environment, with lecture pages, animations, interactive labs, Colab notebooks and a Python track for beginners. Each day starts from its overview, which gives the study path and links to the lecture page, the Colab notebook, the interactive lab and the PDF notes. Progress is stored only in this browser.

Day 1: Foundations of Reinforcement Learning

Learning from interaction: bandits and exploration, Markov decision processes, values and the Bellman equation, and a small world that can be solved exactly.

Day overview | Lecture | Colab notebook | Interactive lab | PDF

Day 2: Value-Based Methods

Learning values without a model: Monte Carlo and temporal differences, the control methods SARSA (state, action, reward, state, action) and Q-learning, the cliff, and a deep Q-network written in NumPy.

Day overview | Lecture | Colab notebook | Interactive lab | PDF

Day 3: Policy Gradient Methods

Optimising the policy directly: the REINFORCE algorithm and its variance, baselines, actor-critic methods, the clipped objective of proximal policy optimisation (PPO), and entropy.

Day overview | Lecture | Colab notebook | Interactive lab | PDF

Day 4: Reinforcement Learning for Language Models

From a language model to an aligned one: fine-tuning, preferences, a learned reward, reinforcement learning from human feedback (RLHF) with a Kullback-Leibler (KL) penalty, and what happens without it.

Day overview | Lecture | Colab notebook | Interactive lab | PDF

Day 5: Direct Preference Optimization and Evaluation

Preference optimisation without a reward model or sampling, compared with reinforcement learning from human feedback (RLHF) on matched data, and the evaluation of aligned models with judges, intervals and seeds.

Day overview | Lecture | Colab notebook | Interactive lab | PDF