VTR UGE 21 value added course, Vel Tech University, Chennai
Prof. Dr. Utku Kose
A five-day course from Bellman equations to direct preference optimisation on one environment, with lecture pages, animations, interactive labs, Colab notebooks and a Python track for beginners. Each day starts from its overview, which gives the study path and links to the lecture page, the Colab notebook, the interactive lab and the PDF notes. Progress is stored only in this browser.
Learning from interaction: bandits and exploration, Markov decision processes, values and the Bellman equation, and a small world that can be solved exactly.
Day overview | Lecture | Colab notebook | Interactive lab | PDF
Learning values without a model: Monte Carlo and temporal differences, the control methods SARSA (state, action, reward, state, action) and Q-learning, the cliff, and a deep Q-network written in NumPy.
Day overview | Lecture | Colab notebook | Interactive lab | PDF
Optimising the policy directly: the REINFORCE algorithm and its variance, baselines, actor-critic methods, the clipped objective of proximal policy optimisation (PPO), and entropy.
Day overview | Lecture | Colab notebook | Interactive lab | PDF
From a language model to an aligned one: fine-tuning, preferences, a learned reward, reinforcement learning from human feedback (RLHF) with a Kullback-Leibler (KL) penalty, and what happens without it.
Day overview | Lecture | Colab notebook | Interactive lab | PDF
Preference optimisation without a reward model or sampling, compared with reinforcement learning from human feedback (RLHF) on matched data, and the evaluation of aligned models with judges, intervals and seeds.
Day overview | Lecture | Colab notebook | Interactive lab | PDF