Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 2 of 5
Prof. Dr. Utku Kose, Süleyman Demirel University
Part A explores the deadly triad on Baird's counterexample. Part B classifies methods as on-policy or off-policy, temporal difference, Monte Carlo or dynamic programming. Part C computes one update of Q-learning and one of SARSA (state, action, reward, state, action) by hand.
Baird's counterexample has seven states, a reward of zero everywhere and therefore a true value function of zero. Each switch below removes one leg of the deadly triad: function approximation, bootstrapping or off-policy sampling. Run every combination and find the only one that diverges.
| Run | Approximation | Bootstrapping | Off-policy | Step size | Final norm | Verdict |
|---|
Continue in Colab, section 7: Baird's counterexample with each leg removed.
Classify each method.
Continue in Colab, section 3: SARSA and Q-learning against the exact optimum.
The current value is Q(s, a) = 2.0. The agent receives the reward -1 and lands in state s'. The best action in s' has the value 4.0, but the agent explores and will take an action whose value is 1.0. The discount factor is 0.9 and the step size 0.5.
Answer each question, rate your confidence and check the answer. Results stay in this browser.
Write a short answer to each question. The text is saved in this browser and is included when you export the learning log.