Reinforcement Learning and Language Model Alignment (VTR UGE 21), day 2 of 5

Value-Based Methods

Prof. Dr. Utku Kose, Süleyman Demirel University

Value-based lab

Part A explores the deadly triad on Baird's counterexample. Part B classifies methods as on-policy or off-policy, temporal difference, Monte Carlo or dynamic programming. Part C computes one update of Q-learning and one of SARSA (state, action, reward, state, action) by hand.

Part A: The deadly-triad lab

Baird's counterexample has seven states, a reward of zero everywhere and therefore a true value function of zero. Each switch below removes one leg of the deadly triad: function approximation, bootstrapping or off-policy sampling. Run every combination and find the only one that diverges.



Scoreboard

RunApproximationBootstrappingOff-policy Step sizeFinal normVerdict

Open in ColabContinue in Colab, section 7: Baird's counterexample with each leg removed.

Part B: Which kind of method?

Classify each method.

Open in ColabContinue in Colab, section 3: SARSA and Q-learning against the exact optimum.

Part C: One Q-learning and one SARSA update by hand

The current value is Q(s, a) = 2.0. The agent receives the reward -1 and lands in state s'. The best action in s' has the value 4.0, but the agent explores and will take an action whose value is 1.0. The discount factor is 0.9 and the step size 0.5.

Open in ColabContinue in Colab, section 4: the two methods on the cliff.