RL Course by David Silver — Lecture 2: Markov Decision Processes
David Silver (Google DeepMind)
Builds Markov decision processes up from Markov chains and reward processes, defining the value function and Bellman equation that the rest of the course's algorithms all solve.
link checked 17 Sept 2026