Paper1992
Q-learning
Christopher Watkins & Peter Dayan
Proves that an agent can converge on optimal action values from raw experience alone, without ever building a model of the environment's dynamics.
link checked 17 Sept 2026FreeAdvanced
Artificial Intelligence · Reinforcement Learning
A topic within Reinforcement Learning, itself one of 11 topics in that field and part of Artificial Intelligence.
2 works
Christopher Watkins & Peter Dayan
Proves that an agent can converge on optimal action values from raw experience alone, without ever building a model of the environment's dynamics.
link checked 17 Sept 2026Mnih et al.
One architecture learns dozens of Atari games from raw pixels and the score, with no per-game features, which is what made reinforcement learning look general.
link checked 17 Sept 2026