Models of Delayed Reinforcement Learning
Christopher J. C. H. Watkins · 1989 · PhD thesis, Psychology Department, Cambridge University
Summary
Develops Q-learning, an algorithm that estimates the value of each action in each state directly from experience, requiring no model of the environment’s transition probabilities.
Why it matters
It provided a way to learn optimal behaviour by trial and error without knowing how the environment works, which is the normal situation in practice. Russell and Norvig credit this thesis as where Q-learning was developed, and it remains the conceptual basis for deep Q-networks and much of modern reinforcement learning.