Skip to content
Kudos AI

Q-Learning

A reinforcement learning algorithm that learns the value of taking each action in each state directly from experience, without a model of the environment.

Five trips updated by hand, with B forced to learn before A can - then the greedy failure drawn as the closed loop it actually is.

Understanding Q-Learning

A value function over states alone is not directly actionable without a model: knowing that a neighbouring state is valuable does not tell you which action reaches it, unless the transition probabilities are known. A Q-function sidesteps this by attaching values to state-action pairs directly, so the best action in a state is simply the one with the highest Q-value.

Russell and Norvig make the consequence explicit: because a learner holding a Q-function needs no transition model either for learning or for action selection, Q-learning is a model-free method. It relates back to the state value function through the identity that the value of a state is the maximum Q-value over the actions available in it.

Learning proceeds by temporal-difference updates. After taking an action and observing the reward and the resulting state, the agent forms a target: the reward received plus the discounted best Q-value available at the new state. The gap between this target and the current estimate is the temporal-difference error, and the estimate is nudged toward the target by a fraction set by the learning rate.

The algorithm is off-policy, which is its most useful structural property. The update always uses the maximum over next actions, so it learns about the optimal policy regardless of how the agent actually behaved while gathering the experience. That permits deliberate exploration, typically ε-greedy, taking the best known action most of the time and a random one occasionally, without corrupting what is learned.

How to Calculate

Q(s, a) ← Q(s, a) + α [ r + γ max_{a′} Q(s′, a′) − Q(s, a) ]

where

Q(s, a)
estimated return from taking a in s, then acting optimally
α
the learning rate, how far the estimate moves toward the target
r
the reward actually observed for this transition
γ
the discount factor on future value
max_{a′} Q(s′, a′)
the best value available from the resulting state

Example of Q-Learning

In a grid world the agent keeps a table with one entry per state-action pair, initialized arbitrarily. It acts ε-greedily, observes the reward and the next cell, and applies the update. Early on the estimates are meaningless and behaviour looks random.

Information propagates backwards from reward. The first time the goal is reached, the Q-value of the action that entered the goal rises. On a later visit to the state before that one, the update sees a now-larger maximum at the next state, so its own value rises too. Value spreads outward from the goal one step per visit.

This backward propagation is also the method’s practical weakness: with sparse rewards it can take a very large number of episodes for signal to reach the states where the critical early decisions are made. Reward shaping and replay techniques exist largely to accelerate it.

Frequently Asked Questions

What does model-free mean here?

That the agent never needs to know or estimate the probability of reaching one state from another. It learns purely from observed transitions and rewards, which matters because in most realistic environments those probabilities are unavailable.

What is the difference between on-policy and off-policy learning?

Off-policy methods such as Q-learning learn about the optimal policy while following a different, exploratory one, because the update takes the maximum over next actions. On-policy methods such as SARSA learn the value of the policy actually being followed, exploration included.

How does deep Q-learning differ?

A table needs one entry per state-action pair, which is impossible for large or continuous state spaces. Deep Q-networks replace the table with a neural network approximating Q, allowing generalization across similar states at the cost of the stability guarantees the tabular version enjoys.

The Bottom Line

Q-learning estimates the value of each action in each state directly from experience, needing no environment model and tolerating exploratory behaviour while still converging toward the optimal policy. Introduced in Watkins’s 1989 thesis, it remains the conceptual basis for much of modern deep reinforcement learning.