Skip to content
Kudos AI

Gridworld RL Lab

Value iteration, policy iteration, and Q-learning on the same gridworld, so a planner that knows the model can be compared directly against a learner that does not.

Python (NumPy)active

A small Markov decision process used as a controlled setting for comparing planning against learning. Value iteration and policy iteration are run first, with full knowledge of the transition model, to compute the optimal value function and policy exactly. Q-learning is then run on the same world with no model at all, learning only from sampled transitions, and its estimates are compared against the planned ground truth - which is what makes the comparison informative: the learner recovers the same optimal policy while its value estimates remain slightly off, and the residual gap is sampling error rather than a bug. The lab also exposes the parameters that actually govern behaviour: the discount factor, the learning rate, and the exploration schedule, with the sweep over epsilon showing why a purely greedy learner can settle on a worse policy. Implementation is in progress and no source repository has been published yet.

Highlights

  • Value iteration and policy iteration run to convergence against a known transition model
  • Q-learning on the identical world with no model, compared against the planned optimum
  • Convergence traced per sweep, so the contraction toward the fixed point is visible
  • Epsilon schedule swept to show a greedy learner converging on a worse policy

Related articles

9 min readReinforcement Learning

Markov Decision Processes

How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.

Reinforcement LearningProbabilityArtificial Intelligence
9 min readReinforcement Learning

Reinforcement Learning and Q-Learning

Learning to act well without a model of the world: temporal-difference updates, the Q-learning rule, exploration versus exploitation, and a run that recovers the planned optimum from experience alone.

Reinforcement LearningMachine LearningArtificial Intelligence