Skip to content
Kudos AI

Q-Learning and Exploration

Learning to act well with no model of the environment: the temporal-difference update, value propagating backwards, and why a purely greedy agent gets stuck.

AdvancedModule 335 min · 160 XP
Five trips updated by hand, with B forced to learn before A can - then the greedy failure drawn as the closed loop it actually is.

This is a premium lesson

Sign in and enrol to read the full lesson, run the code, take the quiz, and earn XP toward the path badge.