Q-Learning and Exploration
Learning to act well with no model of the environment: the temporal-difference update, value propagating backwards, and why a purely greedy agent gets stuck.
AdvancedModule 335 min · 160 XP
This is a premium lesson
Sign in and enrol to read the full lesson, run the code, take the quiz, and earn XP toward the path badge.