9 min readReinforcement Learning
Markov Decision Processes
How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.
Reinforcement LearningProbabilityArtificial Intelligence