Skip to content
Kudos AI

Belief States and the Update

Replacing the unknown state with a distribution over states, the filtering step that maintains it, and the reason a noisy sensor cannot drive that distribution to certainty however long you watch.

AdvancedModule 125 min · 100 XP
A hidden state behind a curtain with a distribution drawn in front of it, the distribution sliding as actions and percepts arrive, and repeated confirming readings pushing it up to a ceiling it never passes.

Every decision problem so far has let the agent see where it is. Markov decision processes gave a policy π(s)\pi(s) - look up the state, read off the action. This lesson removes the lookup: the state is hidden, and only noisy percepts of it arrive.

One addition, and everything changes

A partially observable MDP has everything an MDP has - a transition model P(s′∣s,a)P(s' \mid s, a), actions, a reward function - plus one thing: a sensor model P(e∣s)P(e \mid s) giving the probability of perceiving evidence ee in state ss.

That looks like a small addition and it breaks the central object. A policy cannot be a function of the state any more, because the agent cannot evaluate it. Worse, the optimal action no longer depends only on where the agent is: it depends on how much it knows. Two agents standing in the same state, one certain and one confused, should often do different things.

The object that replaces the state

The fix is to reason about the distribution over states consistent with everything done and perceived so far. That distribution is the belief state, written bb, with b(s)b(s) the probability of being in state ss.

The move that makes this work is almost too simple to notice: the belief state is always observable to the agent. It is a summary of the agent's own history, not of the world. So a policy π(b)\pi(b) is something the agent can actually execute, where π(s)\pi(s) is not.

It is also a sufficient statistic for the history. Because the transition and sensor models are Markov - the next state depends only on the current state and action, the percept only on the current state - everything the past has to say about the future is already inside the current distribution. The agent can therefore discard its history and carry only bb, which is what makes planning possible at all.

Maintaining it

Given a belief, an action aa, and the percept ee that follows, the new belief is

b′(s′)=α  P(e∣s′)∑sP(s′∣s,a) b(s).b'(s') = \alpha \; P(e \mid s') \sum_{s} P(s' \mid s, a) \, b(s) .

Three steps. Predict: push the belief through the transition model for the action taken. Weight: multiply each candidate successor by how likely the observed percept would be there. Normalise: α\alpha makes it sum to one.

This is exactly the forward step of hidden Markov model filtering from the temporal-models path. The only addition is that the transition model now depends on the agent's choice - which means the agent influences not just where it goes but what it will learn.

A world small enough to see

Two states, 0 and 1, with R(0)=0R(0) = 0 and R(1)=1R(1) = 1. Two actions: Stay keeps the state with probability 0.9, Go switches it with probability 0.9. The sensor reports the correct state with probability 0.6 - barely better than a coin.

Because there are two states, a belief is one number: b(1)b(1), with b(0)b(0) its complement. The whole belief space is the interval [0,1][0, 1].

Start at b(1)=1/2b(1) = 1/2 and act. The results are worth staring at:

actionperceptprobabilitynew b(1)b(1)
Stay01/22/5 = 0.400000
Stay11/23/5 = 0.600000
Go01/22/5 = 0.400000
Go11/23/5 = 0.600000

The action made no difference. From an even belief, Stay sends a uniform prediction to a uniform prediction, and so does Go - one persists with 0.9 and the other switches with 0.9, which from a 50/50 split is the same thing. The percept does all the work, and 0.6 against 0.4 gives exactly 3/5.

That coincidence is worth naming because it separates the two jobs an action does in a POMDP: change the world, and change what you know about it. Here the first cancels and only the second is visible. In general both matter at once, and an action can be worth taking purely for what it reveals.

Confidence is easier to lose than to gain

Push a confident belief against one contradicting reading:

b(1)b(1) beforeafter Stay and percept 0
0.52/5 = 0.400000
0.982/109 = 0.752294
0.99446/527 = 0.846300

A belief of 0.99 falls to 0.846 in a single step, but the sensor does less of that than it looks. The Stay leak alone takes 0.99 to a prediction of 0.892, so about ten of the fourteen points are gone before the percept arrives; percept 0 then removes a further 0.046, and even a confirming percept 1 would leave the belief at 0.925. Now push the other way, taking Stay and seeing percept 1 over and over from b(1)=1/2b(1) = 1/2:

0.600,  0.674,  0.727,  0.762,  0.786,  0.801,  0.811,  0.817,…0.600, \; 0.674, \; 0.727, \; 0.762, \; 0.786, \; 0.801, \; 0.811, \; 0.817, \ldots

Eight consecutive confirmations reach only 0.817218. The climb is slowing, and it does not stop at 1.

Below is that whole belief space - the interval - with the belief on it as a single point. Press “Stay, saw 1” and keep pressing. The steps shrink visibly, the climb bends towards the dashed line, and it stops there. Then press “Stay, saw 0” once from up near the ceiling: a single reading from a sensor that is wrong 40% of the time costs more than two confirmations bought. That asymmetry is not the sensor being one-sided: each percept multiplies or divides the odds by the same 0.6/0.4=1.50.6 / 0.4 = 1.5, whatever the belief. It is the Stay leak, which up near the ceiling has already cancelled most of what a confirmation buys.

Interactive: the whole belief, on one line

Two states, so a belief is one number. Try to reach certainty.

00.510.82790.5000
b(1)
0.500000
After acting, before seeing
0.5000
P(next percept = 1)
0.5000
Ceiling
0.827934

From an even belief, Stay and Go do exactly the same thing: one persists with 0.9 and the other switches with 0.9, and from a 50/50 split those are the same operation. The percept does all the work, and 0.6 against 0.4 gives exactly 3/5. That coincidence separates the two jobs an action has in a POMDP - change the world, and change what you know about it. Here the first cancels and only the second is visible.

The ceiling

Iterating that update converges to 0.8279344230.827934423, and solving the fixed-point equation symbolically gives the closed form

b∗=3+10516=0.827934.b^{*} = \frac{3 + \sqrt{105}}{16} = 0.827934 .

Two forces balance there. Each percept pulls the belief outward, because the sensor carries information. Each transition pushes it back toward the middle, because Stay leaks 0.1 of probability to the other state every step. The fixed point is where they exactly cancel, and past it more evidence buys nothing because it is arriving no faster than the world is forgetting.

The consequence is the whole subject. Certainty is unreachable, so an agent that waits to find out where it is waits forever. It has to plan while still uncertain - and the belief, not the state, is what it plans over.

The tempting shortcut, and why it fails

A natural reaction to all this is to skip it: track the belief, take the most probable state, and hand that to an ordinary MDP policy. It is cheap, it is easy to implement, and it is wrong in a specific way worth understanding.

An agent that commits to its best guess treats its own uncertainty as noise to be rounded away. It will therefore never value an action for what the action would reveal. Offer it a cheap sensing move - one that produces no reward but resolves an ambiguity - and it sees an action with no benefit, because the state it believes itself to be in does not change. So it never looks before it leaps, and it will confidently execute a plan that is right in the state it guessed and disastrous in the one it dismissed at 45%.

The belief formulation does not have this hole, because the belief is the state. An action that sharpens the distribution moves the agent to a better belief-state, and the machinery values that move like any other. The value of information, which the decision-theory path treated as its own calculation, arrives here for free as a consequence of using the right state variable.

Before the quiz

A POMDP is an MDP plus a sensor model, and that one addition means a policy cannot be indexed by state. The belief state replaces it: a distribution over states, always observable to the agent, and a sufficient statistic for the whole history. It is maintained by predict-weight-normalise, the same recursion as HMM filtering with the action chosen by the agent. And with a noisy sensor it saturates - here at exactly (3+105)/16(3 + \sqrt{105})/16 - so planning under uncertainty is not a stopgap but the permanent condition.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.