Expected Utility and the MEU Principle
A utility function turns preferences into a number, expected utility averages it over what might happen, and maximising it is the whole criterion - which is not the same as solving the problem.
Every path so far has ended at a belief: a posterior, a filtered estimate, a fitted curve. Beliefs do not tell you what to do. Acting also needs preferences, and decision theory is the arithmetic that combines the two.
Utility and expected utility
A utility function assigns one number to a state, expressing how desirable it is. Because actions in an uncertain world do not determine their outcome, write for the random variable whose values are the possible outcome states. Given evidence , the expected utility of an action is the average utility of its outcomes, weighted by probability:
The maximum expected utility principle is then the whole criterion:
Note what it is not. It is not "maximise the best case", which would ignore risk entirely. It is not "reach a goal", because a decision-theoretic agent has a continuous measure of outcome quality rather than a binary test for success. And it is not "maximise expected money", which is a different quantity and, as the next section shows, frequently the wrong one.
Worked example: the game show
You have won. The host offers $1,000,000 outright, or a fair coin: heads you get nothing, tails you get $2,500,000. The expected monetary value of the gamble is
which beats the sure million. Most people decline anyway. Are they wrong?
Write for possessing total wealth of dollars and let your current wealth be dollars. The two actions have expected utilities
Utility is not proportional to money: the first million changes your life and the second does not change it nearly as much. Suppose you rate your current position at , the middle outcome at , and the top at . Then
so declining is not merely understandable, it is what the principle prescribes.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Nothing here is a criticism of averaging. The average is computed the same way both times. What differs is the scale being averaged, and utility is the one that encodes preference.
The argument is about the shape of a curve, so here is the curve. The three outcomes are points on it, the slider bends it, and the odds never move. The dashed line is the rating the top outcome would need for accepting to be right: 11, at the ratings above. That is worth reading twice - it says the second million and a half would have to add at least as much as the first million did, which is the exact opposite of the diminishing returns the whole argument rests on.
Interactive: the coin, and the curve that decides it
The odds never change. Only the shape of the utility does.
- Expected utility, accept
- 7.00
- Expected utility, decline
- 8.00
- Top rating that would tie
- 11.00
- What MEU prescribes
- decline
- Expected money
- $1,250,000
The gamble is worth $1,250,000 in money against a sure $1,000,000, so money says accept. In utility it is worth 7.00 against 8.00, so the principle says decline. Nothing here is a criticism of averaging - the average is computed the same way both times, and what differs is the scale being averaged. For accepting to be right the top outcome would have to be rated 11.00, which is to say the second million and a half would have to add at least as much as the first million did. At the lesson’s rating it adds 0.33 times as much, which is what diminishing returns MEANS, and why declining is not merely understandable but prescribed.
The same gamble, the opposite answer
The argument is about the shape of your curve, not about caution in general. A billionaire's utility is essentially linear over a few million, so measuring utility in millions gives
and the same principle says take the gamble. Two agents, identical odds, opposite rational choices. Nothing is inconsistent: they have different utility functions, and the theory never claimed there was one.
Why this is not the end of AI
Russell and Norvig observe that MEU "could be seen as defining all of AI", then immediately point out that it does not solve it. Both are right, and the gap is worth being precise about.
MEU says exactly what a rational choice is. What it does not say is how to get the ingredients.
- The probabilities demand a causal model of the world and inference within it - which the earlier path showed is NP-hard in general Bayesian networks.
- The utilities often cannot be read off a state at all. Knowing how good a position is may require searching or planning to see what can be reached from it.
So the principle formalises "do the right thing" and leaves nearly all of the work. That is still valuable: it gives a criterion against which approximations can be judged, which is more than a pile of heuristics offers.
MEU and performance measures. If an agent maximises a utility function that correctly reflects the performance measure by which its behaviour is judged, it achieves the highest possible score for that measure, averaged over the environments it might face. The catch is in correctly reflects.
Before the quiz
Be able to write and state MEU, work the game-show example in both money and utility, explain why a billionaire rationally decides the other way, and name the two ingredients MEU assumes but does not supply.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.