Skip to content
Kudos AI

Expected Utility and the MEU Principle

A utility function turns preferences into a number, expected utility averages it over what might happen, and maximising it is the whole criterion - which is not the same as solving the problem.

IntermediateModule 125 min · 100 XP
A game-show choice weighed twice, once in dollars and once in utility, with the two calculations reaching opposite answers on the same coin flip.

Every path so far has ended at a belief: a posterior, a filtered estimate, a fitted curve. Beliefs do not tell you what to do. Acting also needs preferences, and decision theory is the arithmetic that combines the two.

Utility and expected utility

A utility function U(s)U(s) assigns one number to a state, expressing how desirable it is. Because actions in an uncertain world do not determine their outcome, write Result(a)\mathrm{Result}(a) for the random variable whose values are the possible outcome states. Given evidence e\mathbf{e}, the expected utility of an action is the average utility of its outcomes, weighted by probability:

EU(a∣e)=∑s′P(Result(a)=s′∣a,e) U(s′).EU(a \mid \mathbf{e}) = \sum_{s'} P\big(\mathrm{Result}(a) = s' \mid a, \mathbf{e}\big)\, U(s') .

The maximum expected utility principle is then the whole criterion:

action=argmax⁡aEU(a∣e).\text{action} = \operatorname*{argmax}_{a} EU(a \mid \mathbf{e}) .

Note what it is not. It is not "maximise the best case", which would ignore risk entirely. It is not "reach a goal", because a decision-theoretic agent has a continuous measure of outcome quality rather than a binary test for success. And it is not "maximise expected money", which is a different quantity and, as the next section shows, frequently the wrong one.

Worked example: the game show

You have won. The host offers $1,000,000 outright, or a fair coin: heads you get nothing, tails you get $2,500,000. The expected monetary value of the gamble is

12($0)+12($2,500,000)=$1,250,000,\tfrac{1}{2}(\$0) + \tfrac{1}{2}(\$2{,}500{,}000) = \$1{,}250{,}000 ,

which beats the sure million. Most people decline anyway. Are they wrong?

Write SnS_n for possessing total wealth of nn dollars and let your current wealth be kk dollars. The two actions have expected utilities

EU(Accept)=12U(Sk)+12U(Sk+2,500,000),EU(Decline)=U(Sk+1,000,000).EU(\text{Accept}) = \tfrac{1}{2}U(S_k) + \tfrac{1}{2}U(S_{k+2{,}500{,}000}), \qquad EU(\text{Decline}) = U(S_{k+1{,}000{,}000}) .

Utility is not proportional to money: the first million changes your life and the second does not change it nearly as much. Suppose you rate your current position at 55, the middle outcome at 88, and the top at 99. Then

EU(Accept)=12(5)+12(9)=7<8=EU(Decline),EU(\text{Accept}) = \tfrac{1}{2}(5) + \tfrac{1}{2}(9) = 7 < 8 = EU(\text{Decline}) ,

so declining is not merely understandable, it is what the principle prescribes.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Nothing here is a criticism of averaging. The average is computed the same way both times. What differs is the scale being averaged, and utility is the one that encodes preference.

The argument is about the shape of a curve, so here is the curve. The three outcomes are points on it, the slider bends it, and the odds never move. The dashed line is the rating the top outcome would need for accepting to be right: 11, at the ratings above. That is worth reading twice - it says the second million and a half would have to add at least as much as the first million did, which is the exact opposite of the diminishing returns the whole argument rests on.

Interactive: the coin, and the curve that decides it

The odds never change. Only the shape of the utility does.

01.0M2.5M
Expected utility, accept
7.00
Expected utility, decline
8.00
Top rating that would tie
11.00
What MEU prescribes
decline
Expected money
$1,250,000

The gamble is worth $1,250,000 in money against a sure $1,000,000, so money says accept. In utility it is worth 7.00 against 8.00, so the principle says decline. Nothing here is a criticism of averaging - the average is computed the same way both times, and what differs is the scale being averaged. For accepting to be right the top outcome would have to be rated 11.00, which is to say the second million and a half would have to add at least as much as the first million did. At the lesson’s rating it adds 0.33 times as much, which is what diminishing returns MEANS, and why declining is not merely understandable but prescribed.

The same gamble, the opposite answer

The argument is about the shape of your curve, not about caution in general. A billionaire's utility is essentially linear over a few million, so measuring utility in millions gives

EU(Accept)=12(0)+12(2.5)=1.25>1.0=EU(Decline),EU(\text{Accept}) = \tfrac{1}{2}(0) + \tfrac{1}{2}(2.5) = 1.25 > 1.0 = EU(\text{Decline}) ,

and the same principle says take the gamble. Two agents, identical odds, opposite rational choices. Nothing is inconsistent: they have different utility functions, and the theory never claimed there was one.

Why this is not the end of AI

Russell and Norvig observe that MEU "could be seen as defining all of AI", then immediately point out that it does not solve it. Both are right, and the gap is worth being precise about.

MEU says exactly what a rational choice is. What it does not say is how to get the ingredients.

  • The probabilities P(Result(a)∣a,e)P(\mathrm{Result}(a) \mid a, \mathbf{e}) demand a causal model of the world and inference within it - which the earlier path showed is NP-hard in general Bayesian networks.
  • The utilities U(s′)U(s') often cannot be read off a state at all. Knowing how good a position is may require searching or planning to see what can be reached from it.

So the principle formalises "do the right thing" and leaves nearly all of the work. That is still valuable: it gives a criterion against which approximations can be judged, which is more than a pile of heuristics offers.

MEU and performance measures. If an agent maximises a utility function that correctly reflects the performance measure by which its behaviour is judged, it achieves the highest possible score for that measure, averaged over the environments it might face. The catch is in correctly reflects.

Before the quiz

Be able to write EU(a∣e)EU(a \mid \mathbf{e}) and state MEU, work the game-show example in both money and utility, explain why a billionaire rationally decides the other way, and name the two ingredients MEU assumes but does not supply.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.