Understanding Expected Utility
Decision theory combines what an agent believes with what it wants. A utility function assigns a single number to a state, expressing its desirability, and because an action in an uncertain world does not determine its outcome, the relevant quantity is an average: the expected utility of an action is the sum over possible outcome states of their probability given the action and the evidence, times their utility. The principle of maximum expected utility says a rational agent picks the action maximising that quantity. It is not a rule about best cases, which would ignore risk, nor about reaching goals, since outcome quality is continuous rather than a binary test.
The gap between money and utility is where the principle earns its keep. Offered a certain one million dollars or a fair coin paying nothing or two and a half million, the gamble has the higher expected monetary value at one and a quarter million, yet most people decline. Rating the three outcomes at 5, 8 and 9 gives an expected utility of 7 for accepting against 8 for declining, so declining is what the principle prescribes. Nothing is wrong with the averaging; what differs is the scale being averaged. A billionaire whose curve is nearly linear over a few million computes 1.25 against 1.0 and should accept the same gamble, which shows the disagreement lives in the utility function rather than in caution as such.
Empirically the utility of money is close to logarithmic, an idea due to Bernoulli and confirmed by Grayson’s study of individual preferences. Such a curve is concave, meaning equal increments of wealth add shrinking amounts of utility, and concavity is precisely risk aversion: the utility of a lottery is below the utility of receiving its expected monetary value as a sure thing. The sure amount an agent values equally to a gamble is the certainty equivalent, and the shortfall against the gamble’s expected value is the risk premium. Utilities have no absolute scale, so one is fixed by pinning a best and worst outcome, after which any prize can be assessed by adjusting the probability in a standard lottery until the agent is indifferent.
A decision network extends a Bayesian network with decision nodes for the choices an agent controls and a utility node scoring the outcome, and is evaluated by setting each candidate action in turn, running ordinary probabilistic inference, and taking the action of highest expected utility. That structure supports asking what an observation would be worth before paying for it. The value of perfect information is the expected utility of deciding after the observation, averaged over what the observation might say, minus the expected utility of deciding now. Its defining property is that it is zero whenever no possible reading would change the chosen action, and largest when candidate actions are close and the observation could reorder them.
How to Calculate
EU(a | e) = Σ P(Result(a) = s′ | a, e) U(s′); action = argmax EU(a | e)
where
- U(s′)
- the utility of outcome state s′, a single number expressing its desirability
- Result(a)
- the random variable whose values are the possible outcomes of taking action a
- e
- the evidence the agent has observed so far
- VPI
- value of perfect information: expected utility of deciding after an observation, minus deciding now
Example of Expected Utility
Grayson’s fitted curve for one subject, U = −263.31 + 22.09 log(n + 150,000), gains 11.28 utility for the first hundred thousand dollars, then 7.43, then 5.55, then 4.43. Shrinking gains is concavity, and concavity is risk aversion.
On that same curve a fair coin between nothing and eight hundred thousand dollars has expected utility 20.35, while the certain four hundred thousand is worth 28.67. The certainty equivalent is about two hundred and twenty-seven thousand, so the risk premium is about one hundred and seventy-three thousand dollars.
An oil company faces n indistinguishable blocks, one holding oil worth C, each priced C/n. A definitive survey of a single block is worth exactly C/n for every n, verified in exact arithmetic for n of 2, 3, 5, 10 and 100: the information is worth as much as a block itself.
Advantages and Disadvantages
Pros
- Gives a single precise criterion for rational choice, against which any approximation can be judged.
- Separates belief from preference cleanly, so each can be estimated and criticised on its own terms.
- Prices information before it is bought, which frequently shows a proposed measurement to be worthless.
Cons
- Specifies what to compute without saying how: the probabilities need a causal model and inference that is NP-hard in general.
- Utilities must be elicited, and stated preferences are often inconsistent with revealed ones.
- Outcome utilities may themselves require search or planning, since a state’s value can depend on what is reachable from it.
Frequently Asked Questions
Why not just maximise expected money?
Because utility is not proportional to money once the stakes are comparable to total wealth. The first million changes a life and the second does not change it nearly as much, so averaging dollars gives the wrong ranking over gambles. Expected monetary value is a special case that happens to be adequate when the curve is locally linear.
Is risk aversion irrational?
No. It is what a concave utility curve looks like, and the curve encodes genuine preferences about wealth. What would be irrational is preferring a lottery to a sure thing of higher utility. Risk seeking is also rational in the region where the curve is convex, which is typically deep in debt.
When is buying information not worth it?
Whenever no possible result would change the action taken. If one option is clearly ahead and the observation cannot close the gap, the value of perfect information is exactly zero, however precise or expensive the measurement. Value comes from the ability to make the action depend on the situation.
The Bottom Line
Expected utility is the bridge from belief to action: average the utility of an action’s outcomes over their probabilities and take the largest. Its consequences are practical rather than abstract - a concave curve explains and measures risk aversion, and the same arithmetic prices an observation, usually revealing that information which cannot change the decision is worth nothing at all.