Skip to content
Kudos AI
Lire en français
Sequential Decisions and Reinforcement Learning

Decisions under Uncertainty: Utility and Information

Why a bet with positive expected monetary value can be rational to refuse, what the curvature of a utility function measures, and how to price an observation before buying it - including the common case where the honest price is zero.

7 min readKudos AI

Prerequisites: Bayesian Networks and Probabilistic Inference

A game-show choice weighed twice, once in dollars and once in utility, with the two calculations reaching opposite answers on the same coin flip.

Inference produces beliefs. Beliefs do not tell you what to do. Acting needs something beliefs cannot supply - preferences - and decision theory is the arithmetic that puts the two together.

A. Expected utility

A utility function U(s)U(s) gives one number for how desirable a state is. Since an action in an uncertain world does not fix its outcome, write Result(a)\mathrm{Result}(a) for the random variable of possible outcomes. The expected utility of an action, given evidence e\mathbf{e}, is

EU(a∣e)=∑s′P(Result(a)=s′∣a,e) U(s′),EU(a \mid \mathbf{e}) = \sum_{s'} P\big(\mathrm{Result}(a) = s' \mid a, \mathbf{e}\big)\, U(s') ,

and the maximum expected utility principle is simply argmax⁡aEU(a∣e)\operatorname{argmax}_a EU(a \mid \mathbf{e}).

That is not "maximise the best case", which ignores risk, and not "reach a goal", since outcome quality here is continuous rather than pass-or-fail.

B. Money is not utility

You have won a game show. Take $1,000,000, or flip a fair coin for nothing against $2,500,000. The gamble's expected monetary value is $1,250,000, which beats the sure million. Most people decline.

Rate your current position at 5, the sure million at 8, and the top outcome at 9. Then

EU(Accept)=12(5)+12(9)=7<8=EU(Decline),EU(\text{Accept}) = \tfrac{1}{2}(5) + \tfrac{1}{2}(9) = 7 < 8 = EU(\text{Decline}) ,

so declining is what the principle prescribes. The averaging is identical in both calculations; only the scale differs, and utility is the scale that encodes preference.

The disagreement is about the shape of a curve, not about caution. A billionaire's utility is nearly linear over a few million, giving EU(Accept)=1.25>1.0EU(\text{Accept}) = 1.25 > 1.0, so the same principle says take the gamble. Two agents, identical odds, opposite rational choices.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

The argument is about the shape of a curve, so here is the curve. The three outcomes are points on it, the slider bends it, and the odds never move. The dashed line is the rating the top outcome would need for accepting to be right: 11, at the ratings above. That is worth reading twice - it says the second million and a half would have to add at least as much as the first million did, which is the exact opposite of the diminishing returns the whole argument rests on.

Interactive: the coin, and the curve that decides it

The odds never change. Only the shape of the utility does.

01.0M2.5M
Expected utility, accept
7.00
Expected utility, decline
8.00
Top rating that would tie
11.00
What MEU prescribes
decline
Expected money
$1,250,000

The gamble is worth $1,250,000 in money against a sure $1,000,000, so money says accept. In utility it is worth 7.00 against 8.00, so the principle says decline. Nothing here is a criticism of averaging - the average is computed the same way both times, and what differs is the scale being averaged. For accepting to be right the top outcome would have to be rated 11.00, which is to say the second million and a half would have to add at least as much as the first million did. At the lesson’s rating it adds 0.33 times as much, which is what diminishing returns MEANS, and why declining is not merely understandable but prescribed.

C. Curvature is risk aversion

Grayson found the utility of money close to logarithmic, an idea going back to Bernoulli in 1738. For one subject the fit was U=−263.31+22.09log⁡(n+150,000)U = -263.31 + 22.09\log(n + 150{,}000). Step through it in equal increments of $100,000 and the utility added shrinks: 11.2811.28, then 7.437.43, then 5.555.55, then 4.434.43.

Shrinking gains is concavity, and for a concave curve U(L)<U(SEMV(L))U(L) < U(S_{EMV(L)}) for every lottery LL: facing the gamble is worth less than being handed its expected value. That is risk aversion, and it is measurable. On the same curve, a fair coin between $0 and $800,000 has expected utility 20.3520.35 against 28.6728.67 for the certain $400,000. The sure amount he values equally to the gamble - the certainty equivalent - is about $227,000, so the risk premium is about $173,000.

The curve is not concave everywhere. Someone already $10 million in debt might rationally accept a coin for a further $10 million gain against a $20 million loss, since the outcomes barely differ from where they stand. That gives an S-shape: risk-seeking when desperate, risk-averse over positive wealth.

Utilities have no absolute scale, so fix one by pinning a best and worst outcome at 1 and 0, then elicit any prize by offering it against a standard lottery and adjusting the probability until the agent is indifferent. That probability is the utility.

Where the worst outcome is death, people resist pricing it, and the trade-off happens regardless. A micromort is a one-in-a-million chance of death. Driving 230 miles costs one, so a car's 92,000-mile life costs 400; people pay about $10,000 for a car halving that risk, saving 200, which implies $50 per micromort - matching what studies report directly. The scope is small risks only; nobody accepts $50 million to die outright.

Russell and Norvig describe an agency that rejected an asbestos study for assuming a dollar value for a child's life, then declined the removal - thereby implying a lower value than the one it refused to state.

D. Pricing an observation

A decision network extends a Bayesian network with rectangular decision nodes for the agent's choices and a diamond utility node scoring the outcome. Evaluating it is a loop: set the evidence, then for each action run ordinary inference over the utility node's parents, average, and take the best. The inference engine is unchanged.

That structure lets you ask what an observation is worth before buying it. The value of perfect information is the expected utility of deciding after the observation, averaged over what it might say, minus the expected utility of deciding now:

VPIe(Ej)=(∑kP(Ej=ejk∣e) EU(αejk∣e,Ej=ejk))−EU(α∣e).VPI_{\mathbf{e}}(E_j) = \left(\sum_k P(E_j = e_{jk} \mid \mathbf{e})\, EU\big(\alpha_{e_{jk}} \mid \mathbf{e}, E_j = e_{jk}\big)\right) - EU(\alpha \mid \mathbf{e}) .

An oil company can buy one of nn indistinguishable blocks; exactly one holds oil worth CC and each costs C/nC/n, so expected profit is zero either way. A definitive survey of block 3 changes that. With probability 1/n1/n it finds oil and the company profits (n−1)C/n(n-1)C/n; otherwise the odds among the rest improve from 1/n1/n to 1/(n−1)1/(n-1), worth C/(n(n−1))C/(n(n-1)). Together:

1n⋅(n−1)Cn+n−1n⋅Cn(n−1)=Cn.\frac{1}{n}\cdot\frac{(n-1)C}{n} + \frac{n-1}{n}\cdot\frac{C}{n(n-1)} = \frac{C}{n} .

The survey of one block is worth exactly the price of a block, for every nn.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

E. The case where the price is zero

The same reasoning has a sharper consequence. Two routes: a highway worth expected utility 10, and a dirt track worth 4 that a satellite report could move anywhere in [3,6][3, 6]. The report is worthless. Whatever it says the track tops out at 6, you still take the highway, and expected utility afterwards is still 10, so VPI=0VPI = 0.

Make the routes close instead - 10 against something in [8,13][8, 13] - and over readings 8,10,12,138, 10, 12, 13 you take the better of 10 and the reading each time:

10+10+12+134=11.25.\frac{10 + 10 + 12 + 13}{4} = 11.25 .

Then mind the baseline. Deciding now is not taking the highway: it is taking whichever action has the higher expected utility, and the track is already worth 14(8+10+12+13)=10.75\tfrac{1}{4}(8 + 10 + 12 + 13) = 10.75. So VPI=11.25−10.75=0.50VPI = 11.25 - 10.75 = 0.50, and all of it comes from a single reading. Without the report you drive the track; a reading of 88 sends you to the highway instead, which is worth exactly 14(10−8)=0.50\tfrac{1}{4}(10 - 8) = 0.50. At 1010 the two routes tie, so switching gains nothing, and readings of 1212 and 1313 are good news and worth nothing, because you were taking that route anyway.

Where this leaves you

Information is valuable exactly when it can change what you do. A measurement that cannot alter the choice is worth nothing however precise or expensive, and value peaks when the candidate actions are close and the observation is wide enough to reorder them. That is a cheap thing to check before commissioning a study, and it is the same arithmetic that explains why a sensible person turns down a favourable bet. The training path Making Decisions under Uncertainty works each of these by hand and in code.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

5 min readProbabilistic Reasoning

The Week That Cannot Have Happened

Take the most likely state on each day and write them down in order, and you have a report the model assigns probability exactly zero: on a four-day machine-monitoring example the day-by-day answer is healthy, healthy, failed, failed, and healthy to failed is a transition that cannot occur. What the two questions actually are, why smoothing and Viterbi answer different ones, and what the 0.411 posterior on the best path means for anyone who has to act on it.

Artificial IntelligenceProbability
3 min readProbabilistic Reasoning

A Hundred Thousand Samples, Four Hundred of Them Real

On the burglary network with both neighbours calling, rejection sampling keeps 183 of 100,000 draws and likelihood weighting keeps all of them at an effective sample size of 396. Both estimates are about 10% off a posterior of 0.284172, and the reason is exactly computable: 252 samples carry 76% of the weight and 99.975% of the squared weight.

Artificial IntelligenceProbability
10 min readProbabilistic Reasoning

Learning the Numbers in a Probability Model

Where the numbers in a Bayesian network or a Gaussian actually come from: the three-step maximum-likelihood recipe worked through on discrete and continuous parameters, the Beta prior that repairs what it does to an unseen event, naive Bayes and the single zero count that destroys it, and the EM algorithm for the case where the counts cannot be taken at all - with every figure computed rather than asserted.

ProbabilityStatisticsArtificial Intelligence
← Back to all articles