Skip to content
Kudos AI
Lire en français
Reinforcement Learning

The Parameter Nobody Chooses

The living reward in a grid world is written down once and never discussed, and the optimal policy is a step function of it: eight thresholds between -3 and 0, each flipping exactly one square. The textbook value of -0.04 sits 0.0048 away from the one that decides whether the agent takes the shortcut past the pit, and above -0.0221, when steps are nearly free, the optimal move in one corner is to walk into a wall on purpose.

5 min readKudos AI

Prerequisites: Markov Decision Processes

Bellman sweeps computed live on a small world, the utilities settling square by square, and the policy arrows turning over as the value of staying put changes.

The standard grid world is four squares by three, with a wall at (2,2), a goal worth +1 at (4,3) and a pit worth -1 at (4,2). Every action moves the agent as intended with probability 0.8 and sideways with probability 0.1 each; bumping into a wall leaves it where it is. The discount is 1.

And somewhere in the specification, usually in a sentence that nobody stops on, there is a living reward: every non-terminal square pays RR per step, typically −0.04-0.04. It is there to make dawdling costly. It is not presented as a modelling choice, because it does not feel like one.

It is the most consequential number in the problem.

A. Eight thresholds

Solve the world for every RR between −3-3 and 00 and the optimal policy does not drift. It sits still, then changes in one square, then sits still again. Bisecting each change to six decimal places:

Living rewardSquareWasBecomes
-1.649707(3,2)rightup
-1.564259(3,1)rightup
-0.731138(1,1)rightup
-0.452624(4,1)upleft
-0.084989(2,1)rightleft
-0.044833(3,1)upleft
-0.027357(3,2)upleft
-0.022145(4,1)leftdown

Nine distinct optimal policies over that range, separated by eight numbers, and every figure above was computed twice: by policy iteration, where each evaluation is an exact linear solve, and by value iteration run to a tolerance of 10−1410^{-14}. The two methods agree on every policy and on every utility to within 9×10−149 \times 10^{-14}.

Interactive: the policy as a step function of the living reward

The 4x3 world, 0.8 intended, 0.1 each side, discount 1.

(1,1)0.705(1,2)0.762(1,3)0.812(2,1)0.655(2,3)0.868(3,1)0.611(3,2)0.660(3,3)0.918(4,1)0.388(4,2)-1(4,3)+1-30-0.04-1-0.2
Utility of square (1,1)
0.705308
Policy interval
7 of 9
Squares unlike the textbook
0
Value iteration vs exact
5.4e-15

At R = -0.04 the optimal policy is the one that holds for every living reward between -0.044833 and -0.027357, and U(1,1) is 0.705308. At (3,1) the agent goes left, the long way round, as far from the pit as it can keep.

B. The textbook value is 0.0048 from a cliff

Look at the bolded row. At R=−0.04R = -0.04 the agent standing at (3,1), diagonally below and to the left of the pit, goes left: the long way round, along the bottom and up the far side, without ever stepping onto a square next to the pit. That is the policy printed in every textbook, and the one everybody's intuition learns the world from.

At R=−0.05R = -0.05 it goes up instead, onto (3,2), the square directly left of the pit, taking the short route and accepting a real chance of being blown sideways into the −1-1.

The switch is at R=−0.044833R = -0.044833. The canonical setting sits 0.0048330.004833 above it. A living reward of −0.04-0.04 and a living reward of −0.05-0.05 are the same choice as far as anyone's judgement goes, and they produce two different stories about what a rational agent does in this world.

The utilities move smoothly through all of this: U(1,1)U(1,1) is 0.7053080.705308 at R=−0.04R = -0.04, −1.600186-1.600186 at R=−0.4R = -0.4, and −10.815340-10.815340 at R=−2R = -2. It is the policy, the thing you actually deploy, that is a step function.

C. Two policies that look like bugs

Above -0.022145 the agent walks into a wall on purpose. For every living reward between −0.022145-0.022145 and 00, the cheapest steps in the whole range, the optimal action at (4,1), the bottom-right corner, is down. There is nothing below; the move bumps the floor and the agent stays where it is with probability 0.8, sliding sideways with probability 0.1 each way, one of which also bumps a wall.

It is not a bug. The square directly above (4,1) is the pit. When time costs almost nothing, the cheapest thing to do in that corner is nothing at all, and the action that best achieves nothing is to push against the floor. Any purposeful move risks drifting up into the −1-1.

At the other end, the agent dives into the pit. For R<−1.649707R < -1.649707 the optimal action at (3,2) is to step right, directly into the −1-1 terminal. Once each additional step costs more than 1.649707, the fastest exit wins and the pit is the nearest exit. U(1,1)=−10.815340U(1,1) = -10.815340 at R=−2R = -2: the agent is not confused, it is in a world where existing is expensive.

Both behaviours are correct optimisation. Both would be reported as bugs by anyone who had not looked at the living reward, and neither can be fixed by changing the algorithm.

D. What to do about it

The practical points are small and specific.

  • Sweep the parameter you did not think you were choosing. Nine policies over one interval is not an exotic finding; it is what a piecewise-constant argmax looks like. The sweep is cheap: each solve here is an 11×1111 \times 11 linear system.
  • Report the interval, not the point. "The optimal policy for −0.0448<R<−0.0274-0.0448 < R < -0.0274" is a claim that survives someone rounding the living reward. "The optimal policy for R=−0.04R = -0.04" is a claim about one number, and it is the weaker statement even though it looks more precise.
  • A tuned reward is a fitted parameter. Adjusting a step cost until the agent does the sensible thing is fitting, and the fitted value inherits every caveat a fitted value has. In particular it should not then be reported as part of the environment.
  • Check both edges of the range you believe. If your belief is "time is mildly costly, somewhere around a few percent per step", solve at both ends of "a few percent". Here that interval contains three thresholds.

The deeper version of the point is that the reward function is not a description of the world. It is the entire specification of what the agent is for, and a piece of it that gets written once in a configuration file, never reviewed, and never varied is doing more work than the parts everyone argues about.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

9 min readReinforcement Learning

Markov Decision Processes

How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.

Reinforcement LearningProbabilityArtificial Intelligence
5 min readProbabilistic Reasoning

The Week That Cannot Have Happened

Take the most likely state on each day and write them down in order, and you have a report the model assigns probability exactly zero: on a four-day machine-monitoring example the day-by-day answer is healthy, healthy, failed, failed, and healthy to failed is a transition that cannot occur. What the two questions actually are, why smoothing and Viterbi answer different ones, and what the 0.411 posterior on the best path means for anyone who has to act on it.

Artificial IntelligenceProbability
5 min readSearch and Games

The Fix That Changed the Success Rate Far More Than the Cost

Allowing sideways moves takes hill climbing on 8 queens from 14.75% of runs solved to 94.55%, which reads like a six-fold improvement and is not one: with random restarts the expected cost of a solution goes from 21.9 steps to 23.1, and counted in moves evaluated it falls by 16%, from 1,547 to 1,298. Simulated annealing solves 98.8% and costs 1,622 evaluations. What changed was mostly the statistic, not the work.

Artificial Intelligence
← Back to all articles