The Parameter Nobody Chooses
The living reward in a grid world is written down once and never discussed, and the optimal policy is a step function of it: eight thresholds between -3 and 0, each flipping exactly one square. The textbook value of -0.04 sits 0.0048 away from the one that decides whether the agent takes the shortcut past the pit, and above -0.0221, when steps are nearly free, the optimal move in one corner is to walk into a wall on purpose.
Prerequisites: Markov Decision Processes
The standard grid world is four squares by three, with a wall at (2,2), a goal worth +1 at (4,3) and a pit worth -1 at (4,2). Every action moves the agent as intended with probability 0.8 and sideways with probability 0.1 each; bumping into a wall leaves it where it is. The discount is 1.
And somewhere in the specification, usually in a sentence that nobody stops on, there is a living reward: every non-terminal square pays per step, typically . It is there to make dawdling costly. It is not presented as a modelling choice, because it does not feel like one.
It is the most consequential number in the problem.
A. Eight thresholds
Solve the world for every between and and the optimal policy does not drift. It sits still, then changes in one square, then sits still again. Bisecting each change to six decimal places:
| Living reward | Square | Was | Becomes |
|---|---|---|---|
| -1.649707 | (3,2) | right | up |
| -1.564259 | (3,1) | right | up |
| -0.731138 | (1,1) | right | up |
| -0.452624 | (4,1) | up | left |
| -0.084989 | (2,1) | right | left |
| -0.044833 | (3,1) | up | left |
| -0.027357 | (3,2) | up | left |
| -0.022145 | (4,1) | left | down |
Nine distinct optimal policies over that range, separated by eight numbers, and every figure above was computed twice: by policy iteration, where each evaluation is an exact linear solve, and by value iteration run to a tolerance of . The two methods agree on every policy and on every utility to within .
Interactive: the policy as a step function of the living reward
The 4x3 world, 0.8 intended, 0.1 each side, discount 1.
- Utility of square (1,1)
- 0.705308
- Policy interval
- 7 of 9
- Squares unlike the textbook
- 0
- Value iteration vs exact
- 5.4e-15
At R = -0.04 the optimal policy is the one that holds for every living reward between -0.044833 and -0.027357, and U(1,1) is 0.705308. At (3,1) the agent goes left, the long way round, as far from the pit as it can keep.
B. The textbook value is 0.0048 from a cliff
Look at the bolded row. At the agent standing at (3,1), diagonally below and to the left of the pit, goes left: the long way round, along the bottom and up the far side, without ever stepping onto a square next to the pit. That is the policy printed in every textbook, and the one everybody's intuition learns the world from.
At it goes up instead, onto (3,2), the square directly left of the pit, taking the short route and accepting a real chance of being blown sideways into the .
The switch is at . The canonical setting sits above it. A living reward of and a living reward of are the same choice as far as anyone's judgement goes, and they produce two different stories about what a rational agent does in this world.
The utilities move smoothly through all of this: is at , at , and at . It is the policy, the thing you actually deploy, that is a step function.
C. Two policies that look like bugs
Above -0.022145 the agent walks into a wall on purpose. For every living reward between and , the cheapest steps in the whole range, the optimal action at (4,1), the bottom-right corner, is down. There is nothing below; the move bumps the floor and the agent stays where it is with probability 0.8, sliding sideways with probability 0.1 each way, one of which also bumps a wall.
It is not a bug. The square directly above (4,1) is the pit. When time costs almost nothing, the cheapest thing to do in that corner is nothing at all, and the action that best achieves nothing is to push against the floor. Any purposeful move risks drifting up into the .
At the other end, the agent dives into the pit. For the optimal action at (3,2) is to step right, directly into the terminal. Once each additional step costs more than 1.649707, the fastest exit wins and the pit is the nearest exit. at : the agent is not confused, it is in a world where existing is expensive.
Both behaviours are correct optimisation. Both would be reported as bugs by anyone who had not looked at the living reward, and neither can be fixed by changing the algorithm.
D. What to do about it
The practical points are small and specific.
- Sweep the parameter you did not think you were choosing. Nine policies over one interval is not an exotic finding; it is what a piecewise-constant argmax looks like. The sweep is cheap: each solve here is an linear system.
- Report the interval, not the point. "The optimal policy for " is a claim that survives someone rounding the living reward. "The optimal policy for " is a claim about one number, and it is the weaker statement even though it looks more precise.
- A tuned reward is a fitted parameter. Adjusting a step cost until the agent does the sensible thing is fitting, and the fitted value inherits every caveat a fitted value has. In particular it should not then be reported as part of the environment.
- Check both edges of the range you believe. If your belief is "time is mildly costly, somewhere around a few percent per step", solve at both ends of "a few percent". Here that interval contains three thresholds.
The deeper version of the point is that the reward function is not a description of the world. It is the entire specification of what the agent is for, and a piece of it that gets written once in a configuration file, never reviewed, and never varied is doing more work than the parts everyone argues about.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.