The Baseline That Is Hard to Beat
A rating matrix that is 82% empty, the two offsets that explain most of what is in it, and the latent factors that earn the remaining third of the gain.
Everyone starts a recommender by reaching for the interesting model. Start instead with the boring one, because on a simulated catalogue of 800 users and 300 items it gets you two thirds of the way.
The table
A rating matrix has users down one side, items across the other, and almost nothing in the middle. In the simulation used throughout this path, 17.7% of the cells are filled, and the least-rated half of the catalogue holds only 23% of the ratings.
That sparsity is not just an inconvenience. A missing entry is not missing at random. It is missing because that person never encountered the item, and what people encounter depends on what was popular, what was promoted, and what some earlier recommender chose to show them.
So every estimate built from the filled cells describes a population selected by the very process you are trying to improve. That is the same shape of problem as confounding, and it will come back in the third lesson with a number attached.
Three models, in order of ambition
Predict the global mean for every cell. RMSE 0.9368. It is a real baseline and it is worth computing, because it tells you the scale of everything that follows.
Add two offsets. One per user, capturing who rates generously; one per item, capturing what most people like:
RMSE 0.6761.
Add latent factors. Give each user a short vector and each item a vector , so their dot product can express that this kind of person likes that kind of item:
With eight factors, RMSE 0.5393.
| model | RMSE | share of the total gain |
|---|---|---|
| global mean | 0.9368 | - |
| + user and item offsets | 0.6761 | 66% |
| + 8 latent factors | 0.5393 | 100% |
The offsets carry two thirds. That is the useful shape of the problem: a large part of any rating is not about the match between a person and an item at all. It is about a rater who marks everything highly and an item most people like.
Why the shrinkage is not a formality
Look at the denominator in that update:
Without the , an item rated three times gets an offset fitted to three numbers and is trusted exactly as much as one fitted to three hundred. In a catalogue where counts vary by orders of magnitude, that is how an obscure item with four enthusiastic ratings arrives at the top of every list.
Adding shrinks each offset towards zero in proportion to how little data supports it. The item with three ratings barely moves from the global mean; the item with three hundred is left almost alone. It is the same idea as ridge regression, and here it is doing most of the work of keeping the model sane.
The asymmetry in that output is the whole point. At the thin item keeps 27% of its apparent quality while the well-observed one keeps 97%.
The figure below turns that denominator into a dial, and shows what it decides: not the offsets but the order. Take the constant to zero and a short film four people rated tops the list, because with no shrinkage the count does not enter the calculation at all. The arithmetic is not wrong - that really is the average of what those four said. It is just not an estimate of what the next person will think, and only the constant knows the difference.
Interactive: the constant that decides the leaderboard
Drag the shrinkage constant and watch the top of the list change hands.
- λ
- 8
- Top of the list
- a well-loved classic
- The 4-rating item keeps
- 33%
- The 300-rating item keeps
- 97%
At λ = 8 the four-rating item keeps only 33% of its apparent quality while the three-hundred-rating one keeps 97%, and the top of the list is a well-loved classic, on 300 ratings. That asymmetry is the whole mechanism: each offset is pulled towards zero in proportion to how little data supports it, so a thin row barely moves from the global mean and a thick one is left almost alone. Push λ further and the catalogue collapses towards the mean - which is the trade the constant controls.
What the factors add
The remaining third of the gain comes from the interaction term. Each user gets a vector, each item gets a vector, and the dot product says whether this direction of taste matches this direction of item.
Nobody names those directions in advance. They are whatever explains the residuals, and they are only identified up to rotation - which means any interpretation you place on an individual axis is a story about one arbitrary basis among many. Occasionally a direction lines up with something recognisable. Do not build a product feature on the assumption that it will.
Two practical consequences:
- more factors always fit the training ratings better, and start memorising the sparse rows. The factor count and the shrinkage constant have to be chosen together, on validation data
- the model has nothing to say about a user or item it has never seen, since their vector was never fitted. That is the subject of the next lesson
What this sets up
The numbers here are averages over a test set. The next lesson breaks the same model's error down by how much history each user has, and finds that it is worst for exactly the people a recommender most needs to convince.
References & further reading
- Charu C. Aggarwal, Recommender Systems: The Textbook, Springer, 2016· Kudos AI reference library
- Kevin P. Murphy, Probabilistic Machine Learning: An Introduction, MIT Press (Adaptive Computation and Machine Learning), 2022source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.