Estimators, Bias and Standard Error
An estimator as a random variable with a distribution of its own, the split of its error into bias and variance, and the worked case where the textbook unbiased estimator is the worse of the two.
You have a sample and you want to know something about the population it came from. You compute a number. The whole of this path rests on one shift in how you read that number, so it is worth stating plainly before anything else.
The number you computed is not the number you want
Suppose you want the average height of adults in a city, and you measure 40 people. Their average is cm. That figure is a fact about your 40 people. It is not a fact about the city.
Measure a different 40 and you would get something else - , or . The quantity you actually want, the population mean, sits somewhere fixed and unknown. What moves is your estimate of it.
An estimator is the recipe ("take the average of the sample"). An estimate is what the recipe returns on one particular sample. Because the sample is random, the estimate is a random variable, and it has a distribution - called the sampling distribution. Almost everything in inference is a statement about that distribution rather than about your one number.
Two features of it matter.
Bias asks whether the recipe is aimed correctly. If you repeated the study endlessly and averaged all the estimates, would you land on the truth? If yes, the estimator is unbiased.
Variance asks how tightly the estimates cluster. Two recipes can both be aimed correctly while one scatters far more widely than the other.
Aim and spread are separate questions, and neither alone tells you whether an estimate is any good.
Why dividing by n is wrong, precisely
The sample mean is unbiased for the population mean, which is unsurprising. The sample variance is more interesting, because the obvious recipe is wrong in a way that can be quantified exactly.
The obvious recipe is: take each observation's distance from the sample mean, square it, and average over the observations. It produces answers that are systematically too small - and not by a vague amount, but by the exact factor
At that is nine tenths: on average you recover 90% of the true variance, no matter how many studies you run.
The reason is worth pausing on rather than memorising. You measured distances from the sample mean. But the sample mean is, by construction, the single point that makes those squared distances as small as they can possibly be for this sample. The true mean is somewhere else, so distances from it would have been larger. You have measured spread from an artificially convenient centre.
One degree of freedom was spent locating that centre, and dividing by instead of gives it back exactly. Simulating 400,000 samples of size 10 from a population with variance :
| recipe | average over 400,000 samples | predicted |
|---|---|---|
| divide by | ||
| divide by |
Being right on average is not the same as being close
Here is where the story usually stops, and where it becomes interesting.
Unbiasedness is a statement about an average over infinitely many repetitions. You get one sample. What you actually care about is how far your estimate is likely to land from the truth, and the standard measure of that is mean squared error:
The decomposition is exact, and it means an estimator can trade one for the other. Accept a little bias, buy a larger reduction in variance, and end up closer to the truth on the whole.
That is not a hypothetical. For the two variance recipes at with a true variance of and a normal population, the exact mean squared errors are
The biased estimator is closer to the truth. Dividing by the larger number shrinks every estimate slightly toward zero; the variance that shrinkage saves outweighs the bias it introduces.
This is not a curiosity to file away. It is the same trade that justifies ridge regression, shrinkage estimators and essentially all regularisation: a deliberate, quantified bias that lowers total error. Meeting it here, on a formula you already know, makes it much harder to mistake later for a compromise.
All four of those numbers have closed forms, so the figure below computes them rather than simulating: 3.6 and 4 for the two recipes on average, 3.04 and 3.5556 for their mean squared errors. Dragging the sample size settles something this section states at one size only - for normal data the biased recipe is closer at EVERY n, because (2n-1)(n-1) is less than 2n squared for every n, and the advantage shrinks away as the sample grows.
Interactive: the recipe that is wrong and closer anyway
Closed form throughout. The lesson simulates these; they do not need it.
- Divide by n, on average
- 3.6000
- Divide by n - 1
- 4.0000
- Its mean squared error
- 3.0400
- And the other one’s
- 3.5556
Dividing by n recovers 3.6000 of a true variance of 4 on average - exactly 10 - 1 over 10 of it, because the sample mean is the one point that makes those squared distances as small as they can be. Dividing by n - 1 gives the degree of freedom back and lands on 4 exactly. And yet the biased recipe is CLOSER: its mean squared error is 3.0400 against 3.5556, split into a bias of -0.4000 and a spread of 2.8800. Drag n and the ordering never changes, at any sample size, because (2n-1)(n-1) is less than 2n^2 for every n. Both mean-squared-error formulas assume a normal population, which is where the fourth moment comes from.
The standard error, and what it costs
The standard deviation of the sampling distribution has its own name - the standard error - because it is what "how precise is this estimate?" means. For a sample mean it is
The square root governs the economics of every study ever designed. With a population standard deviation of :
- at , the standard error is
- at , it is
Halving your uncertainty cost you four times the data. Going from to would cost 300 observations more. Precision is purchasable, but the price rises steeply, which is why underpowered studies cannot be rescued by adding a few more subjects - a point the third lesson makes quantitative.
What this buys you
Three ideas, and the rest of the path is built on them.
An estimate is a draw from a distribution, so the honest question is never "what is the answer?" but "how does this recipe behave across the samples I might have drawn?"
Error splits into aim and spread, and the split is exact. Improving an estimator usually means trading between them rather than eliminating either.
Precision improves with , so it is bought at a rising price and must be budgeted before data collection rather than hoped for afterwards.
The next lesson attaches a range to an estimate, and finds that the standard recipe for doing so promises more than it delivers.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.