Understanding Bagging and Random Forests
Averaging independent estimates reduces variance: the mean of many noisy predictions is more stable than any one of them. Bagging, short for bootstrap aggregating, exploits this without needing multiple datasets. It draws repeated bootstrap resamples, each formed by sampling the training data with replacement, fits a model to each, and averages the predictions, or takes a majority vote for classification.
The gain depends entirely on the base learner being unstable. Averaging many nearly identical models achieves nothing. Fully grown decision trees are ideal: they have low bias and high variance, exactly the profile that averaging improves, since the procedure suppresses variance while leaving bias essentially untouched.
Random forests address bagging’s remaining weakness. Bootstrap resamples of the same dataset are similar, so if one predictor is strongly dominant, nearly every tree splits on it first and the trees end up highly correlated. Correlated errors do not average away. The fix is to consider only a random subset of the features as candidates at each split, which forces different trees to rely on different predictors and lowers the correlation between them.
The ensemble also supplies free validation. Each bootstrap resample omits roughly a third of the observations, and each observation can be predicted using only the trees that did not see it. Aggregating these out-of-bag predictions gives an estimate of test error without a separate held-out set. The cost of the whole approach is interpretability: a single tree can be read, hundreds cannot, so importance measures and partial dependence plots are used instead.
Example of Bagging and Random Forests
A single deep tree fitted to a dataset might achieve 100% training accuracy and 72% test accuracy, the classic overfitting signature. Bagging 500 such trees typically lifts test accuracy substantially while leaving training accuracy high, because the averaging cancels the idiosyncratic errors of individual trees.
Suppose one feature is far more predictive than the rest. In plain bagging almost every tree splits on it at the root, so the trees are near-copies and the ensemble barely beats a single tree. A random forest restricting each split to a random subset of features forces most trees to find alternative structure.
The size of that subset is the main tuning knob. Smaller subsets give more decorrelation but weaker individual trees; larger subsets approach plain bagging. Common defaults are the square root of the number of features for classification and about a third for regression, tuned by cross-validation or the out-of-bag estimate.
Frequently Asked Questions
Does adding more trees cause overfitting?
No. Test error decreases and then plateaus as trees are added; it does not rise. The number of trees is a compute budget rather than a complexity parameter. Overfitting in a forest is controlled through tree depth and the feature-subset size instead.
What is the out-of-bag error estimate?
Each bootstrap resample leaves out roughly a third of observations. Predicting each observation using only the trees that did not train on it yields a validation estimate at no extra computational cost, which often makes a separate cross-validation unnecessary.
How do random forests differ from boosting?
Random forests fit independent trees in parallel and average them, targeting variance. Boosting fits trees sequentially, each correcting the previous ensemble’s errors, targeting bias. Boosting can reach higher accuracy but is more sensitive to noise and to hyperparameters.
The Bottom Line
Bagging turns the instability of deep trees into an advantage by averaging over bootstrap resamples, and random forests strengthen the effect by decorrelating the trees through random feature subsets. The result is a strong, low-maintenance default at the cost of a single tree’s interpretability.