Skip to content
Kudos AI

Decision Tree

A model that predicts by applying a sequence of threshold tests on individual features, splitting the data into increasingly homogeneous groups.

Also known as: CART, Classification and regression tree

Understanding Decision Tree

A decision tree partitions the feature space by asking one question at a time. Each internal node tests a single feature against a threshold, each branch follows one answer, and each leaf issues a prediction, the majority class for classification or the mean response for regression. Prediction is simply a walk from root to leaf.

Fitting is greedy. At each node the algorithm considers candidate splits and takes the one that most improves the purity of the resulting groups, measured by information gain, Gini impurity, or reduction in variance. It then recurses on each child. The greediness matters: the split that looks best locally need not belong to the best overall tree, and finding the globally optimal tree is computationally intractable.

Left unchecked the recursion continues until every leaf is pure, producing a tree that memorizes the training data and generalizes poorly. Control comes from limiting depth, requiring a minimum number of observations per leaf, or growing the tree fully and then pruning back the branches that fail to justify themselves on held-out data.

The deeper limitation is instability. Because a split near the root determines everything beneath it, a small change in the training sample can select a different root split and produce an entirely different tree. This high variance is the specific weakness that ensemble methods were designed to remove, by averaging over many trees fitted to perturbed versions of the data.

Example of Decision Tree

A tree predicting loan default might first test whether income is below 30,000. On the low-income branch it might then test debt-to-income ratio above 0.4, reaching a leaf predicting default. On the high-income branch it might test credit history length instead.

The path itself is the explanation: "income below 30,000 and debt-to-income above 0.4" is a rule that can be read directly off the tree and stated to someone with no statistical training. Very few model families offer this.

Note that the tree used different second questions on each branch. This is genuine interaction: the relevance of debt-to-income depends on income level. Trees capture such interactions automatically, whereas a linear model would need them specified in advance as explicit product terms.

Advantages and Disadvantages

Pros

  • Directly interpretable; the decision path is a human-readable rule.
  • Handles numeric and categorical features together, with no scaling required.
  • Captures interactions and non-linear structure without manual feature engineering.

Cons

  • High variance: small data changes can produce a completely different tree.
  • Overfits readily unless depth is constrained or the tree is pruned.
  • A single tree is usually less accurate than an ensemble of trees.

Frequently Asked Questions

What is the difference between Gini impurity and entropy as split criteria?

Both measure how mixed the classes are at a node and both are minimized by pure nodes. Entropy uses a logarithm and Gini a sum of squared proportions, so Gini is slightly cheaper to compute. In practice the two rarely produce meaningfully different trees.

Why do decision trees not need feature scaling?

Splits are threshold comparisons on one feature at a time, and any monotonic rescaling leaves the ordering, and therefore the available splits, unchanged. This is in sharp contrast to distance-based or penalized methods, where scale directly affects the result.

What is pruning?

Growing a deliberately large tree and then removing branches that do not earn their complexity, judged on held-out data. Growing large first and cutting back generally outperforms stopping early, because a weak split can still lead to a strong one beneath it.

The Bottom Line

A decision tree is a greedy sequence of feature tests that is unusually easy to read and unusually unstable. That instability is precisely what random forests and boosting exploit, averaging over many trees to keep the flexibility while discarding the variance.