Skip to content
Kudos AI

Naive Bayes

A classifier that applies Bayes’ theorem while assuming all features are conditionally independent given the class.

Also known as: Naive Bayes classifier, Bayesian classifier

Understanding Naive Bayes

To classify by probability you would like P(class | features), which by Bayes’ theorem is proportional to P(features | class) · P(class). The difficulty is the first factor: modelling the joint distribution over all features simultaneously requires an amount of data that grows exponentially with the number of features.

Naive Bayes cuts through this by assuming the features are conditionally independent given the class. Under that assumption the joint likelihood factorizes into a product of one-dimensional terms, each of which can be estimated from simple counts or a simple density fit. A model with a thousand features needs only a thousand small estimates rather than one intractable joint one. Russell and Norvig describe exactly this structure, a single cause influencing many effects that are conditionally independent given the cause, and note that the model is called "naive" precisely because it is routinely applied where that independence does not really hold.

The assumption is nearly always violated: in text classification, the presence of one word plainly changes the odds of related words appearing. What rescues the method is that classification depends only on which class scores highest, not on the scores being correct. The independence assumption distorts the magnitudes, often severely, while frequently leaving the ordering intact.

The practical consequence is a split verdict. Naive Bayes is a strong, extremely fast baseline for high-dimensional problems such as text, and it remains usable when data is scarce. But its output probabilities should not be read as calibrated confidences: they are habitually pushed toward 0 and 1, because multiplying many correlated terms as though they were independent overstates the accumulated evidence.

How to Calculate

ŷ = argmax_c P(c) · Πⱼ P(xⱼ | c)

where

c
a candidate class
P(c)
the prior probability of that class
P(xⱼ | c)
the likelihood of feature j taking its observed value, given the class
Πⱼ
product over features, valid only under the conditional-independence assumption

Example of Naive Bayes

For spam filtering, each feature is the presence of a particular word. Fitting the model is a counting exercise: for each class, record how often each word appears in messages of that class, and how common the class itself is.

To classify a new message, multiply the class prior by the per-word likelihoods for the words it contains, once for spam and once for not-spam, and choose the larger. In practice the computation is done by summing logarithms rather than multiplying probabilities, since a product of thousands of small numbers underflows to zero in floating point.

A word never seen in the training data for a class would give that class a likelihood of exactly zero, annihilating the entire product regardless of all other evidence. The standard fix is additive (Laplace) smoothing: add a small constant to every count so no probability is ever exactly zero.

Advantages and Disadvantages

Pros

  • Extremely fast to train and to apply, requiring only counts.
  • Works well with very many features and comparatively little data.
  • Simple, transparent, and a genuinely strong baseline for text classification.

Cons

  • The conditional-independence assumption is almost always violated.
  • Predicted probabilities are poorly calibrated, typically far too confident.
  • Correlated features are effectively counted multiple times, compounding their influence.

Frequently Asked Questions

Why does it work when its central assumption is false?

Because the decision depends only on which class has the highest score. Violating independence distorts the scores, but frequently preserves their order, so the predicted class stays correct even when the predicted probability is badly wrong.

What is Laplace smoothing and why is it needed?

It adds a small constant to every count so that no estimated probability is exactly zero. Without it, a single feature value unseen during training for a class forces the entire product to zero, letting one absent word veto all other evidence.

Can naive Bayes handle continuous features?

Yes. The usual approach, Gaussian naive Bayes, fits a normal distribution per feature per class and uses its density in place of a counted probability. Discretizing continuous features into bins is another common option.

The Bottom Line

Naive Bayes buys enormous tractability with an assumption it knows to be wrong, and gets away with it because classification needs only the right ranking. Use it as a fast, robust baseline; do not trust its probabilities as calibrated estimates.