Understanding Anomaly Detection
Every anomaly problem starts with a base rate, and it is usually small enough to break the habits carried over from balanced classification. With 100 anomalies in 20,000 events, calling everything normal is right 19,900 times: 99.50% accuracy for a detector that has learned nothing. Any metric that averages over the majority class inherits the base rate, so the first decision is which metric survives it.
ROC AUC survives it only in appearance. On a realistic detector it reads 0.9468 while the average precision reads 0.3245, and both are correct: the false-positive rate divides by the 19,900 normal events, so a thousand false alarms shift it by five points, while precision divides by the number of alerts, where a thousand false alarms are the whole story. Precision at a fixed k is the number that describes what somebody has to work through - 52% at the top 50, 36% at the top 100, 13.4% at the top 500.
The detectors themselves divide into families by what they assume "anomalous" means, and each family is blind to something. On data with two clusters and twenty anomalies in the sparse gap between them, Mahalanobis distance found all ten of the far-away anomalies, none of the twenty in the gap, and scored a ROC of 0.335 - worse than chance, because the gap sits at the global mean and those points are the least extreme in the dataset. Isolation forest scored 0.903 and also found none of the gap, because a point at the centre of both coordinate ranges takes as many axis-parallel splits to isolate as an ordinary one.
Local methods exist for exactly that case, and they have their own failure. The local outlier factor compares a point's density to its neighbours', which assumes the neighbours are normal. As the gap anomalies grow from 2 to 40 they become each other's neighbourhood, and LOF's recall falls from 100% to 8%. Plain kNN distance, which ignores whether the neighbours are anomalous, degrades from 100% to 25% - the unusual case where the simpler method is the more robust one.
How to Calculate
precision@k = (true anomalies in the top k) / k; threshold* = argmin_t [ C_FN · missed(t) + C_FP · false(t) ]
where
- base rate
- the share of observations that are anomalies; sets what any metric can mean
- C_FN, C_FP
- the cost of a missed anomaly and of a false alarm
- k
- the review capacity, which fixes the threshold rather than the other way round
- contamination
- the anomaly rate in the data the detector meets, which moves
Example of Anomaly Detection
20,000 events at a 0.5% anomaly rate: never firing gives 99.50% accuracy, ROC AUC 0.9468 against average precision 0.3245, and the top 100 scores contain 36 real anomalies.
Two clusters with a sparse gap: Mahalanobis ROC 0.335, kNN distance 0.978, local outlier factor 0.933, isolation forest 0.903 - but of the twenty gap anomalies, the top 60 of the ranking holds 65% for kNN, 30% for LOF, and none for the other two.
With a miss costing 500 and a false alarm 20, the cost-minimising threshold fires 460 alerts, catches 67 of 100, and costs 24,360 against 50,000 for never firing.
Frequently Asked Questions
Should I use a supervised classifier instead, if I have labels?
If you have enough labelled anomalies, yes: a classifier will usually beat an unsupervised detector on the kinds of anomaly it has seen. The reason unsupervised methods persist is that the interesting anomaly is the one that has not happened yet, and a classifier trained on last year's frauds is a detector for last year's frauds.
Where should the threshold go?
From the two costs, not from a round-number quantile. Once a miss and a false alarm are priced, the threshold is arithmetic. A quantile rule can land close by luck - the top 1% cost 26,980 against an optimum of 24,360 in the worked example - but it will not stay close when the cost ratio or the contamination changes.
Why did precision fall when nothing about the model changed?
Because precision depends on the contamination rate, which moves. The same threshold on the same detector gave 25% precision at twice the anomaly rate and 8% at half, with recall steady near 65% throughout. A threshold fixes recall; it cannot fix a promise about how clean the queue will be.
The Bottom Line
Anomaly detection is a metrics problem wearing an algorithms costume. Decide the base rate, the review capacity and the cost of a miss before choosing a method, because those three fix what any detector can be worth - and check which kind of anomaly your data contains, since every family of detector is blind to a kind that another one finds easily.