Anomaly detection using surprisals

Date

28 October 2025

Venue

Various

 

Abstract

I will discuss a probabilistic approach to anomaly detection based on extreme ‘surprisal values’ aka log scores, equal to minus the log density at each observation. The surprisal approach can be used for any collection of data objects, provided a probability density can be defined on the sample space. It can distinguish anomalies from legitimate observations in a heavy tail, and will identify anomalies that are undetected using methods based on distance measures. I will demonstrate the idea in various real data examples including univariate, multivariate and regression contexts, and when exploring more complicated data objects. I will also briefly outline the underlying theory when the density is known, and when it is estimated using a kernel density estimate. In the latter case, an innovative bandwidth selection method is used based on persistent homology.

Given at

Slides

Software

References

Hyndman, Rob J. 2026. That’s Weird: Anomaly Detection Using R. https://OTexts.com/weird.
Hyndman, Rob J, and David T Frazier. 2026. “Anomaly Detection Using Surprisals.” http://robjhyndman.com/publications/surprisals.html.
Hyndman, Rob J, Sevvandi Kandanaarachchi, and Katharine Turner. 2026. “When Lookout Sees Crackle: Anomaly Detection via Kernel Density Estimation.” http://robjhyndman.com/publications/lookout2.html.
Kandanaarachchi, Sevvandi, and Rob J Hyndman. 2022. “Leave-One-Out Kernel Density Estimates for Outlier Detection.” J Computational & Graphical Statistics 31: 586–99. https://robjhyndman.com/publications/lookout/.