Anomaly detection using surprisals
Abstract
I will discuss a probabilistic approach to anomaly detection based on extreme ‘surprisal values’ aka log scores, equal to minus the log density at each observation. The surprisal approach can be used for any collection of data objects, provided a probability density can be defined on the sample space. It can distinguish anomalies from legitimate observations in a heavy tail, and will identify anomalies that are undetected using methods based on distance measures. I will demonstrate the idea in various real data examples including univariate, multivariate and regression contexts, and when exploring more complicated data objects. I will also briefly outline the underlying theory when the density is known, and when it is estimated using a kernel density estimate. In the latter case, an innovative bandwidth selection method is used based on persistent homology.
Given at
- TopTime Conference, Australian National University, Canberra, Australia: 28 October 2025.
- Frontiers of Statistical Inference 2025, MBZUAI, Abu Dhabi, UAE: 26 November 2025.