Paul Boniol

dblp:266/6112 · DBLP profile ↗
in reviewer pool ← Back
29ranked-venue papers in the field
16as first author
25since 2021 · last 2026
0000-0001-8516-0123ORCID · corroborated

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 26 (16 first)Data Mining & Knowledge Discovery · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 From Benchmarks to Production: Transferring Time Series Anomaly Detection Methods for Electricity Production Monitoring
abstract
International audience
Nicolas Vautier, Paul Caron, Nardi Xhepi, Félicie Bizeul, Manel Boumghar, Christophe Degouy, Paul Boniol
ICDE7
2026 A Comprehensive Guide to Time-Series Anomaly Detection
abstract
Anomaly detection is a fundamental data analytics task across scientific fields and industries. In recent years, an increasing interest has been shown in the application of anomaly detection techniques to time series. In this tutorial, we take a holistic view of anomaly detection in time series and comprehensively cover detection algorithms ranging from the 1980s to the most current state-of-the-art techniques. Importantly, the scope of this tutorial extends beyond algorithmic discussion, delving into the latest advancements in benchmarking and evaluation measures for this area. In particular, our interactive systems enable the exploration of methods and benchmarking results, thereby promoting user comprehension. Furthermore, this tutorial extensively explores automated solutions for unsupervised model selection, introduces a new taxonomy, and engages with the challenges and recent findings, particularly the difficulty for these solutions to outperform simple random choice. Driven by the limited generalizability of current detection algorithms, we review recent applications of foundation models for anomaly detection to motivate further research in the area.
John Paparrizos, Paul Boniol, Themis Palpanas
WSDM2
2025 Interpretable Multivariate Anomaly Detector Selection for Automatic Marine Data Quality Control
abstract
International audience
Ngoc-Thanh Nguyen 0002, Astrid Marie Skålvik, Emmanouil Sylligardos, Rogardt Heldal, Patrizio Pelliccione, Paul Boniol, Themis Palpanas, Sverre Jakob Alvsvåg
IEEE Big Data6
2025 Graphint: Graph-Based Time Series Clustering Visualisation Tool
abstract
With the exponential growth of time series data across diverse domains, there is a pressing need for effective analysis tools. Time series clustering is important for identifying patterns in these datasets. However, prevailing methods often encounter obstacles in maintaining data relationships and ensuring interpretability. We present Graphint, an innovative system based on the$k$-Graph methodology that addresses these challenges. Graphint integrates a robust time series clustering algorithm with an interactive tool for comparison and interpretation. More precisely, our system allows users to compare results against competing approaches, identify discriminative subsequences within specified datasets, and visualize the critical information utilized by$k$-Graph to generate outputs. Overall, Graphint offers a comprehensive solution for extracting actionable insights from complex temporal datasets.
Paul Boniol, Donato Tiano, Angela Bonifati, Themis Palpanas
ICDE1
2025 Few Labels are All you Need: A Weakly Supervised Framework for Appliance Localization in Smart-Meter Series
abstract
Improving smart grid system management is crucial in the fight against climate change, and enabling consumers to play an active role in this effort is a significant challenge for electricity suppliers. In this regard, millions of smart meters have been deployed worldwide in the last decade, recording the main electricity power consumed in individual households. This data produces valuable information that can help them reduce their electricity footprint; nevertheless, the collected signal aggregates the consumption of the different appliances running simultaneously in the house, making it difficult to apprehend. Non-Intrusive Load Monitoring (NILM) refers to the challenge of estimating the power consumption, pattern, or on/off state activation of individual appliances using the main smart meter signal. Recent methods proposed to tackle this task are based on a fully supervised deep-learning approach that requires both the aggregate signal and the ground truth of individual appliance power. However, such labels are expensive to collect and extremely scarce in practice, as they require conducting intrusive surveys in households to monitor each appliance. In this paper, we introduce CamAL, a weakly supervised approach for appliance pattern localization that only requires information on the presence of an appliance in a household to be trained. CamAL merges an ensemble of deep-learning classifiers combined with an explainable classification method to be able to localize appliance patterns. Our experimental evaluation, conducted on 4 real-world datasets, demonstrates that CamAL significantly outperforms existing weakly supervised baselines and that current SotA fully supervised NILM approaches require significantly more labels to reach CamAL performances. The source of our experiments is available at: https://github.com/adrienpetralia/CamAL.
Adrien Petralia, Paul Boniol, Philippe Charpentier, Themis Palpanas
ICDE2
2025 DeviceScope: An Interactive App to Detect and Localize Appliance Patterns in Electricity Consumption Time Series
abstract
In recent years, electricity suppliers have installed millions of smart meters worldwide to improve the management of the smart grid system. These meters collect a large amount of electrical consumption data to produce valuable information to help consumers reduce their electricity footprint. However, having non-expert users (e.g., consumers or sales advisors) understand these data and derive usage patterns for different appliances has become a significant challenge for electricity suppliers because these data record the aggregated behavior of all appliances. At the same time, ground-truth labels (which could train appliance detection and localization models) are expensive to collect and extremely scarce in practice. This paper introduces DeviceScope [1], an interactive tool designed to facilitate understanding smart meter data by detecting and localizing individual appliance patterns within a given time period. Our system is based on CamAL (Class Activation Map-based Appliance Localization), a novel weakly supervised approach for appliance localization that only requires the knowledge of the existence of an appliance in a household to be trained.
Adrien Petralia, Paul Boniol, Philippe Charpentier, Themis Palpanas
ICDE2
2025 Advances in Time-Series Anomaly Detection: Algorithms, Benchmarks, and Evaluation Measures
abstract
International audience
John Paparrizos, Paul Boniol, Themis Palpanas
KDD (2)2
2025 Time Series Motif Discovery: A Comprehensive Evaluation
abstract
Motif Discovery involves identifying recurring patterns and locating their occurrences within a time series without prior knowledge about their shape or location. In practice, Motif Discovery faces several data-related challenges, leading to various definitions of the problem and multiple algorithms addressing these challenges to different extents. However, there has been no systematic evaluation and comparison of these diverse approaches. Consequently, this paper presents a comprehensive literature review covering data-related challenges, motif definitions, and algorithms. We also analyze the strengths and limitations of algorithms carefully chosen to represent the literature diversity. The analysis is structured around key research questions identified from our review. Our experimental findings provide practical guidelines for selecting Motif Discovery algorithms suitable for a given task and suggest directions for future research.
Valerio Guerrini, Thibaut Germain, Charles Truong, Laurent Oudre, Paul Boniol
Proc. VLDB Endow.5
2025 -Graph: A Graph Embedding for Interpretable Time Series Clustering
abstract
Time series clustering poses a significant challenge with diverse applications across domains. A prominent drawback of existing solutions lies in their limited interpretability, often confined to presenting users with centroids. In addressing this gap, our work presents$k$-Graph, an unsupervised method explicitly crafted to augment interpretability in time series clustering. Leveraging a graph representation of time series subsequences,$k$-Graph constructs multiple graph representations based on different subsequence lengths. This feature accommodates variable-length time series without requiring users to predetermine subsequence lengths. Our experimental results reveal that$k$-Graph outperforms current state-of-the-art time series clustering algorithms in accuracy, while providing users with meaningful explanations and interpretations of the clustering outcomes.
Paul Boniol, Donato Tiano, Angela Bonifati, Themis Palpanas
IEEE Trans. Knowl. Data Eng.1
2025 VUS: effective and efficient accuracy measures for time-series anomaly detection
Paul Boniol, Ashwin K. Krishna, Marine Bruel, Mingyi Huang, Themis Palpanas, Ruey S. Tsay, Aaron J. Elmore, Michael J. Franklin, John Paparrizos
VLDB J.1
2025 MSAD: A deep dive into model selection for time series anomaly detection
Emmanouil Sylligardos, John Paparrizos, Themis Palpanas, Pierre Senellart, Paul Boniol
VLDB J.5
2024 An Interactive Dive into Time-Series Anomaly Detection
abstract
Anomaly detection is an important problem in data analytics with applications in many domains. In recent years, there has been an increasing interest in anomaly detection tasks applied to time series. In this tutorial, we take a holistic view of anomaly detection in time series, starting from the core definitions and taxonomies related to time series and anomaly types, to an extensive description of the anomaly detection methods proposed by different communities in the literature. We explore the literature and the proposed methods by demonstrating systems that help users understand the core computational steps of some methods and navigate benchmark results. Finally, we describe the problem of model selection for anomaly detection and discuss recent experimental results.
Paul Boniol, John Paparrizos, Themis Palpanas
ICDE1
2024 ADecimo: Model Selection for Time Series Anomaly Detection
abstract
Anomaly detection is a fundamental task for time-series analytics with important implications for the downstream performance of many applications. Despite increasing academic interest and the large number of methods proposed in the literature, recent benchmark and evaluation studies demonstrated that there exists no single best anomaly detection method when applied to heterogeneous time series datasets. Therefore, the only scalable and viable solution to solve anomaly detection over very different time series collected from diverse domains is to propose a model selection method that will choose, based on time series characteristics, the best anomaly detection method to run. This paper describes ADecimo, a modular and extensible web application that helps users understand the performance of time series classification algorithms used as model selection methods for time series anomaly detection. Overall, our system enables users to compare 17 different classifiers over 1980 time series, and decide on the most suitable time series classification method for their own time series and use cases.
Paul Boniol, Emmanouil Sylligardos, John Paparrizos, Panos E. Trahanias, Themis Palpanas
ICDE1
2024 dsymb Playground: An Interactive Tool to Explore Large Multivariate Time Series Datasets
abstract
Exploring and comparing non-stationary multivariate time series is an important problem in many domains and real-world applications. In recent work, we introduced dsymb, a symbolic representation that transforms multivariate time series into interpretable symbolic sequences that comes along with a compatible and efficient distance measure to compare the obtained symbolic sequences. We have shown how dsymbcan handle the non-stationarity of multivariate physiological signals, how interpretable the symbolization is, and how suitable the distance measure is compared to Dynamic Time Warping (DTW) variants. We have also empirically shown that the computation time when using dsymbon a clustering time is significantly smaller than with DTW variants (typically 100 times faster). In this demonstration, we present the dsymbplayground, an interactive web-based tool to interpret and compare a large multivariate time series dataset quickly. We showcase the relevance of this tool in several scenarios based on real-world datasets.
Sylvain W. Combettes, Paul Boniol, Charles Truong, Laurent Oudre
ICDE2
2024 Time-Series Anomaly Detection: Overview and New Trends
abstract
Anomaly detection is a fundamental data analytics task across scientific fields and industries. In recent years, an increasing interest has been shown in the application of anomaly detection techniques to time series. In this tutorial, we take a holistic view of anomaly detection in time series and comprehensively cover detection algorithms ranging from the 1980s to the most current state-of-the-art techniques. Importantly, the scope of this tutorial extends beyond algorithmic discussion, delving into the latest advancements in benchmarking and evaluation measures for this area. In particular, our interactive systems enable the exploration of detection algorithms and benchmarking results, thereby promoting user comprehension. Driven by the absence of a one-size-fits-all anomaly detector for various time series domains and applications, we review recent advancements in automated solutions and propose a new taxonomy to motivate further research.
Paul Boniol, Themis Palpanas, John Paparrizos
Proc. VLDB Endow.2
2023 New Trends in Time Series Anomaly Detection
Paul Boniol, John Paparrizos, Themis Palpanas
EDBT1
2023 Choose Wisely: An Extensive Evaluation of Model Selection for Anomaly Detection in Time Series
abstract
Anomaly detection is a fundamental task for time-series analytics with important implications for the downstream performance of many applications. Despite increasing academic interest and the large number of methods proposed in the literature, recent benchmark and evaluation studies demonstrated that no overall best anomaly detection methods exist when applied to very heterogeneous time series datasets. Therefore, the only scalable and viable solution to solve anomaly detection over very different time series collected from diverse domains is to propose a model selection method that will select, based on time series characteristics, the best anomaly detection method to run. Existing AutoML solutions are, unfortunately, not directly applicable to time series anomaly detection, and no evaluation of time series-based approaches for model selection exists. Towards that direction, this paper studies the performance of time series classification methods used as model selection for anomaly detection. Overall, we compare 17 different classifiers over 1800 time series, and we propose the first extensive experimental evaluation of time series classification as model selection for anomaly detection. Our results demonstrate that model selection methods outperform every single anomaly detection method while being in the same order of magnitude regarding execution time. This evaluation is the first step to demonstrate the accuracy and efficiency of time series classification algorithms for anomaly detection, and represents a strong baseline that can then be used to guide the model selection step in general AutoML pipelines.
Emmanouil Sylligardos, Paul Boniol, John Paparrizos, Panos E. Trahanias, Themis Palpanas
Proc. VLDB Endow.2
2023 Correction to: Unsupervised and scalable subsequence anomaly detection in large data series
Paul Boniol, Michele Linardi, Federico Roncallo, Themis Palpanas, Mohammed Meftah, Emmanuel Remy
VLDB J.1
2022 dCAM: Dimension-wise Class Activation Map for Explaining Multivariate Data Series Classification
abstract
Data series classification is an important and challenging problem in data science. Explaining the classification decisions by finding the discriminant parts of the input that led the algorithm to some decision is a real need in many applications. Convolutional neural networks perform well for the data series classification task; though, the explanations provided by this type of algorithms are poor for the specific case of multivariate data series. Addressing this important limitation is a significant challenge. In this paper, we propose a novel method that solves this problem by highlighting both the temporal and dimensional discriminant information. Our contribution is two-fold: we first describe a convolutional architecture that enables the comparison of dimensions; then, we propose a method that returns dCAM, a Dimension-wise Class Activation Map specifically designed for multivariate time series (and CNN-based models). Experiments with several synthetic and real datasets demonstrate that dCAM is not only more accurate than previous approaches, but the only viable solution for discriminant feature discovery and classification explanation in multivariate time series.
Paul Boniol, Mohammed Meftah, Emmanuel Remy, Themis Palpanas
SIGMOD Conference1
2022 Theseus: Navigating the Labyrinth of Time-Series Anomaly Detection
abstract
The detection of anomalies in time series has gained ample academic and industrial attention, yet, no comprehensive benchmark exists to evaluate time-series anomaly detection methods. Therefore, there is no final verdict on which method performs the best (and under what conditions). Consequently, we often observe methods performing exceptionally well on one dataset but surprisingly poorly on another, creating an illusion of progress. To address these issues, we thoroughly studied over one hundred papers, and summarized our effort in TSB-UAD, a new benchmark to evaluate univariate time series anomaly detection methods. In this paper, we describe Theseus, a modular and extensible web application that helps users navigate through the benchmark, and reason about the merits and drawbacks of both anomaly detection methods and accuracy measures, under different conditions. Overall, our system enables users to compare 12 anomaly detection methods on 1980 time series, using 13 accuracy measures, and decide on the most suitable method and measure for some application.
Paul Boniol, John Paparrizos, Yuhao Kang, Themis Palpanas, Ruey S. Tsay, Aaron J. Elmore, Michael J. Franklin
Proc. VLDB Endow.1
2022 Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly Detection
abstract
Anomaly detection (AD) is a fundamental task for time-series analytics with important implications for the downstream performance of many applications. In contrast to other domains where AD mainly focuses on point-based anomalies (i.e., outliers in standalone observations), AD for time series is also concerned with range-based anomalies (i.e., outliers spanning multiple observations). Nevertheless, it is common to use traditional point-based information retrieval measures, such as Precision, Recall, and F-score, to assess the quality of methods by thresholding the anomaly score to mark each point as an anomaly or not. However, mapping discrete labels into continuous data introduces unavoidable shortcomings, complicating the evaluation of range-based anomalies. Notably, the choice of evaluation measure may significantly bias the experimental outcome. Despite over six decades of attention, there has never been a large-scale systematic quantitative and qualitative analysis of time-series AD evaluation measures. This paper extensively evaluates quality measures for time-series AD to assess their robustness under noise, misalignments, and different anomaly cardinality ratios. Our results indicate that measures producing quality values independently of a threshold (i.e., AUC-ROC and AUC-PR) are more suitable for time-series AD. Motivated by this observation, we first extend the AUC-based measures to account for range-based anomalies. Then, we introduce a new family of parameter-free and threshold-independent measures, VUS (Volume Under the Surface), to evaluate methods while varying parameters. Our findings demonstrate that our four measures are significantly more robust in assessing the quality of time-series AD methods.
John Paparrizos, Paul Boniol, Themis Palpanas, Ruey S. Tsay, Aaron J. Elmore, Michael J. Franklin
Proc. VLDB Endow.2
2022 TSB-UAD: An End-to-End Benchmark Suite for Univariate Time-Series Anomaly Detection
abstract
The detection of anomalies in time series has gained ample academic and industrial attention. However, no comprehensive benchmark exists to evaluate time-series anomaly detection methods. It is common to use (i) proprietary or synthetic data, often biased to support particular claims; or (ii) a limited collection of publicly available datasets. Consequently, we often observe methods performing exceptionally well in one dataset but surprisingly poorly in another, creating an illusion of progress. To address the issues above, we thoroughly studied over one hundred papers to identify, collect, process, and systematically format datasets proposed in the past decades. We summarize our effort in TSB-UAD, a new benchmark to ease the evaluation of univariate time-series anomaly detection methods. Overall, TSB-UAD contains 13766 time series with labeled anomalies spanning different domains with high variability of anomaly types, ratios, and sizes. TSB-UAD includes 18 previously proposed datasets containing 1980 time series and we contribute two collections of datasets. Specifically, we generate 958 time series using a principled methodology for transforming 126 time-series classification datasets into time series with labeled anomalies. In addition, we present data transformations with which we introduce new anomalies, resulting in 10828 time series with varying complexity for anomaly detection. Finally, we evaluate 12 representative methods demonstrating that TSB-UAD is a robust resource for assessing anomaly detection methods. We make our data and code available at www.timeseries.org/TSB-UAD. TSB-UAD provides a valuable, reproducible, and frequently updated resource to establish a leaderboard of univariate time-series anomaly detection methods.
John Paparrizos, Yuhao Kang, Paul Boniol, Ruey S. Tsay, Themis Palpanas, Michael J. Franklin
Proc. VLDB Endow.3
2021 SAND: Streaming Subsequence Anomaly Detection
abstract
With the increasing demand for real-time analytics and decision making, anomaly detection methods need to operate over streams of values and handle drifts in data distribution. Unfortunately, existing approaches have severe limitations: they either require prior domain knowledge or become cumbersome and expensive to use in situations with recurrent anomalies of the same type. In addition, subsequence anomaly detection methods usually require access to the entire dataset and are not able to learn and detect anomalies in streaming settings. To address these problems, we propose SAND, a novel online method suitable for domain-agnostic anomaly detection. SAND aims to detect anomalies based on their distance to a model that represents normal behavior. SAND relies on a novel steaming methodology to incrementally update such model, which adapts to distribution drifts and omits obsolete data. The experimental results on several real-world datasets demonstrate that SAND correctly identifies single and recurrent anomalies without prior knowledge of the characteristics of these anomalies. SAND outperforms by a large margin the current state-of-the-art algorithms in terms of accuracy while achieving orders of magnitude speedups.
Paul Boniol, John Paparrizos, Themis Palpanas, Michael J. Franklin
Proc. VLDB Endow.1
2021 SAND in Action: Subsequence Anomaly Detection for Streams
abstract
Subsequence anomaly detection in long data series is a significant problem. While the demand for real-time analytics and decision making increases, anomaly detection methods have to operate over streams and handle drifts in data distribution. Nevertheless, existing approaches either require prior domain knowledge or become cumbersome and expensive to use in situations with recurrent anomalies of the same type. Moreover, subsequence anomaly detection methods usually require access to the entire dataset and are not able to learn and detect anomalies in streaming settings. To address these limitations, we propose SAND, a novel online system suitable for domain-agnostic anomaly detection. SAND relies on a novel steaming methodology to incrementally update a model that adapts to distribution drifts and omits obsolete data. We demonstrate our system over different streaming scenarios and compare SAND with other subsequence anomaly detection methods.
Paul Boniol, John Paparrizos, Themis Palpanas, Michael J. Franklin
Proc. VLDB Endow.1
2021 Unsupervised and scalable subsequence anomaly detection in large data series
Paul Boniol, Michele Linardi, Federico Roncallo, Themis Palpanas, Mohammed Meftah, Emmanuel Remy
VLDB J.1
2020 SAD: An Unsupervised System for Subsequence Anomaly Detection
abstract
Subsequence anomaly (or outlier) detection in long sequences is an important problem with applications in a wide range of domains. However, current approaches have severe limitations: they either require prior domain knowledge, or become cumbersome and expensive to use in situations with recurrent anomalies of the same type. We recently proposed NorM, a novel approach suitable for domain-agnostic anomaly detection, which addresses the aforementioned problems by detecting anomalies based on their (dis)similarity to a model that represents normal behavior. The experimental results on several real datasets demonstrate that the proposed approach outperforms the current state-of-the art in terms of both accuracy and execution time. In this demonstration, we present a system for unsupervised Subsequence Anomaly Detection (SAD) that uses the NorM method. Through various scenarios with real datasets, we showcase the challenges of the problem, and we demonstrate the advantages of the proposed system.
Paul Boniol, Michele Linardi, Federico Roncallo, Themis Palpanas
ICDE1
2020 Automated Anomaly Detection in Large Sequences
abstract
Subsequence anomaly (or outlier) detection in long sequences is an important problem with applications in a wide range of domains. However, current approaches have severe limitations: they either require prior domain knowledge, or become cumbersome and expensive to use in situations with recurrent anomalies of the same type. In this work, we address these problems, and propose NorM, a novel approach, suitable for domain-agnostic anomaly detection. NorM is based on a new data series primitive, which permits to detect anomalies based on their (dis)similarity to a model that represents normal behavior. The experimental results on several real datasets demonstrate that the proposed approach outperforms by a large margin the current state-of-the art algorithms in terms of accuracy, while being orders of magnitude faster.
Paul Boniol, Michele Linardi, Federico Roncallo, Themis Palpanas
ICDE1
2020 Series2Graph: Graph-based Subsequence Anomaly Detection for Time Series
Paul Boniol, Themis Palpanas
Proc. VLDB Endow.1
2020 GraphAn: Graph-based Subsequence Anomaly Detection
abstract
Subsequence anomaly detection in long sequences is an important problem with applications in a wide range of domains. However, the state-of-the-art approaches have severe limitations: they either require prior domain knowledge, or become cumbersome and inefficient/ineffective in situations with recurrent anomalies of the same type. We recently proposed Series2Graph, a novel method based on a graph representation of a low-dimensionality embedding of subsequences, that detects anomalous subsequences. The experimental results, on the largest set of synthetic and real datasets used to date, demonstrate that the proposed approach correctly identifies single and recurrent anomalies of various types without any prior knowledge of the characteristics of these anomalies, outperforming by a large margin several competing approaches in accuracy, while being up to orders of magnitude faster. In this demonstration, we present GraphAn, a system based on Series2Graph, show-case the challenges of the problem, and demonstrate the advantages of the proposed system.
Paul Boniol, Themis Palpanas, Mohammed Meftah, Emmanuel Remy
Proc. VLDB Endow.1