Simon Klüttermann

dblp:339/0303 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-9698-4339ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unsupervised Symbolic Anomaly Detection
Tim Katzke, Simon Klüttermann, Emmanuel Müller
PAKDD (2)3
2025 Evaluating Anomaly Detection Algorithms: The Role of Hyperparameters and Standardized Benchmarks
abstract
Anomaly detection is a cornerstone of machine learning with applications spanning healthcare, fraud detection, and scientific discovery. Despite extensive research, fair bench-marking remains a significant challenge due to the unsupervised nature of anomaly detection. Hyperparameter selection, a crucial determinant of algorithm performance, is often overlooked or bi-ased, leading to inflated or misleading results. Current practices, including reliance on default configurations, random choices, or limited optimization, hinder reproducibility and impede progress. This work presents a novel pipeline for standardized hyperparameter optimization in anomaly detection. Leveraging a curated collection of nearly 500 datasets, the largest of its kind, our approach systematically optimizes over 80 hyperparameters for 13 widely used anomaly detection algorithms. Our comparison re-veals that the performance variance from hyperparameters often surpasses inter-algorithm differences, emphasizing the need for hyperparameter-specific evaluations. We establish a reproducible foundation for anomaly detection research by providing open-access datasets and code. Our findings not only challenge existing evaluation norms but also pave the way for more robust and reliable comparisons toward better anomaly detection research.
Simon Klüttermann, Emmanuel Müller
DSAA1
2025 Unsupervised Surrogate Anomaly Detection
Simon Klüttermann, Tim Katzke, Emmanuel Müller
ECML/PKDD (1)1
2024 Towards Highly Efficient Anomaly Detection for Predictive Maintenance
abstract
This paper introduces SEAN, a novel anomaly detection algorithm designed for real-time applications in predictive maintenance. SEAN leverages an ensemble-based approach to deliver competitive performance while drastically reducing computational costs. In our comprehensive evaluation across 121 datasets, SEAN consistently outperforms comparable shallow anomaly detection algorithms. Our comparisons reveal that SEAN operates over 20,000 times faster than a similar state-of-the-art deep learning alternative, with negligible sacrifice in detection accuracy. We further demonstrate SEAN's versatility through an ablation study, highlighting how its hyperparameters can be tuned to balance runtime and performance effectively. Additionally, we present a practical C++ export tool that enables the deployment of SEAN on resource-constrained devices, meeting the stringent requirements of on-device predictive maintenance tasks. Our findings underscore SEAN as a powerful and efficient solution for anomaly detection in real-world engineering applications.
Simon Klüttermann, Vanlal Peka, Philipp Doebler, Emmanuel Müller
ICMLA1
2024 On the Effectiveness of Heterogeneous Ensemble Methods for Re-Identification
abstract
In this contribution, we introduce a novel ensemble method for the re-identification of industrial entities, using images of chipwood pallets and galvanized metal plates as dataset examples. Our algorithms replace commonly used, complex siamese neural networks with an ensemble of simplified, rudimentary models, providing wider applicability, especially in hardware-restricted scenarios. Each ensemble sub-model uses different types of extracted features of the given data as its input, allowing for the creation of effective ensembles in a fraction of the training duration needed for more complex state-of-the-art models. We reach state-of-the-art performance at our task, with a Rank-1 accuracy of over 77% and a Rank-10 accuracy of over 99%, and introduce five distinct feature extraction approaches, and study their combination using different ensemble methods.
Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski, Britta Grimme, Moritz Roidl, Emmanuel Müller
ICMLA1
2024 Evaluating Anomaly Detection Algorithms: A Multi-Metric Analysis Across Variable Class Imbalances
abstract
Anomaly detection is critical for ensuring integrity and performance in numerous high-stakes domains, ranging from financial systems to network security. The effectiveness of anomaly detection algorithms is significantly affected by the class imbalances that characterize real-world data. This research conducts an in-depth analysis of five anomaly detection algorithms—Angle-based Outlier Detector (ABOD), K-Nearest Neighbors (KNN), Local Outlier Factor (LOF), Isolation Forest (IsoForest), and One-Class SVM (OCSVM). Our evaluation spans datasets with a spectrum of feature complexities and observation volumes, alongside a targeted resampling of anomaly percentages, shifting from a balanced 50/50 distribution to imbalances ranging from 10% to 40%. A multi-metric evaluation framework is deployed, encompassing F1, ROC AUC, PR AUC, MCC, and Kappa, to deliver a layered assessment of algorithmic performance. Our findings reveal distinct stability in the correlations of F1 with Kappa and MCC, and between MCC and Kappa, signifying their potential as consistent performance indicators across various datasets and models. In contrast, F1’s correlation with ROC and PR AUC, and to a lesser degree PR AUC’s correlation with ROC, displayed notable fluctuations, indicating a differential impact of class distribution on these metrics. The study underscores the imperative of utilizing a multi-metric approach for a comprehensive evaluation of anomaly detection algorithms, ensuring adaptability to the diverse and skewed distributions encountered in practice. The insights from this analysis provide a pathway for practitioners to make informed decisions in selecting and deploying anomaly detection models that can withstand the challenges posed by varying class imbalances.
Mohammad-Sahadet Hossain, Mohammad Sakhawat Hossain, Simon Klüttermann, Emmanuel Müller
IJCNN3
2024 The Phenomenon of Correlated Representations in Contrastive Learning
abstract
Contrastive learning is widely considered to be an important domain of machine learning. Its main premise is to use contrasting samples of data in order to learn common features that allow for accurate data clustering in a representation space. The generation of such a representation or embedding space can be of great value, for instance for re-identification tasks. While working on one such task, we noticed that, paradoxically, decreasing training data resolution led to a considerably higher re-identification accuracy. Upon further analysis, we discovered that this occurs since during training, highly correlated features are learned that are ultimately redundant and limit the model’s performance. This effect is exaggerated with the use of high-resolution data, therefore eventually decreasing the obtained re-identification accuracy. In this contribution, we characterize this phenomenon, which we believe to be a novel problem in contrastive learning. We study the effects of various changes to a common neural network architecture on this phenomenon, linking it to the concept of bias, and propose a set of solutions to mitigate its effect.
Simon Klüttermann, Jérôme Rutinowski, Emmanuel Müller
IJCNN1
2024 Autoencoder Optimization for Anomaly Detection: A Comparative Study with Shallow Algorithms
abstract
This paper presents an innovative guide for optimizing autoencoder performance, specifically targeting anomaly detection tasks. In addressing prevalent issues in deep learning algorithms, our primary focus lies in effectively selecting and controlling the latent space in autoencoders. We comprehensively explore methodologies for determining the optimal latent size, a critical and often overlooked aspect in autoencoder architectures. This endeavor forms part of a broader initiative to enhance autoencoder efficacy, ensuring their performance is on par with or superior to many shallow learning algorithms, a challenge highlighted in studies like ADBENCH. Our approach encompasses a detailed examination and experimentation with various parameters, architectures, and loss functions, all aimed at refining the efficiency and accuracy of autoencoders in anomaly detection for image and tabular data.This research stands out for its dual focus on image and tabular datasets. We thoroughly examine the performance of autoencoders in detecting anomalies, utilizing a variety of autoencoder architectures and diverse hyperparameters.
Vishesh Srivastava, Sadia Mahjabin, Arindam Pal 0003, Simon Klüttermann, Emmanuel Müller
IJCNN5
2024 On the Efficient Explanation of Outlier Detection Ensembles Through Shapley Values
Simon Klüttermann, Chiara Balestra, Emmanuel Müller
PAKDD (3)1
2023 Evaluating and Comparing Heterogeneous Ensemble Methods for Unsupervised Anomaly Detection
abstract
Ensembles are one of the most promising research directions for unsupervised anomaly detection. But combining many different models into such an ensemble requires good combination procedures that are able to combine the strengths of many different submodels. To find, evaluate and understand these procedures, we create the biggest experiment to date, including multiple orders of magnitude more ensembles than each of our competitors. Using this high number of comparisons, we also study the effect different normalization methods have on the combination procedure and extract conditional performances of individual models. We use this, to develop a simple set of best practices to create good and reliable anomaly detection ensembles.
Simon Klüttermann, Emmanuel Müller
IJCNN1
2022 Post-Robustifying Deep Anomaly Detection Ensembles by Model Selection
abstract
Anomaly detection has been a major research area in machine learning with deep ensemble models showing exceptional performance. However, formal verification of robustness for anomaly detection in general, and ensemble models in particular, has been mostly neglected. Moreover, given an already trained, non-robust model, there is no way to adapt it for robustness as a post-processing step as of yet. By harnessing properties of ensemble methods - in particular of the DEAN model - we are the first to post-robustify a model via submodel selection. Beyond this new capability, our method significantly increases verification scalability by employing the inherent properties of ensemble methods. Our experiments show that the DEAN model is most suitable for our method: it proves to be the most robust from the start, allows for post-robustification and keeps a stable runtime across all datasets considered.
Benedikt Böing, Simon Klüttermann, Emmanuel Müller
ICDM2
2022 Towards Graph Representation based Re-Identification of Chipwood Pallet Blocks
abstract
This contribution provides a novel graph representation-based approach for the re-identification of chipwood surface structures and, in the herein observed use case, the re-identification of Euro-pallets. For this purpose, we suggest, in contrast to common re-identification approaches, replacing the usual image representation with a highly compressed graph representation. This allows for the creation of an efficient algorithm while also providing robustness to environmental changes such as rotation and shearing. The resulting method, called IRAG (Image Representation through Anomaly Graphs), is a siamese graph neural network, that is applied on a previously published dataset consisting of images from 502 EPAL pallet blocks. The results of this approach lead to a rank-1 accuracy of 27% when re-identifying pallet blocks. Even though IRAG does not yet reach accuracy values that are comparable to state-of-the-art literature, it is however more efficient, concerning the handling and representation of data. In addition, the experiments in this work demonstrate that the re-identification accuracy of the model is not affected by rotation or shearing, demonstrating the model’s invariance to these environmental changes.
Simon Klüttermann, Jérôme Rutinowski, Christopher Reining, Moritz Roidl, Emmanuel Müller
ICMLA1