VLDB 2026 Research / reviewers in the wild / expert
Georg Stefan Schlake
dblp:295/1795
· DBLP profile ↗
9ranked-venue papers in the field
6as first author
9since 2021 · last 2025
0009-0008-5714-1804ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4 (3 first)Database Systems & Data Management · 2 (1 first)Data Mining & Knowledge Discovery · 2 (2 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Arbitrary Shaped Clustering Validation on the Test Bench
Georg Stefan Schlake, Christian Beecks |
DATA | 1 |
| 2025 | Multi Algorithm Selection and Hyperparameter Optimization for Automated ClusteringabstractA known technique to enable inexperienced users to apply sophisticated machine learning models is Automated Machine Learning (AutoML). AutoML can be used to determine appropriate algorithms and hyperparameters to specific problem settings. Typically, Hyperparameter Optimization (HPO) results in hyperparameters of a single algorithm, optimizing a given objective function. In the domain of automated clustering, the problem inherently has a non-unique solution, often yielding multiple valid algorithms. This multiplicity arises due to the exis-tence of various clustering algorithms, each employing different heuristics and optimization strategies as well as the subjectivity of the clustering notion. Consequently, the search space for finding an optimal clustering solution is highly complex, comprising the selection and parameter tuning of multiple algorithms. In this paper, we introduce the problem of Multi Algorithm Selection, which aims to find an appropriate set of algorithms of unknown size along with their respective hyperparameters, for automated clustering. To this end, we propose various search space strategies and utilize them as inputs for several well-known HPO algorithms to determine an optimal set of clustering algorithms. The resulting pipelines are evaluated on a number of publicly available, synthetic datasets to assess the ability of our methods to find interesting clustering sets. The strategy used to model a complex search space significantly influences the quality of the resulting model. Furthermore, clustering sets generated by means of our proposal yield higher quality than those created using naive approaches. Georg Stefan Schlake, Max Pernklau, Christian Beecks |
DSAA | 1 |
| 2024 | Discovering Propagating Signals in High-Content Multivariate Time Series via Spatio-Temporal Subsequence ClusteringabstractBig data technologies have been applied successfully to diverse application domains in order to facilitate analytics of voluminous and heterogeneous databases at scale. Digital sensory typically provides high-content data comprising multiple data recordings with high frequency. One example of such sensory are multi-electrode arrays (MEA), which are able to measure electric cell activity with high spatial and temporal resolution. The resulting multivariate time series and their inherent subsequences can then be analyzed and compared in aspects of time, space and shape. This analytical process is frequently performed manually by domain experts in combination with data analytical methods that help to identify and track signal beginnings, signal ends and signal propagations.In this paper, we propose an unsupervised approach to discover propagating signals in high-content, multivariate time series databases. To this end, we introduce an efficient spatio-temporal subsequence clustering algorithm that detects and tracks spatial and temporal signal progagations by means of density-based clusters. We present a formal propagation model and show how to adapt the DBSCAN algorithm to our specific application setting on pharmacological data. Our empirical investigation shows that our proposal is able to detect signal propagations with high accuracy and efficiency. Our approach hence scales not only to pharmacological settings but also to other biological, medical, and chemical domains making use of high-content multi-electrode array data. Jan David Hüwel, Georg Stefan Schlake, Kevin Albrechts, Christian Beecks |
IEEE Big Data | 2 |
| 2024 | The Skyline Operator to Find the Needle in the Haystack for Automated ClusteringabstractThe analysis of big datasets is a challenging task. While many data scientists are working in the field of supervised data analysis, there is also a growing demand in the field of unsupervised data analysis, such as clustering. To come up with a solution for this, multiple AutoML approaches for clustering have been proposed. However, most of these approaches try to find the "best" clustering, ignoring the subjective nature of the clustering task. A domain expert, however, might be able to identify an appropriate clustering for his/her application in a small set of clusterings, which have been generated, even if he/she is not capable of creating these clusterings by themselves. To enable domain experts to identify valuable clusterings without becoming an expert in clustering as well, we propose to generate multiple clusterings via AutoML processes and to return a selection of clusterings, from which the user can select the most preferred one. We will investigate the use of the Skyline Operator in this use case, to prune clusterings, which are likely useless, and to find a number of clusterings, which are usable for domain experts. We will investigate, how many clusters can be pruned this way and how many valuable clusters get falsely pruned. Our empirical investigation is carried out on a number of synthetic datasets, where a known ground truth can proxy for the wishes of a domain expert and multiple properties of the clusterings can be known beforehand. Georg Stefan Schlake, Christian Beecks |
IEEE Big Data | 1 |
| 2024 | Automated Exploratory ClusteringabstractClustering is a frequently encountered task in big data analytics, where the goal is to simultaneously group and separate similar and dissimilar objects, respectively. It is also a well known fact, that clustering has a highly subjective nature, in the sense that determining the best clustering is highly dependent on the application setting. Though the recently established research direction of Automated Clustering has originated different algorithmic solutions to the clustering problem, these approaches assume a defined clustering evaluation metric to be optimized. These approaches thus inherently assume that such a thing like a single best clustering exists, which is not always true in real applications where insight into the data comes when inspecting the resulting clusterings.In order to maximize the insight for a data scientists or a domain specialist, we propose to not solely investigate a single best clustering but instead to explore multiple best clusterings according to different evaluation criteria. This will not only help to identify several clusters of interest to the user, but also to maximize the impact gained from following different evaluation criteria. In this paper, we hence propose the concept of Automated Exploratory Clustering, which follows the idea of automatically providing the best clusterings for further exploration. To this end, we formalize the problem of Automated Exploratory Clustering and define a theoretic framework comprising necessary formulations. In addition, we propose an efficient algorithm to compute the most interesting clusterings and benchmark its effectiveness and efficiency. Our approach will help domain experts without expertise in clustering to gain new insights in their datasets and serves as a baseline for future research. Georg Stefan Schlake, Max Pernklau, Christian Beecks |
IEEE Big Data | 1 |
| 2024 | Validating Arbitrary Shaped Clusters - A SurveyabstractClustering is a fundamental method for advanced data analytics. Not only the selection of a suitable clustering method, but also the choice of the resulting clustering, which complies with the application requirements and the data-analytical hypotheses, is a challenge for complex analysis settings. While there exists a multitude of different clustering algorithms for the computation of simple convex up to arbitrary shaped clusters, the question of how to quantify the quality of each individual clustering remains a challenge. In this paper, we investigate the ability of state-of-the-art Clustering Validation Indices (CVI) to assess clustering performance. To this end, we provide a survey of the inner workings of the different CVI and an extensive benchmark on 180 publicly available datasets. Furthermore, we evaluate both the Euclidean distance and the density-based DC-distance to quantify the quality of arbitrary shaped clusters. Our performance evaluation indicates that no singular CVI performs significantly better than the others in general and that the density-based DC-distance is well suited for finding arbitrary shaped clusters even with CVI not specifically designed for this task. Moreover, we discovered that no single CVI effectively performs well for both arbitrary shaped and overlapping clusters at the same time. Our survey provides a comprehensive analysis of CVI from both a theoretical and a practical point of view, and is thus a useful guideline for researchers and practitioners in academia and business. Georg Stefan Schlake, Christian Beecks |
DSAA | 1 |
| 2024 | Identifying Propagating Signals with Spatio-Temporal Clustering in Multivariate Time Series
Jan David Hüwel, Georg Stefan Schlake, Kevin Albrechts, Christian Beecks |
SISAP | 2 |
| 2023 | Towards Automated ClusteringabstractAutomated Machine Learning enables many inexperienced users to generate good classification solutions solely by letting machines generate optimal pipelines. However, there are not as many possibilities to generate meaningful clusterings without further knowledge. As the clustering problem is context and domain dependent, a single solution can never be the sole best clustering for any dataset. To overcome this problem, in this paper we design a framework which uses clustering algorithms, CVIs, the skyline operator and multiple visualization techniques to generate diverse interesting clusterings and present them to a user, enabling him to take an informed decision without needing any experience in the field of clustering. Georg Stefan Schlake, Christian Beecks |
IEEE Big Data | 1 |
| 2022 | A Comparative Performance Analysis of Fast K-Means Clustering Algorithms
Christian Beecks, Fabian Berns, Jan David Hüwel, Andrea Linxen, Georg Stefan Schlake, Tim Düsterhus |
iiWAS | 5 |