VLDB 2026 Research / reviewers in the wild / expert
Collin Leiber
dblp:299/4795
· DBLP profile ↗
13ranked-venue papers in the field
4as first author
13since 2021 · last 2026
0000-0001-5368-5697ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 11 (3 first)Database Systems & Data Management · 1Information Retrieval & Web Search · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Khatri-Rao Clustering for Data SummarizationabstractAs datasets continue to grow in size and complexity, finding succinct yet accurate data summaries poses a key challenge. Centroid-based clustering, a widely adopted approach to address this challenge, finds informative summaries of datasets in terms of few prototypes, each representing a cluster in the data. Despite their wide adoption, the resulting data summaries often contain redundancies, limiting their effectiveness particularly in datasets characterized by a large number of underlying clusters. To overcome this limitation, we introduce the Khatri-Rao clustering paradigm that extends traditional centroid-based clustering to produce more succinct but equally accurate data summaries by postulating that centroids arise from the interaction of two or more succinct sets of protocentroids. We study two approaches to centroid-based clustering, the well-established -Means algorithm and the increasingly popular deep clustering, under the lens of the Khatri-Rao paradigm. To this end, we introduce the Khatri-Rao--Means algorithm and the Khatri-Rao deep clustering framework. Extensive experiments show that Khatri-Rao--Means can strike a more favorable trade-off between succinctness and accuracy in data summarization than standard -Means. Leveraging representation learning, the Khatri-Rao deep clustering framework offerseven greater benefits, reducing even more the size of data summaries given by deep clustering while preserving their accuracy Martino Ciaperoni, Collin Leiber, Aristides Gionis, Heikki Mannila |
EDBT | 2 |
| 2026 | CODAC: Constraint-based Deep Active ClusteringabstractAbstract Constraint-based Deep Active Clustering (CODAC) integrates actively selected pairwise constraints into deep representation learning to efficiently improve existing cluster structures, even under tight query budgets. CODAC encodes the constraint information into the embedding so that the learned representation can generalize to unconstrained data, leading to a rapid improvement of the clustering quality even on large datasets. CODAC makes minimal assumptions regarding the data and can be combined with a wide variety of deep clustering models. It does not require the number of clusters to be known a priori and is even effective if the initial estimate is badly misspecified. Across diverse image, text, and tabular datasets, CODAC consistently attains higher cluster quality with fewer queries than the previous state-of-the-art, and can substantially improve the clustering quality with just 100–200 queries compared to the deep clustering baselines. Anri Patron, Sandra Gilhuber, Kai Puolamäki, Collin Leiber |
Data Min. Knowl. Discov. | 4 |
| 2025 | DCMatch - Identify Matching Architectures in Deep Clustering Through Meta-learning
Mamdouh Aljoud, Gabriel Marques Tavares, Collin Leiber, Thomas Seidl 0001 |
PAKDD (1) | 3 |
| 2025 | Going Offline: An Evaluation of the Offline Phase in Stream Clustering
Philipp Jahn 0001, Walid Durani, Collin Leiber, Anna Beer 0001, Thomas Seidl 0001 |
ECML/PKDD (7) | 3 |
| 2024 | SHADE: Deep Density-based ClusteringabstractDetecting arbitrarily shaped clusters in high-dimensional noisy data is challenging for current clustering methods. We introduce SHADE, the first deep clustering algorithm that incorporates density-connectivity into its loss function. Similar to existing deep clustering algorithms, SHADE supports high-dimensional and large data sets with the expressive power of a deep autoencoder. In contrast to most existing deep clustering methods that rely on a centroid-based clustering objective, SHADE incorporates a novel loss function that captures density-connectivity. It thereby learns a representation that enhances the separation of density-connected clusters. SHADE detects a stable clustering and noise points fully automatically without any user input. It outperforms existing methods in clustering quality, especially on data that contain non-Gaussian clusters, such as video data. Moreover, the embedded space of SHADE is suitable for visualization and interpretation of the clustering results as the individual shapes of the clusters are preserved. Anna Beer 0001, Pascal Weber 0001, Lukas Miklautz, Collin Leiber, Walid Durani, Christian Böhm 0001, Claudia Plant |
ICDM | 4 |
| 2024 | Data with Density-Based Clusters: A Generator for Systematic Evaluation of Clustering Algorithms
Philipp Jahn 0001, Christian M. M. Frey, Anna Beer 0001, Collin Leiber, Thomas Seidl 0001 |
ECML/PKDD (7) | 4 |
| 2023 | Application of Deep Clustering AlgorithmsabstractDeep clustering algorithms have gained popularity for clustering complex, large-scale data sets, but getting started is difficult because of numerous decisions regarding architecture, optimizer, and other hyperparameters. Theoretical foundations must be known to obtain meaningful results. At the same time, ease of use is necessary to get used by a broader audience. Therefore, we require a unified framework that allows for easy execution in diverse settings. While this applies to established clustering methods like k-Means and DBSCAN, deep clustering algorithms lack a standard structure, resulting in significant programming overhead. This complicates empirical evaluations, which are essential in both scientific and practical applications. We present a solution to this problem by providing a theoretical background on deep clustering as well as practical implementation techniques and a unified structure with predefined neural networks. For the latter, we use the Python package ClustPy. The aim is to share best practices and facilitate community participation in deep clustering research. Collin Leiber, Lukas Miklautz, Claudia Plant, Christian Böhm 0001 |
CIKM | 1 |
| 2023 | Non-Redundant Image Clustering of Early Medieval Glass BeadsabstractGlass beads were among the most common grave goods in the Early Middle Ages, with an estimated number in the millions. The color, size, shape and decoration of the beads are diverse leading to many different archaeological classification systems that depend on the subjective decisions of individual experts. The lack of an agreed upon expert categorization leads to a pressing problem in archaeology, as the categorization of archaeological artifacts, like glass beads, is important to learn about cultural trends, manufacturing processes or economic relationships (e.g., trade routes) of historical times. An automated, objective and reproducible classification system is therefore highly desirable. We present a high-quality data set of images of Early Medieval beads and propose a clustering pipeline to learn a classification system in a data-driven way. The pipeline consists of a novel extension of deep embedded non-redundant clustering to identify multiple, meaningful clusterings of glass bead images. During the cluster analysis we address several challenges associated with the data and as a result identify high-quality clusterings that overlap with archaeological domain expertise. To the best of our knowledge this is the first application of non-redundant image clustering for archaeological data. Lukas Miklautz, Andrii Shkabrii, Collin Leiber, Bendeguz Tobias, Benedict Seidl, Elisabeth Weissensteiner, Andreas Rausch 0001, Christian Böhm 0001, Claudia Plant |
DSAA | 3 |
| 2023 | k-SubMix: Common Subspace Clustering on Mixed-Type Data
Mauritius Klein, Collin Leiber, Christian Böhm 0001 |
ECML/PKDD (1) | 2 |
| 2023 | Extension of the Dip-test Repertoire - Efficient and Differentiable p-value Calculation for ClusteringabstractOver the last decade, the Dip-test of unimodality has gained increasing interest in the data mining community as it is a parameter-free statistical test that reliably rates the modality in one-dimensional samples. It returns a so called Dip-value and a corresponding probability for the sample's unimodality (Dip-p-value). These two values share a sigmoidal relationship. However, the specific transformation is dependent on the sample size. Many Dip-based clustering algorithms use bootstrapped look-up tables translating Dip- to Dip-p-values for a certain limited amount of sample sizes. We propose a specifically designed sigmoid function as a substitute for these state-of-the-art look-up tables. This accelerates computation and provides an approximation of the Dip- to Dip-p-value transformation for every single sample size. Further, it is differentiable and can therefore easily be integrated in learning schemes using gradient descent. We showcase this by exploiting our function in a novel subspace clustering algorithm called Dip'n’Sub. We highlight in extensive experiments the various benefits of our proposal. Lena G. M. Bauer, Collin Leiber, Christian Böhm 0001, Claudia Plant |
SDM | 2 |
| 2022 | The DipEncoder: Enforcing Multimodality in AutoencodersabstractHartigan's Dip-test of unimodality gained increasing interest in unsupervised learning over the past few years. It is free from complex parameterization and does not require a distribution assumed a priori. A useful property is that the resulting Dip-values can be derived to find a projection axis that identifies multimodal structures in the data set. In this paper, we show how to apply the gradient not only with respect to the projection axis but also with respect to the data to improve the cluster structure. By tightly coupling the Dip-test with an autoencoder, we obtain an embedding that clearly separates all clusters in the data set. This method, called DipEncoder, is the basis of a novel deep clustering algorithm. Extensive experiments show that the DipEncoder is highly competitive to state-of-the-art methods. Collin Leiber, Lena G. M. Bauer, Michael Neumayr, Claudia Plant, Christian Böhm 0001 |
KDD | 1 |
| 2022 | Automatic Parameter Selection for Non-Redundant ClusteringabstractHigh-dimensional datasets often contain multiple meaningful clusterings in different subspaces. For example, objects can be clustered either by color, weight, or size, revealing different interpretations of the given dataset. A variety of approaches are able to identify such non-redundant clusterings. However, most of these methods require the user to specify the expected number of subspaces and clusters for each subspace. Stating these values is a non-trivial problem and usually requires detailed knowledge of the input dataset. In this paper, we propose a framework that utilizes the Minimum Description Length Principle (MDL) to detect the number of subspaces and clusters per subspace automatically. We describe an efficient procedure that greedily searches the parameter space by splitting and merging subspaces and clusters within subspaces. Additionally, an encoding strategy is introduced that allows us to detect outliers in each subspace. Extensive experiments show that our approach is highly competitive to state-of-the-art methods. Collin Leiber, Dominik Mautz, Claudia Plant, Christian Böhm 0001 |
SDM | 1 |
| 2021 | Dip-based Deep Embedded Clustering with k-EstimationabstractThe combination of clustering with Deep Learning has gained much attention in recent years. Unsupervised neural networks like autoencoders can autonomously learn the essential structures in a data set. This idea can be combined with clustering objectives to learn relevant features automatically. Unfortunately, they are often based on a k-means framework, from which they inherit various assumptions, like spherical-shaped clusters. Another assumption, also found in approaches outside the k-means-family, is knowing the number of clusters a-priori. In this paper, we present the novel clustering algorithm DipDECK, which can estimate the number of clusters simultaneously to improving a Deep Learning-based clustering objective. Additionally, we can cluster complex data sets without assuming only spherically shaped clusters. Our algorithm works by heavily overestimating the number of clusters in the embedded space of an autoencoder and, based on Hartigan's Dip-test - a statistical test for unimodality - analyses the resulting micro-clusters to determine which to merge. We show in extensive experiments the various benefits of our method: (1) we achieve competitive results while learning the clustering-friendly representation and number of clusters simultaneously; (2) our method is robust regarding parameters, stable in performance, and allows for more flexibility in the cluster shape; (3) we outperform relevant competitors in the estimation of the number of clusters. Collin Leiber, Lena G. M. Bauer, Benjamin Schelling, Christian Böhm 0001, Claudia Plant |
KDD | 1 |