EDBT 2026 Demo / reviewers in the wild / expert
Lena G. M. Bauer
dblp:256/6566 · also Lena Greta Marie Bauer
· DBLP profile ↗
4ranked-venue papers in the field
1as first author
3since 2021 · last 2023
—ORCID · none
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 4 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Extension of the Dip-test Repertoire - Efficient and Differentiable p-value Calculation for ClusteringabstractOver the last decade, the Dip-test of unimodality has gained increasing interest in the data mining community as it is a parameter-free statistical test that reliably rates the modality in one-dimensional samples. It returns a so called Dip-value and a corresponding probability for the sample's unimodality (Dip-p-value). These two values share a sigmoidal relationship. However, the specific transformation is dependent on the sample size. Many Dip-based clustering algorithms use bootstrapped look-up tables translating Dip- to Dip-p-values for a certain limited amount of sample sizes. We propose a specifically designed sigmoid function as a substitute for these state-of-the-art look-up tables. This accelerates computation and provides an approximation of the Dip- to Dip-p-value transformation for every single sample size. Further, it is differentiable and can therefore easily be integrated in learning schemes using gradient descent. We showcase this by exploiting our function in a novel subspace clustering algorithm called Dip'n’Sub. We highlight in extensive experiments the various benefits of our proposal. Lena G. M. Bauer, Collin Leiber, Christian Böhm 0001, Claudia Plant |
SDM | 1 |
| 2022 | The DipEncoder: Enforcing Multimodality in AutoencodersabstractHartigan's Dip-test of unimodality gained increasing interest in unsupervised learning over the past few years. It is free from complex parameterization and does not require a distribution assumed a priori. A useful property is that the resulting Dip-values can be derived to find a projection axis that identifies multimodal structures in the data set. In this paper, we show how to apply the gradient not only with respect to the projection axis but also with respect to the data to improve the cluster structure. By tightly coupling the Dip-test with an autoencoder, we obtain an embedding that clearly separates all clusters in the data set. This method, called DipEncoder, is the basis of a novel deep clustering algorithm. Extensive experiments show that the DipEncoder is highly competitive to state-of-the-art methods. Collin Leiber, Lena G. M. Bauer, Michael Neumayr, Claudia Plant, Christian Böhm 0001 |
KDD | 2 |
| 2021 | Dip-based Deep Embedded Clustering with k-EstimationabstractThe combination of clustering with Deep Learning has gained much attention in recent years. Unsupervised neural networks like autoencoders can autonomously learn the essential structures in a data set. This idea can be combined with clustering objectives to learn relevant features automatically. Unfortunately, they are often based on a k-means framework, from which they inherit various assumptions, like spherical-shaped clusters. Another assumption, also found in approaches outside the k-means-family, is knowing the number of clusters a-priori. In this paper, we present the novel clustering algorithm DipDECK, which can estimate the number of clusters simultaneously to improving a Deep Learning-based clustering objective. Additionally, we can cluster complex data sets without assuming only spherically shaped clusters. Our algorithm works by heavily overestimating the number of clusters in the embedded space of an autoencoder and, based on Hartigan's Dip-test - a statistical test for unimodality - analyses the resulting micro-clusters to determine which to merge. We show in extensive experiments the various benefits of our method: (1) we achieve competitive results while learning the clustering-friendly representation and number of clusters simultaneously; (2) our method is robust regarding parameters, stable in performance, and allows for more flexibility in the cluster shape; (3) we outperform relevant competitors in the estimation of the number of clusters. Collin Leiber, Lena G. M. Bauer, Benjamin Schelling, Christian Böhm 0001, Claudia Plant |
KDD | 2 |
| 2020 | Utilizing Structure-Rich Features to Improve Clustering
Benjamin Schelling, Lena G. M. Bauer, Sahar Behzadi, Claudia Plant |
ECML/PKDD (1) | 2 |