VLDB 2026 Research / reviewers in the wild / expert
Andreas Lang 0002
dblp:88/5471-2
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-3212-5548ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the evaluation and optimization of LabeledPAM
Miriama Jánosová, Andreas Lang 0002, Petra Budíková, Erich Schubert, Vlastislav Dohnal |
Inf. Syst. | 2 |
| 2024 | Advancing the PAM Algorithm to Semi-supervised k-Medoids Clustering
Miriama Jánosová, Andreas Lang 0002, Petra Budíková, Erich Schubert, Vlastislav Dohnal |
SISAP | 2 |
| 2023 | Accelerating k-Means Clustering with Cover Trees
Andreas Lang 0002, Erich Schubert |
SISAP | 1 |
| 2022 | BETULA: Fast clustering of large data with improved BIRCH CF-TreesabstractBIRCH clustering is a widely known approach for clustering that has influenced much subsequent research and commercial products. The key contribution of BIRCH is the Clustering Feature tree (CF-Tree), which is a compressed representation of the input data. As new data arrives, the tree is eventually rebuilt to increase the compression. Afterward, the leaves of the tree are used for clustering. Because of the data compression, this method is very scalable. The idea has been adopted, for example, for k-means, data stream, and density-based clustering. Clustering features used by BIRCH are simple summary statistics that can easily be updated with new data: the number of points, the linear sums, and the sum of squared values. Unfortunately, how the sum of squares is then used in BIRCH is prone to catastrophic cancellation. We introduce a replacement cluster feature that does not have this numeric problem, that is not much more expensive to maintain, and which makes many computations simpler and, hence, more efficient. These cluster features can also easily be used in other work derived from BIRCH such as algorithms for streaming data. In the experiments, we demonstrate the numerical problem and compare the performance of the original algorithm compared to the improved cluster features. We furthermore explain how to improve clustering with different algorithms such as hierarchical clustering, k-means, k-means++ and Gaussian mixture modeling by using the variance information stored in the cluster features to obtain high-quality result approximations in a substantially reduced runtime (up to 500× faster in our experiments). Andreas Lang 0002, Erich Schubert |
Inf. Syst. | 1 |
| 2021 | Accelerating Spherical k-Means
Erich Schubert, Andreas Lang 0002, Gloria Feher |
SISAP | 2 |
| 2020 | BETULA: Numerically Stable CF-Trees for BIRCH Clustering
Andreas Lang 0002, Erich Schubert |
SISAP | 1 |
| 2019 | Can Flexible Multi-Core Scheduling Help to Execute Machine Learning Algorithms Resource-Efficiently?abstractMachine learning techniques such as model-based optimization are frequently used to solve expensive problems. Since a sequential execution of these algorithms is time-intensive due to the problem complexity, several attempts have been made to parallelize existing approaches. However, no state-of-the-art technique is able to efficiently exploit the full potential of multi-core architectures up to now. In this work, we propose a flexible extension to the Resource-Aware Model-Based Optimization framework (RAMBO) adopting multi-core scheduling techniques, which allows to use the available resources in a more efficient way and thus reduces the time required to solve expensive optimization problems. Helena Kotthaus, Lea Schönberger, Andreas Lang 0002, Jian-Jia Chen, Peter Marwedel |
SCOPES | 3 |