VLDB 2026 Research / reviewers in the wild / expert
Sebastian Damrich
dblp:252/5237
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0003-1394-6236ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
4 papers |
Graph algorithms and graph theory · 58% Computational geometry · 17% Algorithms and data structures · 16% | |
| Artificial intelligence
7 papers |
Representation and self-supervised learning · 78% Deep learning architectures and training · 15% Probabilistic and Bayesian machine learning · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 100% |
Topics — the 23 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph algorithms and graph theory
spectral graph theory |
1.6 | 3 | 2024 | Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024 Directed Probabilistic Watershed · NeurIPS 2021 Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Machine learning › Representation and self-supervised learning
contrastive learning |
1.5 | 2 | 2025 | TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025 From $t$-SNE to UMAP with contrastive learning · ICLR 2023 |
Graph algorithms and graph theory › spectral graph theory
effective resistance |
0.9 | 2 | 2021 | Directed Probabilistic Watershed · NeurIPS 2021 Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Graph algorithms and graph theory › graph learning
graph-based semi-supervised learning |
0.9 | 2 | 2021 | Directed Probabilistic Watershed · NeurIPS 2021 Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Algorithms and data structures › signal processing algorithms
watershed algorithm |
0.9 | 2 | 2021 | Directed Probabilistic Watershed · NeurIPS 2021 Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis |
0.9 | 1 | 2025 | TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025 |
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural population decoding |
0.9 | 1 | 2025 | TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025 |
Computational geometry › topological data analysis
persistent homology |
0.8 | 1 | 2024 | Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024 |
Information theory
spectral distance measures |
0.8 | 1 | 2024 | Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024 |
Computational geometry
topological data analysis |
0.8 | 1 | 2024 | Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
autoencoder |
0.7 | 1 | 2023 | Geometric Autoencoders - What You See is What You Decode · ICML 2023 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.7 | 1 | 2023 | From $t$-SNE to UMAP with contrastive learning · ICLR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
t-SNE |
0.7 | 1 | 2023 | From $t$-SNE to UMAP with contrastive learning · ICLR 2023 |
Visualization and visual analytics › dimensionality reduction
dimensionality reduction visualization |
0.7 | 1 | 2023 | Geometric Autoencoders - What You See is What You Decode · ICML 2023 |
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing |
0.6 | 1 | 2022 | Visualizing hierarchies in scRNA-seq data using a density tree-biased autoencoder · Bioinform. 2022 |
Algorithms and data structures › symbolic computation › computational algebra › algebraic algorithms
algebraic path problem |
0.6 | 1 | 2022 | The Algebraic Path Problem for Graph Metrics · ICML 2022 |
Graph algorithms and graph theory
shortest path |
0.6 | 1 | 2022 | The Algebraic Path Problem for Graph Metrics · ICML 2022 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.5 | 1 | 2021 | On UMAP's True Loss Function · NeurIPS 2021 |
Visualization and visual analytics
dimensionality reduction |
0.5 | 1 | 2021 | On UMAP's True Loss Function · NeurIPS 2021 |
Visualization and visual analytics › dimensionality reduction
UMAP |
0.5 | 1 | 2021 | On UMAP's True Loss Function · NeurIPS 2021 |
Graph algorithms and graph theory
directed graph |
0.5 | 1 | 2021 | Directed Probabilistic Watershed · NeurIPS 2021 |
Graph algorithms and graph theory › graph theory › spanning forest
minimum spanning forest |
0.4 | 1 | 2019 | Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Graph algorithms and graph theory › graph theory
spanning forest |
0.4 | 1 | 2019 | Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learning · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.4gibbs distribution · 1.8neighbor embedding · 1.7regularization · 1.3differential geometry · 1.3stochastic gradient descent · 1.0negative sampling · 1.0matrix-tree theorem · 0.9matrix tree theorem · 0.9hypersphere embedding · 0.9cosine similarity · 0.9k-nearest neighbor graph · 0.8effective resistance · 0.8diffusion distance · 0.8vector quantization · 0.6semiring theory · 0.6density-based maximum spanning tree · 0.6bimonoid theory · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Importance of Embedding Norms in Self-Supervised LearningabstractSelf-supervised learning (SSL) allows training data representations without a supervised signal and has become an important paradigm in machine learning. Most SSL methods employ the cosine similarity between embedding vectors and hence effectively embed data on a hypersphere. While this seemingly implies that embedding norms cannot play any role in SSL, a few recent works have suggested that embedding norms have properties related to network convergence and confidence. In this paper, we resolve this apparent contradiction and systematically establish the embedding norm's role in SSL training. Using theoretical analysis, simulations, and experiments, we show that embedding norms (i) govern SSL convergence rates and (ii) encode network confidence, with smaller norms corresponding to unexpected samples. Additionally, we show that manipulating embedding norms can have large effects on convergence speed.
Our findings demonstrate that SSL embedding norms are integral to understanding and optimizing network behavior. Andrew Draganov, Sharvaree Vadgama, Sebastian Damrich, Jan Niklas Böhm, Lucas Maes, Dmitry Kobak, Erik J. Bekkers |
ICML | 3 |
| 2025 | TRACE: Contrastive learning for multi-trial time series data in neuroscienceabstractModern neural recording techniques such as two-photon imaging or Neuropixel probes allow to acquire vast time-series datasets with responses of hundreds or thousands of neurons. Contrastive learning is a powerful self-supervised framework for learning representations of complex datasets. Existing applications for neural time series rely on generic data augmentations and do not exploit the multi-trial data structure inherent in many neural datasets. Here we present TRACE, a new contrastive learning framework that averages across different subsets of trials to generate positive pairs. TRACE allows to directly learn a two-dimensional embedding, combining ideas from contrastive learning and neighbor embeddings. We show that TRACE outperforms other methods, resolving fine response differences in simulated data. Further, using in vivo recordings, we show that the representations learned by TRACE capture both biologically relevant continuous variation, cell-type-related cluster structure, and can assist data quality control. Lisa Schmors, Dominic Gonschorek, Jan Niklas Böhm, Yongrong Qiu, Na Zhou, Dmitry Kobak, Andreas S. Tolias, Fabian H. Sinz, Jacob Reimer, Katrin Franke, Sebastian Damrich, Philipp Berens |
NeurIPS | 11 |
| 2024 | Persistent Homology for High-dimensional Data Based on Spectral MethodsabstractPersistent homology is a popular computational tool for analyzing the topology of point clouds, such as the presence of loops or voids. However, many real-world datasets with low intrinsic dimensionality reside in an ambient space of much higher dimensionality. We show that in this case traditional persistent homology becomes very sensitive to noise and fails to detect the correct topology. The same holds true for existing refinements of persistent homology. As a remedy, we find that spectral distances on the k-nearest-neighbor graph of the data, such as diffusion distance and effective resistance, allow to detect the correct topology even in the presence of high-dimensional noise. Moreover, we derive a novel closed-form formula for effective resistance, and describe its relation to diffusion distances. Finally, we apply these methods to high-dimensional single-cell RNA-sequencing data and show that spectral distances allow robust detection of cell cycle loops. Sebastian Damrich, Philipp Berens, Dmitry Kobak |
NeurIPS | 1 |
| 2023 | From $t$-SNE to UMAP with contrastive learning
Sebastian Damrich, Jan Niklas Böhm, Fred A. Hamprecht, Dmitry Kobak |
ICLR | 1 |
| 2023 | Geometric Autoencoders - What You See is What You DecodeabstractVisualization is a crucial step in exploratory data analysis. One possible approach is to train an autoencoder with low-dimensional latent space. Large network depth and width can help unfolding the data. However, such expressive networks can achieve low reconstruction error even when the latent representation is distorted. To avoid such misleading visualizations, we propose first a differential geometric perspective on the decoder, leading to insightful diagnostics for an embedding’s distortion, and second a new regularizer mitigating such distortion. Our “Geometric Autoencoder” avoids stretching the embedding spuriously, so that the visualization captures the data structure more faithfully. It also flags areas where little distortion could not be achieved, thus guarding against misinterpretation. Philipp Nazari, Sebastian Damrich, Fred A. Hamprecht |
ICML | 2 |
| 2022 | The Algebraic Path Problem for Graph MetricsabstractFinding paths with optimal properties is a foundational problem in computer science. The notions of shortest paths (minimal sum of edge costs), minimax paths (minimal maximum edge weight), reliability of a path and many others all arise as special cases of the "algebraic path problem" (APP). Indeed, the APP formalizes the relation between different semirings such as min-plus, min-max and the distances they induce. We here clarify, for the first time, the relation between the potential distance and the log-semiring. We also define a new unifying family of algebraic structures that include all above-mentioned path problems as well as the commute cost and others as special or limiting cases. The family comprises not only semirings but also strong bimonoids (that is, semirings without distributivity). We call this new and very general distance the "log-norm distance". Finally, we derive some sufficient conditions which ensure that the APP associated with a semiring defines a metric over an arbitrary graph. Enrique Fita Sanmartin, Sebastian Damrich, Fred A. Hamprecht |
ICML | 2 |
| 2022 | Visualizing hierarchies in scRNA-seq data using a density tree-biased autoencoderabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) allows studying the development of cells in unprecedented detail. Given that many cellular differentiation processes are hierarchical, their scRNA-seq data are expected to be approximately tree-shaped in gene expression space. Inference and representation of this tree structure in two dimensions is highly desirable for biological interpretation and exploratory analysis. RESULTS: Our two contributions are an approach for identifying a meaningful tree structure from high-dimensional scRNA-seq data, and a visualization method respecting the tree structure. We extract the tree structure by means of a density-based maximum spanning tree on a vector quantization of the data and show that it captures biological information well. We then introduce density-tree biased autoencoder (DTAE), a tree-biased autoencoder that emphasizes the tree structure of the data in low dimensional space. We compare to other dimension reduction methods and demonstrate the success of our method both qualitatively and quantitatively on real and toy data. AVAILABILITY AND IMPLEMENTATION: Our implementation relying on PyTorch and Higra is available at github.com/hci-unihd/DTAE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Quentin Garrido, Sebastian Damrich, Alexander Jäger, Dario Cerletti, Manfred Claassen, Laurent Najman, Fred A. Hamprecht |
Bioinform. | 2 |
| 2021 | On UMAP's True Loss FunctionabstractUMAP has supplanted $t$-SNE as state-of-the-art for visualizing high-dimensional datasets in many disciplines, but the reason for its success is not well understood. In this work, we investigate UMAP's sampling based optimization scheme in detail. We derive UMAP's true loss function in closed form and find that it differs from the published one in a dataset size dependent way. As a consequence, we show that UMAP does not aim to reproduce its theoretically motivated high-dimensional UMAP similarities. Instead, it tries to reproduce similarities that only encode the $k$ nearest neighbor graph, thereby challenging the previous understanding of UMAP's effectiveness. Alternatively, we consider the implicit balancing of attraction and repulsion due to the negative sampling to be key to UMAP's success. We corroborate our theoretical findings on toy and single cell RNA sequencing data. Sebastian Damrich, Fred A. Hamprecht |
NeurIPS | 1 |
| 2021 | Directed Probabilistic WatershedabstractThe Probabilistic Watershed is a semi-supervised learning algorithm applied on undirected graphs. Given a set of labeled nodes (seeds), it defines a Gibbs probability distribution over all possible spanning forests disconnecting the seeds. It calculates, for every node, the probability of sampling a forest connecting a certain seed with the considered node. We propose the "Directed Probabilistic Watershed", an extension of the Probabilistic Watershed algorithm to directed graphs. Building on the Probabilistic Watershed, we apply the Matrix Tree Theorem for directed graphs and define a Gibbs probability distribution over all incoming directed forests rooted at the seeds. Similar to the undirected case, this turns out to be equivalent to the Directed Random Walker. Furthermore, we show that in the limit case in which the Gibbs distribution has infinitely low temperature, the labeling of the Directed Probabilistic Watershed is equal to the one induced by the incoming directed forest of minimum cost. Finally, for illustration, we compare the empirical performance of the proposed method with other semi-supervised segmentation methods for directed graphs. Enrique Fita Sanmartin, Sebastian Damrich, Fred A. Hamprecht |
NeurIPS | 2 |
| 2019 | Probabilistic Watershed: Sampling all spanning forests for seeded segmentation and semi-supervised learningabstractThe seeded Watershed algorithm / minimax semi-supervised learning on a graph computes a minimum spanning forest which connects every pixel / unlabeled node to a seed / labeled node. We propose instead to consider all possible spanning forests and calculate, for every node, the probability of sampling a forest connecting a certain seed with that node. We dub this approach "Probabilistic Watershed". Leo Grady (2006) already noted its equivalence to the Random Walker / Harmonic energy minimization. We here give a simpler proof of this equivalence and establish the computational feasibility of the Probabilistic Watershed with Kirchhoff's matrix tree theorem. Furthermore, we show a new connection between the Random Walker probabilities and the triangle inequality of the effective resistance. Finally, we derive a new and intuitive interpretation of the Power Watershed. Enrique Fita Sanmartin, Sebastian Damrich, Fred A. Hamprecht |
NeurIPS | 2 |