Dmitry Kobak

dblp:236/5191 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-5639-7209ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Representation and self-supervised learning · 62% Learning theory · 17% Deep learning architectures and training · 12%
Theoretical computer science
1 paper
Computational geometry · 50% Graph algorithms and graph theory · 25% Information theory · 25%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
2.232025
TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025
From $t$-SNE to UMAP with contrastive learning · ICLR 2023
Unsupervised visualization of image datasets using contrastive learning · ICLR 2023
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
t-SNE
1.222023
From $t$-SNE to UMAP with contrastive learning · ICLR 2023
Attraction-Repulsion Spectrum in Neighbor Embeddings · J. Mach. Learn. Res. 2022
Visualization and visual analytics
dimensionality reduction
1.222023
Unsupervised visualization of image datasets using contrastive learning · ICLR 2023
Attraction-Repulsion Spectrum in Neighbor Embeddings · J. Mach. Learn. Res. 2022
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis
0.912025
TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural population decoding
0.912025
TRACE: Contrastive learning for multi-trial time series data in neuroscience · NeurIPS 2025
Computational geometry › topological data analysis
persistent homology
0.812024
Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024
Information theory
spectral distance measures
0.812024
Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024
Graph algorithms and graph theory
spectral graph theory
0.812024
Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024
Computational geometry
topological data analysis
0.812024
Persistent Homology for High-dimensional Data Based on Spectral Methods · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction
0.712023
From $t$-SNE to UMAP with contrastive learning · ICLR 2023
Visualization and visual analytics › dimensionality reduction
neighbor embedding
0.612022
Attraction-Repulsion Spectrum in Neighbor Embeddings · J. Mach. Learn. Res. 2022
Machine learning › Learning theory
high-dimensional statistics
0.412020
The Optimal Ridge Penalty for Real-world High-dimensional Data Can Be Zero or Negative due to the Implicit Ridge Regularization · J. Mach. Learn. Res. 2020
Machine learning › Learning theory › statistical learning theory › regularization theory
minimum-norm least-squares
0.412020
The Optimal Ridge Penalty for Real-world High-dimensional Data Can Be Zero or Negative due to the Implicit Ridge Regularization · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression
0.412020
The Optimal Ridge Penalty for Real-world High-dimensional Data Can Be Zero or Negative due to the Implicit Ridge Regularization · J. Mach. Learn. Res. 2020
Machine learning › Learning theory
inductive bias
0.212024
Scaling Down Deep Learning with MNIST-1D · ICML 2024

Methods — techniques the papers use, named apart from their topics

contrastive learning · 3.7neighbor embedding · 1.7t-SNE · 1.3negative sampling · 1.1attractive-repulsive force · 1.1hypersphere embedding · 0.9cosine similarity · 0.9self-supervised learning · 0.8meta-learning · 0.8lottery ticket hypothesis · 0.8k-nearest neighbor graph · 0.8effective resistance · 0.8diffusion distance · 0.8
YearPublicationVenuePosition
2025 On the Importance of Embedding Norms in Self-Supervised Learning
abstract
Self-supervised learning (SSL) allows training data representations without a supervised signal and has become an important paradigm in machine learning. Most SSL methods employ the cosine similarity between embedding vectors and hence effectively embed data on a hypersphere. While this seemingly implies that embedding norms cannot play any role in SSL, a few recent works have suggested that embedding norms have properties related to network convergence and confidence. In this paper, we resolve this apparent contradiction and systematically establish the embedding norm's role in SSL training. Using theoretical analysis, simulations, and experiments, we show that embedding norms (i) govern SSL convergence rates and (ii) encode network confidence, with smaller norms corresponding to unexpected samples. Additionally, we show that manipulating embedding norms can have large effects on convergence speed. Our findings demonstrate that SSL embedding norms are integral to understanding and optimizing network behavior.
Andrew Draganov, Sharvaree Vadgama, Sebastian Damrich, Jan Niklas Böhm, Lucas Maes, Dmitry Kobak, Erik J. Bekkers
ICML6
2025 TRACE: Contrastive learning for multi-trial time series data in neuroscience
abstract
Modern neural recording techniques such as two-photon imaging or Neuropixel probes allow to acquire vast time-series datasets with responses of hundreds or thousands of neurons. Contrastive learning is a powerful self-supervised framework for learning representations of complex datasets. Existing applications for neural time series rely on generic data augmentations and do not exploit the multi-trial data structure inherent in many neural datasets. Here we present TRACE, a new contrastive learning framework that averages across different subsets of trials to generate positive pairs. TRACE allows to directly learn a two-dimensional embedding, combining ideas from contrastive learning and neighbor embeddings. We show that TRACE outperforms other methods, resolving fine response differences in simulated data. Further, using in vivo recordings, we show that the representations learned by TRACE capture both biologically relevant continuous variation, cell-type-related cluster structure, and can assist data quality control.
Lisa Schmors, Dominic Gonschorek, Jan Niklas Böhm, Yongrong Qiu, Na Zhou, Dmitry Kobak, Andreas S. Tolias, Fabian H. Sinz, Jacob Reimer, Katrin Franke, Sebastian Damrich, Philipp Berens
NeurIPS6
2024 Scaling Down Deep Learning with MNIST-1D
abstract
Although deep learning models have taken on commercial and political relevance, key aspects of their training and operation remain poorly understood. This has sparked interest in science of deep learning projects, many of which require large amounts of time, money, and electricity. But how much of this research really needs to occur at scale? In this paper, we introduce MNIST-1D: a minimalist, procedurally generated, low-memory, and low-compute alternative to classic deep learning benchmarks. Although the dimensionality of MNIST-1D is only 40 and its default training set size only 4000, MNIST-1D can be used to study inductive biases of different deep architectures, find lottery tickets, observe deep double descent, metalearn an activation function, and demonstrate guillotine regularization in self-supervised learning. All these experiments can be conducted on a GPU or often even on a CPU within minutes, allowing for fast prototyping, educational use cases, and cutting-edge research on a low budget.
Sam Greydanus, Dmitry Kobak
ICML2
2024 Persistent Homology for High-dimensional Data Based on Spectral Methods
abstract
Persistent homology is a popular computational tool for analyzing the topology of point clouds, such as the presence of loops or voids. However, many real-world datasets with low intrinsic dimensionality reside in an ambient space of much higher dimensionality. We show that in this case traditional persistent homology becomes very sensitive to noise and fails to detect the correct topology. The same holds true for existing refinements of persistent homology. As a remedy, we find that spectral distances on the k-nearest-neighbor graph of the data, such as diffusion distance and effective resistance, allow to detect the correct topology even in the presence of high-dimensional noise. Moreover, we derive a novel closed-form formula for effective resistance, and describe its relation to diffusion distances. Finally, we apply these methods to high-dimensional single-cell RNA-sequencing data and show that spectral distances allow robust detection of cell cycle loops.
Sebastian Damrich, Philipp Berens, Dmitry Kobak
NeurIPS3
2024 The art of seeing the elephant in the room: 2D embeddings of single-cell data do make sense
abstract
A recent paper claimed that t-SNE and UMAP embeddings of single-cell datasets are "specious" and fail to capture true biological structure. The authors argued that such embeddings are as arbitrary and as misleading as forcing the data into an elephant shape. Here we show that this conclusion was based on inadequate and limited metrics of embedding quality. More appropriate metrics quantifying neighborhood and class preservation reveal the elephant in the room: while t-SNE and UMAP embeddings of single-cell data do not preserve high-dimensional distances, they can nevertheless provide biologically relevant information.
Jan Lause, Philipp Berens, Dmitry Kobak
PLoS Comput. Biol.3
2023 Unsupervised visualization of image datasets using contrastive learning
Jan Niklas Böhm, Philipp Berens, Dmitry Kobak
ICLR3
2023 From $t$-SNE to UMAP with contrastive learning
Sebastian Damrich, Jan Niklas Böhm, Fred A. Hamprecht, Dmitry Kobak
ICLR4
2022 t-SNE Highlights Phylogenetic and Temporal Patterns of SARS-CoV-2 Spike and Nucleocapsid Protein Evolution
Gaik Tamazian, Andrey B. Komissarov, Dmitry Kobak, Dmitry Polyakov, Evgeny Andronov, Sergei Nechaev, Sergey Kryzhevich, Yuri Porozov, Eugene Stepanov
ISBRA3
2022 Wasserstein t-SNE
abstract
Abstract Scientific datasets often have hierarchical structure: for example, in surveys, individual participants (samples) might be grouped at a higher level (units) such as their geographical region. In these settings, the interest is often in exploring the structure on the unit level rather than on the sample level. Units can be compared based on the distance between their means, however this ignores the within-unit distribution of samples. Here we develop an approach for exploratory analysis of hierarchical datasets using the Wasserstein distance metric that takes into account the shapes of within-unit distributions. We use t-SNE to construct 2D embeddings of the units, based on the matrix of pairwise Wasserstein distances between them. The distance matrix can be efficiently computed by approximating each unit with a Gaussian distribution, but we also provide a scalable method to compute exact Wasserstein distances. We use synthetic data to demonstrate the effectiveness of our Wassersteint-SNE, and apply it to data from the 2017 German parliamentary election, considering polling stations as samples and voting districts as units. The resulting embedding uncovers meaningful structure in the data.
Fynn Bachmann, Philipp Hennig, Dmitry Kobak
ECML/PKDD (1)3
2022 Attraction-Repulsion Spectrum in Neighbor Embeddings
abstract
Neighbor embeddings are a family of methods for visualizing complex high-dimensional data sets using kNN graphs. To find the low-dimensional embedding, these algorithms combine an attractive force between neighboring pairs of points with a repulsive force between all points. One of the most popular examples of such algorithms is t-SNE. Here we empirically show that changing the balance between the attractive and the repulsive forces in t-SNE using the exaggeration parameter yields a spectrum of embeddings, which is characterized by a simple trade-off: stronger attraction can better represent continuous manifold structures, while stronger repulsion can better represent discrete cluster structures and yields higher kNN recall. We find that UMAP embeddings correspond to t-SNE with increased attraction; mathematical analysis shows that this is because the negative sampling optimization strategy employed by UMAP strongly lowers the effective repulsion. Likewise, ForceAtlas2, commonly used for visualizing developmental single-cell transcriptomic data, yields embeddings corresponding to t-SNE with the attraction increased even more. At the extreme of this spectrum lie Laplacian eigenmaps. Our results demonstrate that many prominent neighbor embedding algorithms can be placed onto the attraction-repulsion spectrum, and highlight the inherent trade-offs between them.
Jan Niklas Böhm, Philipp Berens, Dmitry Kobak
J. Mach. Learn. Res.3
2021 Interpretable Gender Classification from Retinal Fundus Images Using BagNets
Indu Ilanchezian, Dmitry Kobak, Hanna Faber, Focke Ziemssen, Philipp Berens, Murat Seçkin Ayhan
MICCAI (3)2
2020 The Optimal Ridge Penalty for Real-world High-dimensional Data Can Be Zero or Negative due to the Implicit Ridge Regularization
abstract
A conventional wisdom in statistical learning is that large models require strong regularization to prevent overfitting. Here we show that this rule can be violated by linear regression in the underdetermined $n\ll p$ situation under realistic conditions. Using simulations and real-life high-dimensional datasets, we demonstrate that an explicit positive ridge penalty can fail to provide any improvement over the minimum-norm least squares estimator. Moreover, the optimal value of ridge penalty in this situation can be negative. This happens when the high-variance directions in the predictor space can predict the response variable, which is often the case in the real-world high-dimensional data. In this regime, low-variance directions provide an implicit ridge regularization and can make any further positive ridge penalty detrimental. We prove that augmenting any linear model with random covariates and using minimum-norm estimator is asymptotically equivalent to adding the ridge penalty. We use a spiked covariance model as an analytically tractable example and prove that the optimal ridge penalty in this case is negative when $n\ll p$.
Dmitry Kobak, Jonathan Lomond, Benoit Sanchez
J. Mach. Learn. Res.1
2019 Heavy-Tailed Kernels Reveal a Finer Cluster Structure in t-SNE Visualisations
abstract
Abstract T-distributed stochastic neighbour embedding (t-SNE) is a widely used data visualisation technique. It differs from its predecessor SNE by the low-dimensional similarity kernel: the Gaussian kernel was replaced by the heavy-tailed Cauchy kernel, solving the ‘crowding problem’ of SNE. Here, we develop an efficient implementation of t-SNE for a t-distribution kernel with an arbitrary degree of freedom $$\nu $$ , with $$\nu \rightarrow \infty $$ corresponding to SNE and $$\nu =1$$ corresponding to the standard t-SNE. Using theoretical analysis and toy examples, we show that $$\nu <1$$ can further reduce the crowding problem and reveal finer cluster structure that is invisible in standard t-SNE. We further demonstrate the striking effect of heavier-tailed kernels on large real-life data sets such as MNIST, single-cell RNA-sequencing data, and the HathiTrust library. We use domain knowledge to confirm that the revealed clusters are meaningful. Overall, we argue that modifying the tail heaviness of the t-SNE kernel can yield additional insight into the cluster structure of the data.
Dmitry Kobak, George C. Linderman, Stefan Steinerberger, Yuval Kluger, Philipp Berens
ECML/PKDD (1)1