VLDB 2026 Research / reviewers in the wild / expert
Allon M. Klein
dblp:217/5464
· DBLP profile ↗
2ranked-venue papers
0as first author
1since 2021 · last 2024
0000-0001-8913-7879ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
low-dimensional embedding |
0.8 | 1 | 2024 | Approximating mutual information of high-dimensional variables using learned representations · NeurIPS 2024 |
Information theory › information measures › mutual information
mutual information estimation |
0.8 | 1 | 2024 | Approximating mutual information of high-dimensional variables using learned representations · NeurIPS 2024 |
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics |
0.3 | 1 | 2018 | SPRING: a kinetic interface for visualizing high dimensional single-cell expression data · Bioinform. 2018 |
Bioinformatics and computational biology › single-cell analysis
single-cell trajectory analysis |
0.1 | 1 | 2018 | SPRING: a kinetic interface for visualizing high dimensional single-cell expression data · Bioinform. 2018 |
Methods — techniques the papers use, named apart from their topics
nonparametric MI estimation · 1.5protein language models · 0.8protein language model · 0.8k-nearest-neighbor graph · 0.3force-directed layout · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Approximating mutual information of high-dimensional variables using learned representationsabstractMutual information (MI) is a general measure of statistical dependence with widespread application across the sciences. However, estimating MI between multi-dimensional variables is challenging because the number of samples necessary to converge to an accurate estimate scales unfavorably with dimensionality. In practice, existing techniques can reliably estimate MI in up to tens of dimensions, but fail in higher dimensions, where sufficient sample sizes are infeasible. Here, we explore the idea that underlying low-dimensional structure in high-dimensional data can be exploited to faithfully approximate MI in high-dimensional settings with realistic sample sizes. We develop a method that we call latent MI (LMI) approximation, which applies a nonparametric MI estimator to low-dimensional representations learned by a simple, theoretically-motivated model architecture. Using several benchmarks, we show that unlike existing techniques, LMI can approximate MI well for variables with $> 10^3$ dimensions if their dependence structure is captured by low-dimensional representations. Finally, we showcase LMI on two open problems in biology. First, we approximate MI between protein language model (pLM) representations of interacting proteins, and find that pLMs encode non-trivial information about protein-protein interactions. Second, we quantify cell fate information contained in single-cell RNA-seq (scRNA-seq) measurements of hematopoietic stem cells, and find a sharp transition during neutrophil differentiation when fate information captured by scRNA-seq increases dramatically. An implementation of LMI is available at *latentmi.readthedocs.io.* Gokul Gowri, Xiao-Kang Lun, Allon M. Klein |
NeurIPS | 3 |
| 2018 | SPRING: a kinetic interface for visualizing high dimensional single-cell expression dataabstractMotivation: Single-cell gene expression profiling technologies can map the cell states in a tissue or organism. As these technologies become more common, there is a need for computational tools to explore the data they produce. In particular, visualizing continuous gene expression topologies can be improved, since current tools tend to fragment gene expression continua or capture only limited features of complex population topologies. Results: Force-directed layouts of k-nearest-neighbor graphs can visualize continuous gene expression topologies in a manner that preserves high-dimensional relationships and captures complex population topologies. We describe SPRING, a pipeline for data filtering, normalization and visualization using force-directed layouts and show that it reveals more detailed biological relationships than existing approaches when applied to branching gene expression trajectories from hematopoietic progenitor cells and cells of the upper airway epithelium. Visualizations from SPRING are also more reproducible than those of stochastic visualization methods such as tSNE, a state-of-the-art tool. We provide SPRING as an interactive web-tool with an easy to use GUI. Availability and implementation: https://kleintools.hms.harvard.edu/tools/spring.html, https://github.com/AllonKleinLab/SPRING/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Caleb Weinreb, Samuel L. Wolock, Allon M. Klein |
Bioinform. | 3 |