Allon M. Klein

dblp:217/5464 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2024
0000-0001-8913-7879ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Representation and self-supervised learning · 100%
Theoretical computer science
1 paper
Information theory · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
low-dimensional embedding
0.812024
Approximating mutual information of high-dimensional variables using learned representations · NeurIPS 2024
Information theory › information measures › mutual information
mutual information estimation
0.812024
Approximating mutual information of high-dimensional variables using learned representations · NeurIPS 2024
Bioinformatics and computational biology › single-cell analysis
single-cell transcriptomics
0.312018
SPRING: a kinetic interface for visualizing high dimensional single-cell expression data · Bioinform. 2018
Bioinformatics and computational biology › single-cell analysis
single-cell trajectory analysis
0.112018
SPRING: a kinetic interface for visualizing high dimensional single-cell expression data · Bioinform. 2018

Methods — techniques the papers use, named apart from their topics

nonparametric MI estimation · 1.5protein language models · 0.8protein language model · 0.8k-nearest-neighbor graph · 0.3force-directed layout · 0.3
YearPublicationVenuePosition
2024 Approximating mutual information of high-dimensional variables using learned representations
abstract
Mutual information (MI) is a general measure of statistical dependence with widespread application across the sciences. However, estimating MI between multi-dimensional variables is challenging because the number of samples necessary to converge to an accurate estimate scales unfavorably with dimensionality. In practice, existing techniques can reliably estimate MI in up to tens of dimensions, but fail in higher dimensions, where sufficient sample sizes are infeasible. Here, we explore the idea that underlying low-dimensional structure in high-dimensional data can be exploited to faithfully approximate MI in high-dimensional settings with realistic sample sizes. We develop a method that we call latent MI (LMI) approximation, which applies a nonparametric MI estimator to low-dimensional representations learned by a simple, theoretically-motivated model architecture. Using several benchmarks, we show that unlike existing techniques, LMI can approximate MI well for variables with $> 10^3$ dimensions if their dependence structure is captured by low-dimensional representations. Finally, we showcase LMI on two open problems in biology. First, we approximate MI between protein language model (pLM) representations of interacting proteins, and find that pLMs encode non-trivial information about protein-protein interactions. Second, we quantify cell fate information contained in single-cell RNA-seq (scRNA-seq) measurements of hematopoietic stem cells, and find a sharp transition during neutrophil differentiation when fate information captured by scRNA-seq increases dramatically. An implementation of LMI is available at *latentmi.readthedocs.io.*
Gokul Gowri, Xiao-Kang Lun, Allon M. Klein
NeurIPS3
2018 SPRING: a kinetic interface for visualizing high dimensional single-cell expression data
abstract
Motivation: Single-cell gene expression profiling technologies can map the cell states in a tissue or organism. As these technologies become more common, there is a need for computational tools to explore the data they produce. In particular, visualizing continuous gene expression topologies can be improved, since current tools tend to fragment gene expression continua or capture only limited features of complex population topologies. Results: Force-directed layouts of k-nearest-neighbor graphs can visualize continuous gene expression topologies in a manner that preserves high-dimensional relationships and captures complex population topologies. We describe SPRING, a pipeline for data filtering, normalization and visualization using force-directed layouts and show that it reveals more detailed biological relationships than existing approaches when applied to branching gene expression trajectories from hematopoietic progenitor cells and cells of the upper airway epithelium. Visualizations from SPRING are also more reproducible than those of stochastic visualization methods such as tSNE, a state-of-the-art tool. We provide SPRING as an interactive web-tool with an easy to use GUI. Availability and implementation: https://kleintools.hms.harvard.edu/tools/spring.html, https://github.com/AllonKleinLab/SPRING/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Caleb Weinreb, Samuel L. Wolock, Allon M. Klein
Bioinform.3