Constantin Ahlmann-Eltze

dblp:235/0344 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
1since 2021 · last 2021
0000-0002-3762-068XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
count data modeling
0.512021
glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count data · Bioinform. 2021
Bioinformatics and computational biology › gene expression analysis
differential expression analysis
0.512021
glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count data · Bioinform. 2021
Bioinformatics and computational biology
single-cell analysis
0.512021
glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count data · Bioinform. 2021
Bioinformatics and computational biology › single-cell analysis
single-cell RNA sequencing
0.512021
glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count data · Bioinform. 2021

Methods — techniques the papers use, named apart from their topics

out-of-core computation · 0.5generalized linear model · 0.5gamma-poisson distribution · 0.5
YearPublicationVenuePosition
2021 glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count data
abstract
MOTIVATION: The Gamma-Poisson distribution is a theoretically and empirically motivated model for the sampling variability of single cell RNA-sequencing counts and an essential building block for analysis approaches including differential expression analysis, principal component analysis and factor analysis. Existing implementations for inferring its parameters from data often struggle with the size of single cell datasets, which can comprise millions of cells; at the same time, they do not take full advantage of the fact that zero and other small numbers are frequent in the data. These limitations have hampered uptake of the model, leaving room for statistically inferior approaches such as logarithm(-like) transformation. RESULTS: We present a new R package for fitting the Gamma-Poisson distribution to data with the characteristics of modern single cell datasets more quickly and more accurately than existing methods. The software can work with data on disk without having to load them into RAM simultaneously. AVAILABILITYAND IMPLEMENTATION: The package glmGamPoi is available from Bioconductor for Windows, macOS and Linux, and source code is available on github.com/const-ae/glmGamPoi under a GPL-3 license. The scripts to reproduce the results of this paper are available on github.com/const-ae/glmGamPoi-Paper. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Constantin Ahlmann-Eltze, Wolfgang Huber
Bioinform.1
2018 MixDir: Scalable Bayesian Clustering for High-Dimensional Categorical Data
abstract
Multivariate analysis of high-dimensional datasets with multiple categorical variables (e.g. surveys, questionnaires) is a challenging task but can reveal patterns of responses that are masked from univariate analyses. In this paper we propose a novel variational inference algorithm to cluster high-dimensional categorical observations into latent classes. Variational inference is an approximate Bayesian inference algorithm, which combines fast optimization methods with the ability to propagate the uncertainty to the clustering (soft clustering). The model is robust to misspecification of the number of latent classes and can infer a reasonable number from the data. We assess the performance on synthetic and real world data and show that our algorithm has similar performance to the best other tested method if the correct number of classes is known and outperforms the other methods if it the number of classes needs to be inferred. An R-package implementing our algorithm is available at the Comprehensive R Archive Network.
Constantin Ahlmann-Eltze, Christopher Yau
DSAA1