EDBT 2026 Demo / reviewers in the wild / expert
Joshua D. Welch
dblp:123/6991
· DBLP profile ↗
11ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5869-2391ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bayesian inference of RNA velocity incorporating timepoints, lineage bifurcations, and count dataabstractExperimental approaches for measuring single-cell gene expression can observe each cell at only one time point, requiring computational approaches for reconstructing the dynamics of gene expression during cell fate transitions. RNA velocity is a promising computational approach for this problem, but existing inference methods fail to capture key aspects of real data, limiting their utility. To address these limitations, we developed VeloVAE, a Bayesian model for RNA velocity inference. VeloVAE uses variational Bayesian inference to estimate the posterior distribution of latent time, latent cell state, and kinetic rate parameters for each cell. Our approach can incorporate prior distributions on rate parameters and time points; model lineage bifurcations using branching differential equations; and directly model discrete count data. We show that VeloVAE significantly outperforms previous approaches in terms of data fit, accuracy of inferred differentiation directions, and transcription rate estimation. These improvements allow VeloVAE to accurately model gene expression dynamics in complex biological systems, including hematopoiesis, induced pluripotent stem cell reprogramming, the developing mouse brain, and the entire mouse embryo. We find that the latent time automatically inferred using all cells can even outperform pseudotime inferred using manually chosen cell subsets and root cells. Our work provides important new tools for modeling sequential changes in gene expression from single-cell expression data. Yichen Gu, David T. Blaauw, Joshua D. Welch |
PLoS Comput. Biol. | 4 |
| 2025 | Integrating single-cell multimodal epigenomic data using 1D convolutional neural networksabstractMOTIVATION: Recent experimental developments enable single-cell multimodal epigenomic profiling, which measures multiple histone modifications and chromatin accessibility within the same cell. Such parallel measurements provide exciting new opportunities to investigate how epigenomic modalities vary together across cell types and states. A pivotal step in using these types of data is integrating the epigenomic modalities to learn a unified representation of each cell, but existing approaches are not designed to model the unique nature of this data type. Our key insight is to model single-cell multimodal epigenome data as a multichannel sequential signal. RESULTS: We developed ConvNet-VAEs, a novel framework that uses one-dimensional (1D) convolutional variational autoencoders (VAEs) for single-cell multimodal epigenomic data integration. We evaluated ConvNet-VAEs on nano-CUT&Tag and single-cell nanobody-tethered transposition followed by sequencing data generated from juvenile mouse brain and human bone marrow. We found that ConvNet-VAEs can perform dimension reduction and batch correction better than previous architectures while using significantly fewer parameters. Furthermore, the performance gap between convolutional and fully connected architectures increases with the number of modalities, and deeper convolutional architectures can increase the performance, while the performance degrades for deeper fully connected architectures. Our results indicate that convolutional autoencoders are a promising method for integrating current and future single-cell multimodal epigenomic datasets. AVAILABILITY AND IMPLEMENTATION: The source code of VAE models and a demo in Jupyter notebook are available at https://github.com/welch-lab/ConvNetVAE. Joshua D. Welch |
Bioinform. | 2 |
| 2025 | CytoSimplex: visualizing single-cell fates and transitions on a simplexabstractSUMMARY: Cells differentiate to their final fates along unique trajectories, often involving multi-potent progenitors that can produce multiple terminally differentiated cell types. Recent developments in single-cell transcriptomic and epigenomic measurement provide tremendous opportunities for mapping these trajectories. The visualization of single-cell data often relies on dimension reduction methods such as UMAP to simplify high-dimensional single-cell data down to an understandable 2D form. However, these dimension reduction methods are not constructed to allow direct interpretation of the reduced dimensions in terms of cell differentiation. To address these limitations, we developed a new approach that places each cell from a single-cell dataset within a simplex whose vertices correspond to terminally differentiated cell types. Our approach can quantify and visualize current cell fate commitment and future cell potential. We developed CytoSimplex, a standalone open-source package implemented in R and Python that provides simple and intuitive visualizations of cell differentiation in 2D ternary and 3D quaternary plots. We believe that CytoSimplex can help researchers gain a better understanding of cell type transitions in specific tissues and characterize developmental processes. AVAILABILITY AND IMPLEMENTATION: The R version of CytoSimplex is available on Github at https://github.com/welch-lab/CytoSimplex. The Python version of CytoSimplex is available on Github at https://github.com/welch-lab/pyCytoSimplex. Yichen Gu, Noriaki Ono, Joshua D. Welch |
Bioinform. | 6 |
| 2024 | Mapping Cell Fate Transition in Space and Time
Yichen Gu, Joshua D. Welch |
RECOMB | 4 |
| 2022 | Variational Mixtures of ODEs for Inferring Cellular Gene Expression DynamicsabstractA key problem in computational biology is discovering the gene expression changes that regulate cell fate transitions, in which one cell type turns into another. However, each individual cell cannot be tracked longitudinally, and cells at the same point in real time may be at different stages of the transition process. This can be viewed as a problem of learning the behavior of a dynamical system from observations whose times are unknown. Additionally, a single progenitor cell type often bifurcates into multiple child cell types, further complicating the problem of modeling the dynamics. To address this problem, we developed an approach called variational mixtures of ordinary differential equations. By using a simple family of ODEs informed by the biochemistry of gene expression to constrain the likelihood of a deep generative model, we can simultaneously infer the latent time and latent state of each cell and predict its future gene expression state. The model can be interpreted as a mixture of ODEs whose parameters vary continuously across a latent space of cell states. Our approach dramatically improves data fit, latent time inference, and future cell state estimation of single-cell gene expression data compared to previous approaches. Yichen Gu, David T. Blaauw, Joshua D. Welch |
ICML | 3 |
| 2022 | Single-Cell Multi-omic Velocity Infers Dynamic and Decoupled Gene Regulation
Maria Virgilio, Kathleen L. Collins, Joshua D. Welch |
RECOMB | 4 |
| 2022 | PyLiger: scalable single-cell multi-omic data integration in PythonabstractMOTIVATION: LIGER (Linked Inference of Genomic Experimental Relationships) is a widely used R package for single-cell multi-omic data integration. However, many users prefer to analyze their single-cell datasets in Python, which offers an attractive syntax and highly optimized scientific computing libraries for increased efficiency. RESULTS: We developed PyLiger, a Python package for integrating single-cell multi-omic datasets. PyLiger offers faster performance than the previous R implementation (2-5× speedup), interoperability with AnnData format, flexible on-disk or in-memory analysis capability and new functionality for gene ontology enrichment analysis. The on-disk capability enables analysis of arbitrarily large single-cell datasets using fixed memory. AVAILABILITY AND IMPLEMENTATION: PyLiger is available on Github at https://github.com/welch-lab/pyliger and on the Python Package Index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joshua D. Welch |
Bioinform. | 2 |
| 2020 | Iterative Refinement of Cellular Identity from Single-Cell Data Using Online Learning
Joshua D. Welch |
RECOMB | 2 |
| 2017 | E Pluribus Unum: United States of Single Cells
Joshua D. Welch, Alexander J. Hartemink, Jan F. Prins |
RECOMB | 1 |
| 2016 | SLICER: Inferring Branched, Nonlinear Cellular Trajectories from Single Cell RNA-seq Data
Joshua D. Welch, Ziqing Liu, Paul Lerou, Jeremy E. Purvis, Alexander J. Hartemink, Jan F. Prins |
RECOMB | 1 |
| 2010 | WordSeeker: concurrent bioinformatics software for discovering genome-wide patterns and word-based genomic signaturesabstractBACKGROUND: An important focus of genomic science is the discovery and characterization of all functional elements within genomes. In silico methods are used in genome studies to discover putative regulatory genomic elements (called words or motifs). Although a number of methods have been developed for motif discovery, most of them lack the scalability needed to analyze large genomic data sets. METHODS: This manuscript presents WordSeeker, an enumerative motif discovery toolkit that utilizes multi-core and distributed computational platforms to enable scalable analysis of genomic data. A controller task coordinates activities of worker nodes, each of which (1) enumerates a subset of the DNA word space and (2) scores words with a distributed Markov chain model. RESULTS: A comprehensive suite of performance tests was conducted to demonstrate the performance, speedup and efficiency of WordSeeker. The scalability of the toolkit enabled the analysis of the entire genome of Arabidopsis thaliana; the results of the analysis were integrated into The Arabidopsis Gene Regulatory Information Server (AGRIS). A public version of WordSeeker was deployed on the Glenn cluster at the Ohio Supercomputer Center. CONCLUSION: WordSeeker effectively utilizes concurrent computing platforms to enable the identification of putative functional elements in genomic data sets. This capability facilitates the analysis of the large quantity of sequenced genomic data. Jens Lichtenberg, Kyle Kurz, Rami Al-ouran, Lev Neiman, Lee J. Nau, Joshua D. Welch, Edwin Jacox, Thomas Bitterman, Klaus H. Ecker, Laura Elnitski, Frank Drews, Stephen Lee, Lonnie R. Welch |
BMC Bioinform. | 7 |