EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyang Chen 0007
dblp:98/8121-7
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-7356-8129ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-modality representation and multi-sample integration of spatially resolved omics dataabstractSpatially resolved sequencing technologies have revolutionized our understanding of biological regulatory processes within tissue microenvironments by simultaneously capturing the states of genomic regions, genes, and proteins alongside the spatial organization of cells. However, inherent heterogeneity across modalities and samples poses substantial challenges for the integrative analysis of spatial omics data, underscoring the urgent need for advanced computational methods. In this study, we propose PRESENT, a contrastive learning-based integrative framework for cross-modality representation of spatial multi-omics data. PRESENT employs omics-specific encoders consisting of graph attention networks and Bayesian neural networks coupled with distribution-aware decoders to model distinct modalities, and an inter-omics alignment module for multi-omics integration. By effectively incorporating spatial dependencies with multi-omics information across diverse species and technologies, PRESENT facilitates the accurate identification of spatial domains and the elucidation of underlying regulatory mechanisms. Furthermore, PRESENT can be extended to multi-sample integration via a two-stage training workflow, which incorporates inter-batch alignment loss, intra-batch preserving loss, batch-adversarial learning, and cyclic graph refinement strategies to eliminate batch effects while retaining biological signals. Extensive experiments on tissue samples across different anatomical regions and developmental stages demonstrate that PRESENT enables the characterization of hierarchical tissue structures from a spatiotemporal perspective. Zhen Li 0056, Xuejian Cui, Xiaoyang Chen 0007, Zijing Gao, Yuyao Liu, Yan Pan 0014, Shengquan Chen, Hairong Lv, Lei Zhai, Rui Jiang 0001 |
Briefings Bioinform. | 3 |
| 2024 | Cofea: correlation-based feature selection for single-cell chromatin accessibility dataabstractSingle-cell chromatin accessibility sequencing (scCAS) technologies have enabled characterizing the epigenomic heterogeneity of individual cells. However, the identification of features of scCAS data that are relevant to underlying biological processes remains a significant gap. Here, we introduce a novel method Cofea, to fill this gap. Through comprehensive experiments on 5 simulated and 54 real datasets, Cofea demonstrates its superiority in capturing cellular heterogeneity and facilitating downstream analysis. Applying this method to identification of cell type-specific peaks and candidate enhancers, as well as pathway enrichment analysis and partitioned heritability analysis, we illustrate the potential of Cofea to uncover functional biological process. Xiaoyang Chen 0007, Shuang Song 0006, Lin Hou 0003, Shengquan Chen, Rui Jiang 0001 |
Briefings Bioinform. | 2 |
| 2024 | scPRAM accurately predicts single-cell gene expression perturbation response based on attention mechanismabstractMOTIVATION: With the rapid advancement of single-cell sequencing technology, it becomes gradually possible to delve into the cellular responses to various external perturbations at the gene expression level. However, obtaining perturbed samples in certain scenarios may be considerably challenging, and the substantial costs associated with sequencing also curtail the feasibility of large-scale experimentation. A repertoire of methodologies has been employed for forecasting perturbative responses in single-cell gene expression. However, existing methods primarily focus on the average response of a specific cell type to perturbation, overlooking the single-cell specificity of perturbation responses and a more comprehensive prediction of the entire perturbation response distribution. RESULTS: Here, we present scPRAM, a method for predicting perturbation responses in single-cell gene expression based on attention mechanisms. Leveraging variational autoencoders and optimal transport, scPRAM aligns cell states before and after perturbation, followed by accurate prediction of gene expression responses to perturbations for unseen cell types through attention mechanisms. Experiments on multiple real perturbation datasets involving drug treatments and bacterial infections demonstrate that scPRAM attains heightened accuracy in perturbation prediction across cell types, species, and individuals, surpassing existing methodologies. Furthermore, scPRAM demonstrates outstanding capability in identifying differentially expressed genes under perturbation, capturing heterogeneity in perturbation responses across species, and maintaining stability in the presence of data noise and sample size variations. AVAILABILITY AND IMPLEMENTATION: https://github.com/jiang-q19/scPRAM and https://doi.org/10.5281/zenodo.10935038. Qun Jiang, Shengquan Chen, Xiaoyang Chen 0007, Rui Jiang 0001 |
Bioinform. | 3 |
| 2024 | EpiCarousel: memory- and time-efficient identification of metacells for atlas-level single-cell chromatin accessibility dataabstractSUMMARY: Recent technical advancements in single-cell chromatin accessibility sequencing (scCAS) have brought new insights to the characterization of epigenetic heterogeneity. As single-cell genomics experiments scale up to hundreds of thousands of cells, the demand for computational resources for downstream analysis grows intractably large and exceeds the capabilities of most researchers. Here, we propose EpiCarousel, a tailored Python package based on lazy loading, parallel processing, and community detection for memory- and time-efficient identification of metacells, i.e. the emergence of homogenous cells, in large-scale scCAS data. Through comprehensive experiments on five datasets of various protocols, sample sizes, dimensions, number of cell types, and degrees of cell-type imbalance, EpiCarousel outperformed baseline methods in systematic evaluation of memory usage, computational time, and multiple downstream analyses including cell type identification. Moreover, EpiCarousel executes preprocessing and downstream cell clustering on the atlas-level dataset with 707 043 cells and 1 154 611 peaks within 2 h consuming <75 GB of RAM and provides superior performance for characterizing cell heterogeneity than state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The EpiCarousel software is well-documented and freely available at https://github.com/biox-nku/epicarousel. It can be seamlessly interoperated with extensive scCAS analysis toolkits. Xiaoyang Chen 0007, Songming Tang, Shengquan Chen |
Bioinform. | 5 |
| 2023 | simCAS: an embedding-based method for simulating single-cell chromatin accessibility sequencing dataabstractMOTIVATION: Single-cell chromatin accessibility sequencing (scCAS) technology provides an epigenomic perspective to characterize gene regulatory mechanisms at single-cell resolution. With an increasing number of computational methods proposed for analyzing scCAS data, a powerful simulation framework is desirable for evaluation and validation of these methods. However, existing simulators generate synthetic data by sampling reads from real data or mimicking existing cell states, which is inadequate to provide credible ground-truth labels for method evaluation. RESULTS: We present simCAS, an embedding-based simulator, for generating high-fidelity scCAS data from both cell- and peak-wise embeddings. We demonstrate simCAS outperforms existing simulators in resembling real data and show that simCAS can generate cells of different states with user-defined cell populations and differentiation trajectories. Additionally, simCAS can simulate data from different batches and encode user-specified interactions of chromatin regions in the synthetic data, which provides ground-truth labels more than cell states. We systematically demonstrate that simCAS facilitates the benchmarking of four core tasks in downstream analysis: cell clustering, trajectory inference, data integration, and cis-regulatory interaction inference. We anticipate simCAS will be a reliable and flexible simulator for evaluating the ongoing computational methods applied on scCAS data. AVAILABILITY AND IMPLEMENTATION: simCAS is freely available at https://github.com/Chen-Li-17/simCAS. Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001, Xuegong Zhang |
Bioinform. | 2 |
| 2021 | stPlus: a reference-based method for the accurate enhancement of spatial transcriptomicsabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. RESULTS: We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial gene expression via a weighted k-nearest-neighbor. stPlus outperforms baseline methods with higher gene-wise and cell-wise Spearman correlation coefficients. We also introduce a clustering-based approach to assess the enhancement performance systematically. Using the data enhanced by stPlus, cell populations can be better identified than using the measured data. The predicted expression of genes unique to scRNA-seq data can also well characterize spatial cell heterogeneity. Besides, stPlus is robust and scalable to datasets of diverse gene detection sensitivity levels, sample sizes and number of spatially measured genes. We anticipate stPlus will facilitate the analysis of spatial transcriptomics. AVAILABILITY AND IMPLEMENTATION: stPlus with detailed documents is freely accessible at http://health.tsinghua.edu.cn/software/stPlus/ and the source code is openly available on https://github.com/xy-chen16/stPlus. Shengquan Chen, Boheng Zhang, Xiaoyang Chen 0007, Xuegong Zhang, Rui Jiang 0001 |
Bioinform. | 3 |
| 2020 | EnClaSC: a novel ensemble approach for accurate and robust cell-type classification of single-cell transcriptomesabstractBACKGROUND: In recent years, the rapid development of single-cell RNA-sequencing (scRNA-seq) techniques enables the quantitative characterization of cell types at a single-cell resolution. With the explosive growth of the number of cells profiled in individual scRNA-seq experiments, there is a demand for novel computational methods for classifying newly-generated scRNA-seq data onto annotated labels. Although several methods have recently been proposed for the cell-type classification of single-cell transcriptomic data, such limitations as inadequate accuracy, inferior robustness, and low stability greatly limit their wide applications. RESULTS: We propose a novel ensemble approach, named EnClaSC, for accurate and robust cell-type classification of single-cell transcriptomic data. Through comprehensive validation experiments, we demonstrate that EnClaSC can not only be applied to the self-projection within a specific dataset and the cell-type classification across different datasets, but also scale up well to various data dimensionality and different data sparsity. We further illustrate the ability of EnClaSC to effectively make cross-species classification, which may shed light on the studies in correlation of different species. EnClaSC is freely available at https://github.com/xy-chen16/EnClaSC . CONCLUSIONS: EnClaSC enables highly accurate and robust cell-type classification of single-cell transcriptomic data via an ensemble learning method. We expect to see wide applications of our method to not only transcriptome studies, but also the classification of more general data. Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001 |
BMC Bioinform. | 1 |