VLDB 2026 Research / reviewers in the wild / expert
Shengquan Chen
dblp:210/3652
· DBLP profile ↗
18ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-3503-9306ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Yi Liu 0069, Yuekang Li, Ling Shi 0002, Gelei Deng, Shengquan Chen, Kailong Wang 0001 |
ICPR (2) | 6 |
| 2026 | Cross-modality representation and multi-sample integration of spatially resolved omics dataabstractSpatially resolved sequencing technologies have revolutionized our understanding of biological regulatory processes within tissue microenvironments by simultaneously capturing the states of genomic regions, genes, and proteins alongside the spatial organization of cells. However, inherent heterogeneity across modalities and samples poses substantial challenges for the integrative analysis of spatial omics data, underscoring the urgent need for advanced computational methods. In this study, we propose PRESENT, a contrastive learning-based integrative framework for cross-modality representation of spatial multi-omics data. PRESENT employs omics-specific encoders consisting of graph attention networks and Bayesian neural networks coupled with distribution-aware decoders to model distinct modalities, and an inter-omics alignment module for multi-omics integration. By effectively incorporating spatial dependencies with multi-omics information across diverse species and technologies, PRESENT facilitates the accurate identification of spatial domains and the elucidation of underlying regulatory mechanisms. Furthermore, PRESENT can be extended to multi-sample integration via a two-stage training workflow, which incorporates inter-batch alignment loss, intra-batch preserving loss, batch-adversarial learning, and cyclic graph refinement strategies to eliminate batch effects while retaining biological signals. Extensive experiments on tissue samples across different anatomical regions and developmental stages demonstrate that PRESENT enables the characterization of hierarchical tissue structures from a spatiotemporal perspective. Zhen Li 0056, Xuejian Cui, Xiaoyang Chen 0007, Zijing Gao, Yuyao Liu, Yan Pan 0014, Shengquan Chen, Hairong Lv, Lei Zhai, Rui Jiang 0001 |
Briefings Bioinform. | 7 |
| 2025 | BIOTIC: a Bayesian framework to integrate single-cell multi-omics for transcription factor activity inference and improve identity characterization of cellsabstractUnderstanding cell destiny requires unraveling the intricate mechanism of gene regulation, where transcription factors (TFs) play a pivotal role. However, the actual contribution of TFs, that is TF activity, is not only determined by TF expression, but also accessibility of corresponding chromatin regions. Therefore, we introduce BIOTIC, an advanced Bayesian model with a well-established gene regulation structure that harnesses the power of single-cell multi-omics data to model the gene expression process under the control of regulatory elements, thereby defining the regulatory activity of TFs with variational inference. We demonstrated that the TF activity inferred by BIOTIC can serve as a characterization of cell identity, and outperforms baseline methods for the tasks of cell typing, cell development tracking, and batch effect correction. Additionally, BIOTIC trained on multi-omics data can flexibly be applied to the scenario where merely single-cell transcriptome sequencing is available, to infer TF activity and annotate the cell type by mapping the query cell into the reference TF activity space, as an emerging application of cell atlases. The structure of BIOTIC has been determined to be adaptable for the inclusion of additional biological factors, allowing for flexible and more comprehensive gene regulation analysis. BIOTIC introduces a pioneering biological-mechanism-driven framework to infer TF activity and elucidate cell identity states at gene regulatory level, paving the way for a deeper understanding of the complex interplay between TFs and gene expression in living systems. Shengquan Chen, Xiaobing Huang, Ying Wang 0005 |
Briefings Bioinform. | 4 |
| 2025 | Graph neural networks for single-cell omics data: a review of approaches and applicationsabstractRapid advancement of sequencing technologies now allows for the utilization of precise signals at single-cell resolution in various omics studies. However, the massive volume, ultra-high dimensionality, and high sparsity nature of single-cell data have introduced substantial difficulties to traditional computational methods. The intricate non-Euclidean networks of intracellular and intercellular signaling molecules within single-cell datasets, coupled with the complex, multimodal structures arising from multi-omics joint analysis, pose significant challenges to conventional deep learning operations reliant on Euclidean geometries. Graph neural networks (GNNs) have extended deep learning to non-Euclidean data, allowing cells and their features in single-cell datasets to be modeled as nodes within a graph structure. GNNs have been successfully applied across a broad range of tasks in single-cell data analysis. In this survey, we systematically review 107 successful applications of GNNs and their six variants in various single-cell omics tasks. We begin by outlining the fundamental principles of GNNs and their six variants, followed by a systematic review of GNN-based models applied in single-cell epigenomics, transcriptomics, spatial transcriptomics, proteomics, and multi-omics. In each section dedicated to a specific omics type, we have summarized the publicly available single-cell datasets commonly utilized in the articles reviewed in that section, totaling 77 datasets. Finally, we summarize the potential shortcomings of current research and explore directions for future studies. We anticipate that this review will serve as a guiding resource for researchers to deepen the application of GNNs in single-cell omics. Heyang Hua, Shengquan Chen |
Briefings Bioinform. | 3 |
| 2025 | Facilitating single-cell chromatin accessibility research with a user-friendly database
Heyang Hua, Haitian Liang, Shengquan Chen |
Frontiers Comput. Sci. | 4 |
| 2024 | Cofea: correlation-based feature selection for single-cell chromatin accessibility dataabstractSingle-cell chromatin accessibility sequencing (scCAS) technologies have enabled characterizing the epigenomic heterogeneity of individual cells. However, the identification of features of scCAS data that are relevant to underlying biological processes remains a significant gap. Here, we introduce a novel method Cofea, to fill this gap. Through comprehensive experiments on 5 simulated and 54 real datasets, Cofea demonstrates its superiority in capturing cellular heterogeneity and facilitating downstream analysis. Applying this method to identification of cell type-specific peaks and candidate enhancers, as well as pathway enrichment analysis and partitioned heritability analysis, we illustrate the potential of Cofea to uncover functional biological process. Xiaoyang Chen 0007, Shuang Song 0006, Lin Hou 0003, Shengquan Chen, Rui Jiang 0001 |
Briefings Bioinform. | 5 |
| 2024 | scPRAM accurately predicts single-cell gene expression perturbation response based on attention mechanismabstractMOTIVATION: With the rapid advancement of single-cell sequencing technology, it becomes gradually possible to delve into the cellular responses to various external perturbations at the gene expression level. However, obtaining perturbed samples in certain scenarios may be considerably challenging, and the substantial costs associated with sequencing also curtail the feasibility of large-scale experimentation. A repertoire of methodologies has been employed for forecasting perturbative responses in single-cell gene expression. However, existing methods primarily focus on the average response of a specific cell type to perturbation, overlooking the single-cell specificity of perturbation responses and a more comprehensive prediction of the entire perturbation response distribution. RESULTS: Here, we present scPRAM, a method for predicting perturbation responses in single-cell gene expression based on attention mechanisms. Leveraging variational autoencoders and optimal transport, scPRAM aligns cell states before and after perturbation, followed by accurate prediction of gene expression responses to perturbations for unseen cell types through attention mechanisms. Experiments on multiple real perturbation datasets involving drug treatments and bacterial infections demonstrate that scPRAM attains heightened accuracy in perturbation prediction across cell types, species, and individuals, surpassing existing methodologies. Furthermore, scPRAM demonstrates outstanding capability in identifying differentially expressed genes under perturbation, capturing heterogeneity in perturbation responses across species, and maintaining stability in the presence of data noise and sample size variations. AVAILABILITY AND IMPLEMENTATION: https://github.com/jiang-q19/scPRAM and https://doi.org/10.5281/zenodo.10935038. Qun Jiang, Shengquan Chen, Xiaoyang Chen 0007, Rui Jiang 0001 |
Bioinform. | 2 |
| 2024 | EpiCarousel: memory- and time-efficient identification of metacells for atlas-level single-cell chromatin accessibility dataabstractSUMMARY: Recent technical advancements in single-cell chromatin accessibility sequencing (scCAS) have brought new insights to the characterization of epigenetic heterogeneity. As single-cell genomics experiments scale up to hundreds of thousands of cells, the demand for computational resources for downstream analysis grows intractably large and exceeds the capabilities of most researchers. Here, we propose EpiCarousel, a tailored Python package based on lazy loading, parallel processing, and community detection for memory- and time-efficient identification of metacells, i.e. the emergence of homogenous cells, in large-scale scCAS data. Through comprehensive experiments on five datasets of various protocols, sample sizes, dimensions, number of cell types, and degrees of cell-type imbalance, EpiCarousel outperformed baseline methods in systematic evaluation of memory usage, computational time, and multiple downstream analyses including cell type identification. Moreover, EpiCarousel executes preprocessing and downstream cell clustering on the atlas-level dataset with 707 043 cells and 1 154 611 peaks within 2 h consuming <75 GB of RAM and provides superior performance for characterizing cell heterogeneity than state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The EpiCarousel software is well-documented and freely available at https://github.com/biox-nku/epicarousel. It can be seamlessly interoperated with extensive scCAS analysis toolkits. Xiaoyang Chen 0007, Songming Tang, Shengquan Chen |
Bioinform. | 7 |
| 2024 | SCREEN: predicting single-cell gene expression perturbation responses via optimal transport
Qun Jiang, Shengquan Chen |
Frontiers Comput. Sci. | 5 |
| 2024 | Accurate Annotation for Differentiating and Imbalanced Cell Types in Single-Cell Chromatin Accessibility DataabstractRapid advances in single-cell chromatin accessibility sequencing (scCAS) technologies have enabled the characterization of epigenomic heterogeneity and increased the demand for automatic annotation of cell types. However, there are few computational methods tailored for cell type annotation in scCAS data and the existing methods perform poorly for differentiating and imbalanced cell types. Here, we propose CASCADE, a novel annotation method based on simulation- and denoising-based strategies. With comprehensive experiments on a number of scCAS datasets, we showed that CASCADE can effectively distinguish the patterns of different cell types and mitigate the effect of high noise levels, and thus achieve significantly better annotation performance for differentiating and imbalanced cell types. Besides, we performed model ablation experiments to show the contribution of modules in CASCADE and conducted extensive experiments to demonstrate the robustness of CASCADE to batch effect, imbalance degree, data sparsity, and number of cell types. Moreover, CASCADE significantly outperformed baseline methods for accurately annotating the cell types in newly sequenced data. We anticipate that CASCADE will greatly assist with characterizing cell heterogeneity in scCAS data analysis. Yuhang Jia, Rui Jiang 0001, Shengquan Chen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | RefTM: reference-guided topic modeling of single-cell chromatin accessibility dataabstractSingle-cell analysis is a valuable approach for dissecting the cellular heterogeneity, and single-cell chromatin accessibility sequencing (scCAS) can profile the epigenetic landscapes for thousands of individual cells. It is challenging to analyze scCAS data, because of its high dimensionality and a higher degree of sparsity compared with scRNA-seq data. Topic modeling in single-cell data analysis can lead to robust identification of the cell types and it can provide insight into the regulatory mechanisms. Reference-guided approach may facilitate the analysis of scCAS data by utilizing the information in existing datasets. We present RefTM (Reference-guided Topic Modeling of single-cell chromatin accessibility data), which not only utilizes the information in existing bulk chromatin accessibility and annotated scCAS data, but also takes advantage of topic models for single-cell data analysis. RefTM simultaneously models: (1) the shared biological variation among reference data and the target scCAS data; (2) the unique biological variation in scCAS data; (3) other variations from known covariates in scCAS data. Shengquan Chen, Zhixiang Lin |
Briefings Bioinform. | 2 |
| 2023 | ASTER: accurately estimating the number of cell types in single-cell chromatin accessibility dataabstractSUMMARY: Recent innovations in single-cell chromatin accessibility sequencing (scCAS) have revolutionized the characterization of epigenomic heterogeneity. Estimation of the number of cell types is a crucial step for downstream analyses and biological implications. However, efforts to perform estimation specifically for scCAS data are limited. Here, we propose ASTER, an ensemble learning-based tool for accurately estimating the number of cell types in scCAS data. ASTER outperformed baseline methods in systematic evaluation on 27 datasets of various protocols, sizes, numbers of cell types, degrees of cell-type imbalance, cell states and qualities, providing valuable guidance for scCAS data analysis. AVAILABILITY AND IMPLEMENTATION: ASTER along with detailed documentation is freely accessible at https://aster.readthedocs.io/ under the MIT License. It can be seamlessly integrated into existing scCAS analysis workflows. The source code is available at https://github.com/biox-nku/aster. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shengquan Chen, Rongxiang Wang, Wenxin Long |
Bioinform. | 1 |
| 2023 | simCAS: an embedding-based method for simulating single-cell chromatin accessibility sequencing dataabstractMOTIVATION: Single-cell chromatin accessibility sequencing (scCAS) technology provides an epigenomic perspective to characterize gene regulatory mechanisms at single-cell resolution. With an increasing number of computational methods proposed for analyzing scCAS data, a powerful simulation framework is desirable for evaluation and validation of these methods. However, existing simulators generate synthetic data by sampling reads from real data or mimicking existing cell states, which is inadequate to provide credible ground-truth labels for method evaluation. RESULTS: We present simCAS, an embedding-based simulator, for generating high-fidelity scCAS data from both cell- and peak-wise embeddings. We demonstrate simCAS outperforms existing simulators in resembling real data and show that simCAS can generate cells of different states with user-defined cell populations and differentiation trajectories. Additionally, simCAS can simulate data from different batches and encode user-specified interactions of chromatin regions in the synthetic data, which provides ground-truth labels more than cell states. We systematically demonstrate that simCAS facilitates the benchmarking of four core tasks in downstream analysis: cell clustering, trajectory inference, data integration, and cis-regulatory interaction inference. We anticipate simCAS will be a reliable and flexible simulator for evaluating the ongoing computational methods applied on scCAS data. AVAILABILITY AND IMPLEMENTATION: simCAS is freely available at https://github.com/Chen-Li-17/simCAS. Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001, Xuegong Zhang |
Bioinform. | 3 |
| 2021 | stPlus: a reference-based method for the accurate enhancement of spatial transcriptomicsabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. RESULTS: We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial gene expression via a weighted k-nearest-neighbor. stPlus outperforms baseline methods with higher gene-wise and cell-wise Spearman correlation coefficients. We also introduce a clustering-based approach to assess the enhancement performance systematically. Using the data enhanced by stPlus, cell populations can be better identified than using the measured data. The predicted expression of genes unique to scRNA-seq data can also well characterize spatial cell heterogeneity. Besides, stPlus is robust and scalable to datasets of diverse gene detection sensitivity levels, sample sizes and number of spatially measured genes. We anticipate stPlus will facilitate the analysis of spatial transcriptomics. AVAILABILITY AND IMPLEMENTATION: stPlus with detailed documents is freely accessible at http://health.tsinghua.edu.cn/software/stPlus/ and the source code is openly available on https://github.com/xy-chen16/stPlus. Shengquan Chen, Boheng Zhang, Xiaoyang Chen 0007, Xuegong Zhang, Rui Jiang 0001 |
Bioinform. | 1 |
| 2020 | Research and Improvement of Community Discovery Algorithm Based on Spark for Large Scale Complicated NetworksabstractCommunity discovery algorithm is one of the important topics in complex network research. However, there exist some problems in the traditional community discovery algorithm, such as the oscillation of label propagation, the convergence of iteration, the formation of a large community with a single label, namely “monster community”, or the poor effect of community discovery due to equal treatment of nodes. On the other hand, with the advent of the era of data, the computing power of single computer can't meet the demand of the rapid growth of complex network scale. Based on the above knowledge, this paper proposes the research and improvement of community discovery algorithm based on spark for large-scale complex networks. This paper first weights a complex network with no weight. Then, this paper chooses the classic efficient community discovery algorithm - label propagation algorithm to optimize label initialization, label propagation and label update strategy, iterative convergence strategy and so on, and establishes a new community discovery algorithm model. Then, the algorithm is connected to Spark, the algorithm is synchronized through GraphX programming, and a Spark experiment platform is established. Finally, some classic complex network data and some large-scale complex network data are tested and compared with some classic community discovery algorithms to verify the proposed algorithm is validated and verified by a large-scale complex network data set based on Spark GraphX platform does greatly improve the computational performance of community discovery in complex networks. Shengquan Chen, Lingfeng Lu, Chenkun Meng |
TrustCom | 2 |
| 2020 | EnClaSC: a novel ensemble approach for accurate and robust cell-type classification of single-cell transcriptomesabstractBACKGROUND: In recent years, the rapid development of single-cell RNA-sequencing (scRNA-seq) techniques enables the quantitative characterization of cell types at a single-cell resolution. With the explosive growth of the number of cells profiled in individual scRNA-seq experiments, there is a demand for novel computational methods for classifying newly-generated scRNA-seq data onto annotated labels. Although several methods have recently been proposed for the cell-type classification of single-cell transcriptomic data, such limitations as inadequate accuracy, inferior robustness, and low stability greatly limit their wide applications. RESULTS: We propose a novel ensemble approach, named EnClaSC, for accurate and robust cell-type classification of single-cell transcriptomic data. Through comprehensive validation experiments, we demonstrate that EnClaSC can not only be applied to the self-projection within a specific dataset and the cell-type classification across different datasets, but also scale up well to various data dimensionality and different data sparsity. We further illustrate the ability of EnClaSC to effectively make cross-species classification, which may shed light on the studies in correlation of different species. EnClaSC is freely available at https://github.com/xy-chen16/EnClaSC . CONCLUSIONS: EnClaSC enables highly accurate and robust cell-type classification of single-cell transcriptomic data via an ensemble learning method. We expect to see wide applications of our method to not only transcriptome studies, but also the classification of more general data. Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001 |
BMC Bioinform. | 2 |
| 2019 | VPAC: Variational projection for accurate clustering of single-cell transcriptomic dataabstractBACKGROUND: Single-cell RNA-sequencing (scRNA-seq) technologies have advanced rapidly in recent years and enabled the quantitative characterization at a microscopic resolution. With the exponential growth of the number of cells profiled in individual scRNA-seq experiments, the demand for identifying putative cell types from the data has become a great challenge that appeals for novel computational methods. Although a variety of algorithms have recently been proposed for single-cell clustering, such limitations as low accuracy, inferior robustness, and inadequate stability greatly impede the scope of applications of these methods. RESULTS: We propose a novel model-based algorithm, named VPAC, for accurate clustering of single-cell transcriptomic data through variational projection, which assumes that single-cell samples follow a Gaussian mixture distribution in a latent space. Through comprehensive validation experiments, we demonstrate that VPAC can not only be applied to datasets of discrete counts and normalized continuous data, but also scale up well to various data dimensionality, different dataset size and different data sparsity. We further illustrate the ability of VPAC to detect genes with strong unique signatures of a specific cell type, which may shed light on the studies in system biology. We have released a user-friendly python package of VPAC in Github ( https://github.com/ShengquanChen/VPAC ). Users can directly import our VPAC class and conduct clustering without tedious installation of dependency packages. CONCLUSIONS: VPAC enables highly accurate clustering of single-cell transcriptomic data via a statistical model. We expect to see wide applications of our method to not only transcriptome studies for fully understanding the cell identity and functionality, but also the clustering of more general data. Shengquan Chen, Kui Hua, Hongfei Cui, Rui Jiang 0001 |
BMC Bioinform. | 1 |
| 2017 | Predicting enhancers with deep convolutional neural networksabstractBACKGROUND: With the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. RESULTS: To facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database. CONCLUSIONS: DeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences. Xu Min, Wanwen Zeng, Shengquan Chen, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001 |
BMC Bioinform. | 3 |