Rui Chen 0021

dblp:02/1003-21 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-4341-2908ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Leveraging scHi-C Data for Integrated Single-Cell Omics Analysis
abstract
The integration of single-cell multi-omics data is essential for deciphering complex gene regulatory programs. While single-cell Hi-C (scHi-C) provides insight into 3D genome architecture, its utilization in multi-omics integration remains underexplored. Here, we use mouse brain single-cell multi-omics integration as a case study to demonstrate the dual utility of scHiC data within a knowledge graph-based integrative framework. First, we use scHi-C as a gene regulatory prior to construct a Hi-C-driven guidance graph. This approach enhances integration of scRNA-seq and scATAC-seq data, resulting in an improved alignment score (FOSCTTM$=0.0356)$. Second, we show that the framework can directly integrate scHi-C as a primary data modality with scRNA-seq. We validate this method on a paired scRNA-seq and scHi-C dataset, achieving 85.0% mapping accuracy, and demonstrate its power in a cross-modal labeltransfer application. This direct integration successfully refines a broader neuronal cluster into finer, distinct hippocampal granule and pyramidal subtypes. Our work presents an effective strategy for leveraging scHi-C data in multi-omics integration, providing a powerful tool to dissect cellular regulatory programs through combining 3D genome organization with gene expression.
Weixin Liu 0001, Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin
BIBM2
2025 Integrating Single Cell RNA Sequencing Data and Protein Embeddings to Infer Cell-Cell Communication in Alzheimer's Disease
abstract
Cell-cell communication (CCC) plays a critical role in the pathogenesis of Alzheimer's disease (AD), yet most computational methods for CCC inference rely exclusively on transcriptomic data, showing low consistency across datasets or methods. In this study, we present a Protein Embedding-Infused Cell Talk (PEICTalk) inference method that integrates singlecell RNA-sequencing (scRNA-seq) data with protein embeddings from the ESM-2 language model to improve the accuracy and interpretability of CCC inference. We applied PEICTalk to two large-scale scRNA-seq datasets, after harmonizing celltype annotations using the SEA-AD taxonomy via MapMyCells. PEICTalk outperformed standard tools such as CellChat and CellPhoneDB, identifying more biologically meaningful ligandreceptor interactions with higher cross-dataset reproducibility. Ablation analysis confirmed the crucial value of protein-level information. This integrative strategy offers a robust foundation for uncovering novel intercellular mechanisms in AD and may serve as a blueprint for future CCC studies.
Yuting Tan 0005, Rui Chen 0021, Anshul Tiwari, Zhexing Wen, Xue Zhong, Zhijun Yin, Bingshan Li
BIBM3
2025 Tensor decomposition of multi-dimensional splicing events across multiple tissues to identify splicing-mediated risk genes associated with complex traits
abstract
Identifying risk genes associated with complex traits remains challenging. Integrating gene expression data with Genome-Wide Association Study (GWAS) through Transcriptome-Wide Association Study (TWAS) methods has discovered candidate risk genes for various complex traits. Splicing, which explains a comparable heritability of complex traits as gene expression, is under-explored due to its multidimensionality. To leverage multiple splicing events in a gene and shared splicing across tissues, we develop Multi-tissue Splicing Gene (MTSG), which employs tensor decomposition and sparse Canonical Correlation Analysis (sCCA) to extract meaningful information from high-dimensional multiple splicing events across multiple tissues. We build MTSG models using GTEx data and apply them to GWAS summary statistics of Alzheimer's disease (AD) (111,326 cases and 677,663 controls) and schizophrenia (SCZ) (36,989 cases and 113,075 controls). We identify 174 and 497 significant splicing-mediated risk genes for AD and SCZ, respectively, at Bonferroni correction. For AD, our results demonstrate significant enrichment of AD related pathways and identify additional AD risk genes not detected in the single-tissue analysis, while preserving most top genes identified in the brain frontal cortex. Consistently, for SCZ, genes identified by our brain-wide MTSG model, built from a cluster of 13 brain tissues, exhibit stronger enrichment in SCZ-relevant genes and MTSG identifies unique SCZ risk genes compared to single-tissue models. These results showcase that our MTSG models capture distinctive splicing events across tissues, which might be overlooked when using single tissue alone. Our MTSG models can be applied to other complex traits to help identify splicing-mediated disease risk genes.
Rui Chen 0021, Hakmook Kang, Yuting Tan 0005, Anshul Tiwari, Zhexing Wen, Xue Zhong, Bingshan Li
PLoS Comput. Biol.2
2024 Improving Genetic Perturbation Response Prediction with an Enhanced Biological Knowledge Graph
abstract
Perturb-seq is a technique that combines scRNA-seq and CRISPR to explore cellular system operations and disease-associated genes, providing profound insights into the mechanisms behind biological processes. Although powerful, such a method is limited by its scalability for its cost-intensive and time-consuming nature, which calls for in silico prediction of genetic perturbation responses. Among all computational methods, GEARS represents the state-of-the-art by explicitly modeling the response of each gene to the perturbed gene, exploiting gene-gene relationships derived from Gene Ontology annotations. However, our evaluation of Gene Ontology annotations indicated that they are insufficient as the sole source of prior knowledge for predicting genetic perturbation responses. Therefore, they cannot fully support predicting genetic perturbation responses. We addressed this gap by constructing an augmented gene ontology network that incorporates extensive knowledge of diseases, drugs, and genes to capture nuanced gene-gene relationships not indicated by Gene Ontology alone. By replacing only the Gene Ontology graph in GEARS, our method outperforms GEARS in both single gene and combinational perturbation predictions. These findings suggest the effectiveness and importance of incorporating finer prior knowledge in predicting genetic perturbation responses, thereby encouraging future works on improving knowledge representation for single-cell perturbation prediction.
Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin
BIBM2
2023 From multi-omics data to the cancer druggable gene discovery: a novel machine learning-based approach
abstract
The development of targeted drugs allows precision medicine in cancer treatment and optimal targeted therapies. Accurate identification of cancer druggable genes helps strengthen the understanding of targeted cancer therapy and promotes precise cancer treatment. However, rare cancer-druggable genes have been found due to the multi-omics data's diversity and complexity. This study proposes deep forest for cancer druggable genes discovery (DF-CAGE), a novel machine learning-based method for cancer-druggable gene discovery. DF-CAGE integrated the somatic mutations, copy number variants, DNA methylation and RNA-Seq data across ˜10 000 TCGA profiles to identify the landscape of the cancer-druggable genes. We found that DF-CAGE discovers the commonalities of currently known cancer-druggable genes from the perspective of multi-omics data and achieved excellent performance on OncoKB, Target and Drugbank data sets. Among the ˜20 000 protein-coding genes, DF-CAGE pinpointed 465 potential cancer-druggable genes. We found that the candidate cancer druggable genes (CDG) are clinically meaningful and divided the CDG into known, reliable and potential gene sets. Finally, we analyzed the omics data's contribution to identifying druggable genes. We found that DF-CAGE reports druggable genes mainly based on the copy number variations (CNVs) data, the gene rearrangements and the mutation rates in the population. These findings may enlighten the future study and development of new drugs.
Hai Yang 0002, Lipeng Gan, Rui Chen 0021, Dongdong Li 0003, Jing Zhang 0041, Zhe Wang 0002
Briefings Bioinform.3
2022 TVAR: assessing tissue-specific functional effects of non-coding variants with deep learning
abstract
MOTIVATION: Analysis of whole-genome sequencing (WGS) for genetics is still a challenge due to the lack of accurate functional annotation of non-coding variants, especially the rare ones. As eQTLs have been extensively implicated in the genetics of human diseases, we hypothesize that rare non-coding variants discovered in WGS play a regulatory role in predisposing disease risk. RESULTS: With thousands of tissue- and cell-type-specific epigenomic features, we propose TVAR. This multi-label learning-based deep neural network predicts the functionality of non-coding variants in the genome based on eQTLs across 49 human tissues in the GTEx project. TVAR learns the relationships between high-dimensional epigenomics and eQTLs across tissues, taking the correlation among tissues into account to understand shared and tissue-specific eQTL effects. As a result, TVAR outputs tissue-specific annotations, with an average AUROC of 0.77 across these tissues. We evaluate TVAR's performance on four complex diseases (coronary artery disease, breast cancer, Type 2 diabetes and Schizophrenia), using TVAR's tissue-specific annotations, and observe its superior performance in predicting functional variants for both common and rare variants, compared with five existing state-of-the-art tools. We further evaluate TVAR's G-score, a scoring scheme across all tissues, on ClinVar, fine-mapped GWAS loci, Massive Parallel Reporter Assay (MPRA) validated variants and observe the consistently better performance of TVAR compared with other competing tools. AVAILABILITY AND IMPLEMENTATION: The TVAR source code and its scores on the ClinVar catalog, fine mapped GWAS Loci, high confidence eQTLs from GTEx dataset, and MPRA validated functional variants are available at https://github.com/haiyang1986/TVAR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Xue Zhong, Bingshan Li
Bioinform.2
2022 A Bayesian framework to integrate multi-level genome-scale data for Autism risk gene prioritization
abstract
BACKGROUND: Autism spectrum disorder (ASD) is a group of complex neurodevelopment disorders with a strong genetic basis. Large scale sequencing studies have identified over one hundred ASD risk genes. Nevertheless, the vast majority of ASD risk genes remain to be discovered, as it is estimated that more than 1000 genes are likely to be involved in ASD risk. Prioritization of risk genes is an effective strategy to increase the power of identifying novel risk genes in genetics studies of ASD. As ASD risk genes are likely to exhibit distinct properties from multiple angles, we reason that integrating multiple levels of genomic data is a powerful approach to pinpoint genuine ASD risk genes. RESULTS: We present BNScore, a Bayesian model selection framework to probabilistically prioritize ASD risk genes through explicitly integrating evidence from sequencing-identified ASD genes, biological annotations, and gene functional network. We demonstrate the validity of our approach and its improved performance over existing methods by examining the resulting top candidate ASD risk genes against sets of high-confidence benchmark genes and large-scale ASD genome-wide association studies. We assess the tissue-, cell type- and development stage-specific expression properties of top prioritized genes, and find strong expression specificity in brain tissues, striatal medium spiny neurons, and fetal developmental stages. CONCLUSIONS: In summary, we show that by integrating sequencing findings, functional annotation profiles, and gene-gene functional network, our proposed BNScore provides competitive performance compared to current state-of-the-art methods in prioritizing ASD genes. Our method offers a general and flexible strategy to risk gene prioritization that can potentially be applied to other complex traits as well.
Ying Ji 0002, Rui Chen 0021, Quan Wang 0004, Bingshan Li
BMC Bioinform.2
2021 Subtype-GAN: a deep learning approach for integrative cancer subtyping of multi-omics data
abstract
MOTIVATION: The discovery of cancer subtyping can help explore cancer pathogenesis, determine clinical actionability in treatment, and improve patients' survival rates. However, due to the diversity and complexity of multi-omics data, it is still challenging to develop integrated clustering algorithms for tumor molecular subtyping. RESULTS: We propose Subtype-GAN, a deep adversarial learning approach based on the multiple-input multiple-output neural network to model the complex omics data accurately. With the latent variables extracted from the neural network, Subtype-GAN uses consensus clustering and the Gaussian Mixture model to identify tumor samples' molecular subtypes. Compared with other state-of-the-art subtyping approaches, Subtype-GAN achieved outstanding performance on the benchmark datasets consisting of ∼4000 TCGA tumors from 10 types of cancer. We found that on the comparison dataset, the clustering scheme of Subtype-GAN is not always similar to that of the deep learning method AE but is identical to that of NEMO, MCCA, VAE and other excellent approaches. Finally, we applied Subtype-GAN to the BRCA dataset and automatically obtained the number of subtypes and the subtype labels of 1031 BRCA tumors. Through the detailed analysis, we found that the identified subtypes are clinically meaningful and show distinct patterns in the feature space, demonstrating the practicality of Subtype-GAN. AVAILABILITYAND IMPLEMENTATION: The source codes, the clustering results of Subtype-GAN across the benchmark datasets are available at https://github.com/haiyang1986/Subtype-GAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Dongdong Li 0003, Zhe Wang 0002
Bioinform.2
2020 DRAMS: A tool to detect and re-align mixed-up samples for integrative studies of multi-omics data
abstract
Studies of complex disorders benefit from integrative analyses of multiple omics data. Yet, sample mix-ups frequently occur in multi-omics studies, weakening statistical power and risking false findings. Accurately aligning sample information, genotype, and corresponding omics data is critical for integrative analyses. We developed DRAMS (https://github.com/Yi-Jiang/DRAMS) to Detect and Re-Align Mixed-up Samples to address the sample mix-up problem. It uses a logistic regression model followed by a modified topological sorting algorithm to identify the potential true IDs based on data relationships of multi-omics. According to tests using simulated data, the more types of omics data used or the smaller the proportion of mix-ups, the better that DRAMS performs. Applying DRAMS to real data from the PsychENCODE BrainGVEX project, we detected and corrected 201 (12.5% of total data generated) mix-ups. Of the 21 mix-ups involving errors of racial identity, DRAMS re-assigned all data to the correct racial group in the 1000 Genomes project. In doing so, quantitative trait loci (QTL) (FDR<0.01) increased by an average of 1.62-fold. The use of DRAMS in multi-omics studies will strengthen statistical power of the study and improve quality of the results. Even though very limited studies have multi-omics data in place, we expect such data will increase quickly with the needs of DRAMS.
Gina Giase, Kay Grennan, Annie W. Shieh, Lide Han, Quan Wang 0004, Rui Chen 0021, Kevin P. White, Chao Chen 0041, Bingshan Li, Chunyu Liu 0001
PLoS Comput. Biol.9
2019 De novo pattern discovery enables robust assessment of functional consequences of non-coding variants
abstract
MOTIVATION: Given the complexity of genome regions, prioritize the functional effects of non-coding variants remains a challenge. Although several frameworks have been proposed for the evaluation of the functionality of non-coding variants, most of them used 'black boxes' methods that simplify the task as the pathogenicity/benign classification problem, which ignores the distinct regulatory mechanisms of variants and leads to less desirable performance. In this study, we developed DVAR, an unsupervised framework that leverage various biochemical and evolutionary evidence to distinguish the gene regulatory categories of variants and assess their comprehensive functional impact simultaneously. RESULTS: DVAR performed de novo pattern discovery in high-dimensional data and identified five regulatory clusters of non-coding variants. Leveraging the new insights into the multiple functional patterns, it measures both the between-class and the within-class functional implication of the variants to achieve accurate prioritization. Compared to other two-class learning methods, it showed improved performance in identification of clinically significant variants, fine-mapped GWAS variants, eQTLs and expression-modulating variants. Moreover, it has superior performance on disease causal variants verified by genome-editing (like CRISPR-Cas9), which could provide a pre-selection strategy for genome-editing technologies across the whole genome. Finally, evaluated in BioVU and UK Biobank, two large-scale DNA biobanks linked to complete electronic health records, DVAR demonstrated its effectiveness in prioritizing non-coding variants associated with medical phenotypes. AVAILABILITY AND IMPLEMENTATION: The C++ and Python source codes, the pre-computed DVAR-cluster labels and DVAR-scores across the whole genome are available at https://www.vumc.org/cgg/dvar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Guangze Zheng 0002, Xue Zhong, Nancy J. Cox, Bingshan Li
Bioinform.2
2015 A haplotype-based framework for group-wise transmission/disequilibrium tests for rare variant association analysis
abstract
MOTIVATION: A major focus of current sequencing studies for human genetics is to identify rare variants associated with complex diseases. Aside from reduced power of detecting associated rare variants, controlling for population stratification is particularly challenging for rare variants. Transmission/disequilibrium tests (TDT) based on family designs are robust to population stratification and admixture, and therefore provide an effective approach to rare variant association studies to eliminate spurious associations. To increase power of rare variant association analysis, gene-based collapsing methods become standard approaches for analyzing rare variants. Existing methods that extend this strategy to rare variants in families usually combine TDT statistics at individual variants and therefore lack the flexibility of incorporating other genetic models. RESULTS: In this study, we describe a haplotype-based framework for group-wise TDT (gTDT) that is flexible to encompass a variety of genetic models such as additive, dominant and compound heterozygous (CH) (i.e. recessive) models as well as other complex interactions. Unlike existing methods, gTDT constructs haplotypes by transmission when possible and inherently takes into account the linkage disequilibrium among variants. Through extensive simulations we showed that type I error was correctly controlled for rare variants under all models investigated, and this remained true in the presence of population stratification. Under a variety of genetic models, gTDT showed increased power compared with the single marker TDT. Application of gTDT to an autism exome sequencing data of 118 trios identified potentially interesting candidate genes with CH rare variants. AVAILABILITY AND IMPLEMENTATION: We implemented gTDT in C++ and the source code and the detailed usage are available on the authors' website (https://medschool.vanderbilt.edu/cgg). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rui Chen 0021, Xiaowei Zhan, Xue Zhong, James S. Sutcliffe, Nancy J. Cox, Edwin H. Cook Jr., Wei Chen 0074, Bingshan Li
Bioinform.1