VLDB 2026 Research / reviewers in the wild / expert
Xue Zhong
dblp:66/5354
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging scHi-C Data for Integrated Single-Cell Omics AnalysisabstractThe integration of single-cell multi-omics data is essential for deciphering complex gene regulatory programs. While single-cell Hi-C (scHi-C) provides insight into 3D genome architecture, its utilization in multi-omics integration remains underexplored. Here, we use mouse brain single-cell multi-omics integration as a case study to demonstrate the dual utility of scHiC data within a knowledge graph-based integrative framework. First, we use scHi-C as a gene regulatory prior to construct a Hi-C-driven guidance graph. This approach enhances integration of scRNA-seq and scATAC-seq data, resulting in an improved alignment score (FOSCTTM$=0.0356)$. Second, we show that the framework can directly integrate scHi-C as a primary data modality with scRNA-seq. We validate this method on a paired scRNA-seq and scHi-C dataset, achieving 85.0% mapping accuracy, and demonstrate its power in a cross-modal labeltransfer application. This direct integration successfully refines a broader neuronal cluster into finer, distinct hippocampal granule and pyramidal subtypes. Our work presents an effective strategy for leveraging scHi-C data in multi-omics integration, providing a powerful tool to dissect cellular regulatory programs through combining 3D genome organization with gene expression. Weixin Liu 0001, Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin |
BIBM | 5 |
| 2025 | Integrating Single Cell RNA Sequencing Data and Protein Embeddings to Infer Cell-Cell Communication in Alzheimer's DiseaseabstractCell-cell communication (CCC) plays a critical role in the pathogenesis of Alzheimer's disease (AD), yet most computational methods for CCC inference rely exclusively on transcriptomic data, showing low consistency across datasets or methods. In this study, we present a Protein Embedding-Infused Cell Talk (PEICTalk) inference method that integrates singlecell RNA-sequencing (scRNA-seq) data with protein embeddings from the ESM-2 language model to improve the accuracy and interpretability of CCC inference. We applied PEICTalk to two large-scale scRNA-seq datasets, after harmonizing celltype annotations using the SEA-AD taxonomy via MapMyCells. PEICTalk outperformed standard tools such as CellChat and CellPhoneDB, identifying more biologically meaningful ligandreceptor interactions with higher cross-dataset reproducibility. Ablation analysis confirmed the crucial value of protein-level information. This integrative strategy offers a robust foundation for uncovering novel intercellular mechanisms in AD and may serve as a blueprint for future CCC studies. Yuting Tan 0005, Rui Chen 0021, Anshul Tiwari, Zhexing Wen, Xue Zhong, Zhijun Yin, Bingshan Li |
BIBM | 6 |
| 2025 | Tensor decomposition of multi-dimensional splicing events across multiple tissues to identify splicing-mediated risk genes associated with complex traitsabstractIdentifying risk genes associated with complex traits remains challenging. Integrating gene expression data with Genome-Wide Association Study (GWAS) through Transcriptome-Wide Association Study (TWAS) methods has discovered candidate risk genes for various complex traits. Splicing, which explains a comparable heritability of complex traits as gene expression, is under-explored due to its multidimensionality. To leverage multiple splicing events in a gene and shared splicing across tissues, we develop Multi-tissue Splicing Gene (MTSG), which employs tensor decomposition and sparse Canonical Correlation Analysis (sCCA) to extract meaningful information from high-dimensional multiple splicing events across multiple tissues. We build MTSG models using GTEx data and apply them to GWAS summary statistics of Alzheimer's disease (AD) (111,326 cases and 677,663 controls) and schizophrenia (SCZ) (36,989 cases and 113,075 controls). We identify 174 and 497 significant splicing-mediated risk genes for AD and SCZ, respectively, at Bonferroni correction. For AD, our results demonstrate significant enrichment of AD related pathways and identify additional AD risk genes not detected in the single-tissue analysis, while preserving most top genes identified in the brain frontal cortex. Consistently, for SCZ, genes identified by our brain-wide MTSG model, built from a cluster of 13 brain tissues, exhibit stronger enrichment in SCZ-relevant genes and MTSG identifies unique SCZ risk genes compared to single-tissue models. These results showcase that our MTSG models capture distinctive splicing events across tissues, which might be overlooked when using single tissue alone. Our MTSG models can be applied to other complex traits to help identify splicing-mediated disease risk genes. Rui Chen 0021, Hakmook Kang, Yuting Tan 0005, Anshul Tiwari, Zhexing Wen, Xue Zhong, Bingshan Li |
PLoS Comput. Biol. | 8 |
| 2024 | Improving Genetic Perturbation Response Prediction with an Enhanced Biological Knowledge GraphabstractPerturb-seq is a technique that combines scRNA-seq and CRISPR to explore cellular system operations and disease-associated genes, providing profound insights into the mechanisms behind biological processes. Although powerful, such a method is limited by its scalability for its cost-intensive and time-consuming nature, which calls for in silico prediction of genetic perturbation responses. Among all computational methods, GEARS represents the state-of-the-art by explicitly modeling the response of each gene to the perturbed gene, exploiting gene-gene relationships derived from Gene Ontology annotations. However, our evaluation of Gene Ontology annotations indicated that they are insufficient as the sole source of prior knowledge for predicting genetic perturbation responses. Therefore, they cannot fully support predicting genetic perturbation responses. We addressed this gap by constructing an augmented gene ontology network that incorporates extensive knowledge of diseases, drugs, and genes to capture nuanced gene-gene relationships not indicated by Gene Ontology alone. By replacing only the Gene Ontology graph in GEARS, our method outperforms GEARS in both single gene and combinational perturbation predictions. These findings suggest the effectiveness and importance of incorporating finer prior knowledge in predicting genetic perturbation responses, thereby encouraging future works on improving knowledge representation for single-cell perturbation prediction. Rui Chen 0021, Yuting Tan 0005, Xue Zhong, Bingshan Li, Zhijun Yin |
BIBM | 4 |
| 2022 | TVAR: assessing tissue-specific functional effects of non-coding variants with deep learningabstractMOTIVATION: Analysis of whole-genome sequencing (WGS) for genetics is still a challenge due to the lack of accurate functional annotation of non-coding variants, especially the rare ones. As eQTLs have been extensively implicated in the genetics of human diseases, we hypothesize that rare non-coding variants discovered in WGS play a regulatory role in predisposing disease risk. RESULTS: With thousands of tissue- and cell-type-specific epigenomic features, we propose TVAR. This multi-label learning-based deep neural network predicts the functionality of non-coding variants in the genome based on eQTLs across 49 human tissues in the GTEx project. TVAR learns the relationships between high-dimensional epigenomics and eQTLs across tissues, taking the correlation among tissues into account to understand shared and tissue-specific eQTL effects. As a result, TVAR outputs tissue-specific annotations, with an average AUROC of 0.77 across these tissues. We evaluate TVAR's performance on four complex diseases (coronary artery disease, breast cancer, Type 2 diabetes and Schizophrenia), using TVAR's tissue-specific annotations, and observe its superior performance in predicting functional variants for both common and rare variants, compared with five existing state-of-the-art tools. We further evaluate TVAR's G-score, a scoring scheme across all tissues, on ClinVar, fine-mapped GWAS loci, Massive Parallel Reporter Assay (MPRA) validated variants and observe the consistently better performance of TVAR compared with other competing tools. AVAILABILITY AND IMPLEMENTATION: The TVAR source code and its scores on the ClinVar catalog, fine mapped GWAS Loci, high confidence eQTLs from GTEx dataset, and MPRA validated functional variants are available at https://github.com/haiyang1986/TVAR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Xue Zhong, Bingshan Li |
Bioinform. | 6 |
| 2019 | De novo pattern discovery enables robust assessment of functional consequences of non-coding variantsabstractMOTIVATION: Given the complexity of genome regions, prioritize the functional effects of non-coding variants remains a challenge. Although several frameworks have been proposed for the evaluation of the functionality of non-coding variants, most of them used 'black boxes' methods that simplify the task as the pathogenicity/benign classification problem, which ignores the distinct regulatory mechanisms of variants and leads to less desirable performance. In this study, we developed DVAR, an unsupervised framework that leverage various biochemical and evolutionary evidence to distinguish the gene regulatory categories of variants and assess their comprehensive functional impact simultaneously. RESULTS: DVAR performed de novo pattern discovery in high-dimensional data and identified five regulatory clusters of non-coding variants. Leveraging the new insights into the multiple functional patterns, it measures both the between-class and the within-class functional implication of the variants to achieve accurate prioritization. Compared to other two-class learning methods, it showed improved performance in identification of clinically significant variants, fine-mapped GWAS variants, eQTLs and expression-modulating variants. Moreover, it has superior performance on disease causal variants verified by genome-editing (like CRISPR-Cas9), which could provide a pre-selection strategy for genome-editing technologies across the whole genome. Finally, evaluated in BioVU and UK Biobank, two large-scale DNA biobanks linked to complete electronic health records, DVAR demonstrated its effectiveness in prioritizing non-coding variants associated with medical phenotypes. AVAILABILITY AND IMPLEMENTATION: The C++ and Python source codes, the pre-computed DVAR-cluster labels and DVAR-scores across the whole genome are available at https://www.vumc.org/cgg/dvar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Guangze Zheng 0002, Xue Zhong, Nancy J. Cox, Bingshan Li |
Bioinform. | 7 |
| 2017 | Cancer driver gene discovery through an integrative genomics approach in a non-parametric Bayesian frameworkabstractMotivation: Comprehensive catalogue of genes that drive tumor initiation and progression in cancer is key to advancing diagnostics, therapeutics and treatment. Given the complexity of cancer, the catalogue is far from complete yet. Increasing evidence shows that driver genes exhibit consistent aberration patterns across multiple-omics in tumors. In this study, we aim to leverage complementary information encoded in each of the omics data to identify novel driver genes through an integrative framework. Specifically, we integrated mutations, gene expression, DNA copy numbers, DNA methylation and protein abundance, all available in The Cancer Genome Atlas (TCGA) and developed iDriver, a non-parametric Bayesian framework based on multivariate statistical modeling to identify driver genes in an unsupervised fashion. iDriver captures the inherent clusters of gene aberrations and constructs the background distribution that is used to assess and calibrate the confidence of driver genes identified through multi-dimensional genomic data. Results: We applied the method to 4 cancer types in TCGA and identified candidate driver genes that are highly enriched with known drivers. (e.g.: P < 3.40 × 10 -36 for breast cancer). We are particularly interested in novel genes and observed multiple lines of supporting evidence. Using systematic evaluation from multiple independent aspects, we identified 45 candidate driver genes that were not previously known across these 4 cancer types. The finding has important implications that integrating additional genomic data with multivariate statistics can help identify cancer drivers and guide the next stage of cancer genomics research. Availability and Implementation: The C ++ source code is freely available at https://medschool.vanderbilt.edu/cgg/ . Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Xue Zhong, Hushan Yang, Bingshan Li |
Bioinform. | 3 |
| 2015 | A haplotype-based framework for group-wise transmission/disequilibrium tests for rare variant association analysisabstractMOTIVATION: A major focus of current sequencing studies for human genetics is to identify rare variants associated with complex diseases. Aside from reduced power of detecting associated rare variants, controlling for population stratification is particularly challenging for rare variants. Transmission/disequilibrium tests (TDT) based on family designs are robust to population stratification and admixture, and therefore provide an effective approach to rare variant association studies to eliminate spurious associations. To increase power of rare variant association analysis, gene-based collapsing methods become standard approaches for analyzing rare variants. Existing methods that extend this strategy to rare variants in families usually combine TDT statistics at individual variants and therefore lack the flexibility of incorporating other genetic models. RESULTS: In this study, we describe a haplotype-based framework for group-wise TDT (gTDT) that is flexible to encompass a variety of genetic models such as additive, dominant and compound heterozygous (CH) (i.e. recessive) models as well as other complex interactions. Unlike existing methods, gTDT constructs haplotypes by transmission when possible and inherently takes into account the linkage disequilibrium among variants. Through extensive simulations we showed that type I error was correctly controlled for rare variants under all models investigated, and this remained true in the presence of population stratification. Under a variety of genetic models, gTDT showed increased power compared with the single marker TDT. Application of gTDT to an autism exome sequencing data of 118 trios identified potentially interesting candidate genes with CH rare variants. AVAILABILITY AND IMPLEMENTATION: We implemented gTDT in C++ and the source code and the detailed usage are available on the authors' website (https://medschool.vanderbilt.edu/cgg). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rui Chen 0021, Xiaowei Zhan, Xue Zhong, James S. Sutcliffe, Nancy J. Cox, Edwin H. Cook Jr., Wei Chen 0074, Bingshan Li |
Bioinform. | 4 |
| 2015 | A Bayesian framework for de novo mutation calling in parents-offspring triosabstractMOTIVATION: Spontaneous (de novo) mutations play an important role in the disease etiology of a range of complex diseases. Identifying de novo mutations (DNMs) in sporadic cases provides an effective strategy to find genes or genomic regions implicated in the genetics of disease. High-throughput next-generation sequencing enables genome- or exome-wide detection of DNMs by sequencing parents-proband trios. It is challenging to sift true mutations through massive amount of noise due to sequencing error and alignment artifacts. One of the critical limitations of existing methods is that for all genomic regions the same pre-specified mutation rate is assumed, which has a significant impact on the DNM calling accuracy. RESULTS: In this study, we developed and implemented a novel Bayesian framework for DNM calling in trios (TrioDeNovo), which overcomes these limitations by disentangling prior mutation rates from evaluation of the likelihood of the data so that flexible priors can be adjusted post-hoc at different genomic sites. Through extensively simulations and application to real data we showed that this new method has improved sensitivity and specificity over existing methods, and provides a flexible framework to further improve the efficiency by incorporating proper priors. The accuracy is further improved using effective filtering based on sequence alignment characteristics. AVAILABILITY AND IMPLEMENTATION: The C++ source code implementing TrioDeNovo is freely available at https://medschool.vanderbilt.edu/cgg. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaowei Zhan, Xue Zhong, Yongzhuang Liu, Yujun Han, Wei Chen 0074, Bingshan Li |
Bioinform. | 3 |
| 2014 | CLIP-EZ: a computational tool for HITS-CLIP data analysisabstractBackground Mapping the binding regions of mRNA-binding proteins is critical to the understanding of their regulatory roles in cellular processes. Recent development in experimental technologies combines high throughput sequencing with crosslink immunoprecipitation (HITS-CLIP), which has the merit of detecting RNA-protein interaction sites at a high resolution to single nucleotide level. Analysis of such data typically involves many steps, and teasing out true signals from noise requires crosslink induced mutations (CIMS) analysis, peak identification, integration of the two signal types, and some downstream analysis such as motif finding and conservation evaluation. To our knowledge, there is a lack of a single computational tool that can perform all the tasks as mentioned in an easily accessible manner. Despite the fact that there are several tools available, each performs an individual task. Xue Zhong, Qi Liu 0024, Shyr Yu |
BMC Bioinform. | 1 |
| 2006 | Growth of self-canceling code in evolutionary systemsabstractThis research examines the behavior of inoperative code (introns) in the evolution of genetically robust solutions. Genetically robust solutions are solutions that are less likely to be degraded by genetic operators, such as crossover. Previous work has shown that there is significant evolutionary pressure in favor of genetically robust solutions and that evolving programs adopt a number of strategies to increase genetic robustness, notably an increase in inoperative 'genes' (individual genetic units that don't influence fitness) and a preference for 'genes' with a relatively small effect on fitness.Here we examine the role of genes that cancel each other out. We find that allowing such 'canceling genes' leads to an overall increase in the rate of code growth, both through the inclusion of self-canceling code and through a general increase in introns. Finally, we find that the evolution generally follows a two-step process. Initially the operative code evolves rapidly to achieve a (near) optimal fitness. Then, the inoperative code begins to evolve most rapidly to increase robustness. In an extreme case of a problem that can be solved with no operative genes, individuals evolve by losing all operative genes and then losing all inoperative genes. Xue Zhong, Terence Soule |
GECCO | 1 |