Quan Wang 0004

dblp:86/5728-4 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-7056-0519ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 THOR: a TMB heterogeneity-adaptive optimization model predicts immunotherapy response using clonal genomic features in group-structured data
abstract
With the increasing number of indications for immune checkpoint inhibitors in early and advanced cancers, the prospect of a tumor-agnostic biomarker to prioritize patients is compelling. Tumor mutation burden (TMB) is a widely endorsed biomarker that quantifies nonsynonymous mutations within tumor DNA, essential for neoantigen production, which, in turn, correlates with the immune response and guides decision-making. However, the general clinical application of TMB-relying on simple mutational counts targeted at a single endpoint-does not adequately capture the complex clonal structure of tumors nor the multifaceted nature of prognostic indicators. This recognition has spurred the exploration of sophisticated high-dimensional regression techniques. Unfortunately, the limited cohort sizes in immunotherapy trials have hindered the full potential of these advanced methods. Our approach considers patient subgroups as related yet distinct entities, enabling precise tailoring and refinement to address subgroup-specific dynamics. Given the deficiencies and the constraints, we introduce a TMB heterogeneity-optimized regression (THOR). This innovative model enhances the predictive capabilities of TMB by integrating tumor clonality and a diverse spectrum of clinical endpoints, further augmented by fusion techniques across subgroups to facilitate robust data sharing and interpretation. Our simulations validate THOR's superiority in parameter estimation for statistical inference. Clinically, we assess the utility of THOR in a structured cohort of 238 cancer patients undergoing immunotherapy, supplemented by 2212 patients across 19 subgroups from public datasets. The forecast of the responses and comparison of survival hazards demonstrate that THOR significantly enhances patient stratification and prognostic predictions by incorporating complex immunogenetic biology and subgroup-specific dynamics.
Yanfang Guan, Xin Lai 0003, Yuqian Liu, Zhili Chang, Quan Wang 0004, Jian Zhao 0034, Shuanying Yang, Jiayin Wang 0002
Briefings Bioinform.7
2024 DeepIRES: a hybrid deep learning model for accurate identification of internal ribosome entry sites in cellular and viral mRNAs
abstract
The internal ribosome entry site (IRES) is a cis-regulatory element that can initiate translation in a cap-independent manner. It is often related to cellular processes and many diseases. Thus, identifying the IRES is important for understanding its mechanism and finding potential therapeutic strategies for relevant diseases since identifying IRES elements by experimental method is time-consuming and laborious. Many bioinformatics tools have been developed to predict IRES, but all these tools are based on structure similarity or machine learning algorithms. Here, we introduced a deep learning model named DeepIRES for precisely identifying IRES elements in messenger RNA (mRNA) sequences. DeepIRES is a hybrid model incorporating dilated 1D convolutional neural network blocks, bidirectional gated recurrent units, and self-attention module. Tenfold cross-validation results suggest that DeepIRES can capture deeper relationships between sequence features and prediction results than other baseline models. Further comparison on independent test sets illustrates that DeepIRES has superior and robust prediction capability than other existing methods. Moreover, DeepIRES achieves high accuracy in predicting experimental validated IRESs that are collected in recent studies. With the application of a deep learning interpretable analysis, we discover some potential consensus motifs that are related to IRES activities. In summary, DeepIRES is a reliable tool for IRES prediction and gives insights into the mechanism of IRES elements.
Jian Zhao 0034, Zhewei Chen, Meng Zhang 0046, Lingxiao Zou, Quan Wang 0004
Briefings Bioinform.7
2023 A Statistical Explainable Learning Model Optimizing Co-localization of Multidimensional Positivity Thresholds in Immunotherapy Decision-Supporting
abstract
Tumor Mutation Burden (TMB) serves as a recognized stratified biomarker for immunotherapy. However, its one-dimensional representation of non-synonymous genetic alterations has been contentious. Specifically, the uniform quantification of mutations by TMB, coupled with measurement inaccuracies, complicates the accurate determination of a positive threshold for classifying patients. Parallel to this, assessing immunotherapy benefits requires the joint analysis of multiscale endpoints, namely discrete tumor response and sequential time-to-event, presenting a pressing challenge for clinical computation. Recognizing the intertwined nature of these challenges, we address the inter-sample bias inherent in multidimensional mutation biomarkers within the framework of multiscale endpoint fusion analysis, aiming for a more robust and comprehensive patient stratification. By combining the concept of corrected-score with a soft-threshold strategy, and utilizing the attention mechanism alongside the multiple instance learning, we propose a statistically explainable learning model optimizing co-localization of multidimensional positivity thresholds for immunotherapy categorical decision-supporting.
Jian Zhao 0034, Jiayin Wang 0002, Quan Wang 0004
BIBM5
2022 TVAR: assessing tissue-specific functional effects of non-coding variants with deep learning
abstract
MOTIVATION: Analysis of whole-genome sequencing (WGS) for genetics is still a challenge due to the lack of accurate functional annotation of non-coding variants, especially the rare ones. As eQTLs have been extensively implicated in the genetics of human diseases, we hypothesize that rare non-coding variants discovered in WGS play a regulatory role in predisposing disease risk. RESULTS: With thousands of tissue- and cell-type-specific epigenomic features, we propose TVAR. This multi-label learning-based deep neural network predicts the functionality of non-coding variants in the genome based on eQTLs across 49 human tissues in the GTEx project. TVAR learns the relationships between high-dimensional epigenomics and eQTLs across tissues, taking the correlation among tissues into account to understand shared and tissue-specific eQTL effects. As a result, TVAR outputs tissue-specific annotations, with an average AUROC of 0.77 across these tissues. We evaluate TVAR's performance on four complex diseases (coronary artery disease, breast cancer, Type 2 diabetes and Schizophrenia), using TVAR's tissue-specific annotations, and observe its superior performance in predicting functional variants for both common and rare variants, compared with five existing state-of-the-art tools. We further evaluate TVAR's G-score, a scoring scheme across all tissues, on ClinVar, fine-mapped GWAS loci, Massive Parallel Reporter Assay (MPRA) validated variants and observe the consistently better performance of TVAR compared with other competing tools. AVAILABILITY AND IMPLEMENTATION: The TVAR source code and its scores on the ClinVar catalog, fine mapped GWAS Loci, high confidence eQTLs from GTEx dataset, and MPRA validated functional variants are available at https://github.com/haiyang1986/TVAR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Xue Zhong, Bingshan Li
Bioinform.3
2022 A Bayesian framework to integrate multi-level genome-scale data for Autism risk gene prioritization
abstract
BACKGROUND: Autism spectrum disorder (ASD) is a group of complex neurodevelopment disorders with a strong genetic basis. Large scale sequencing studies have identified over one hundred ASD risk genes. Nevertheless, the vast majority of ASD risk genes remain to be discovered, as it is estimated that more than 1000 genes are likely to be involved in ASD risk. Prioritization of risk genes is an effective strategy to increase the power of identifying novel risk genes in genetics studies of ASD. As ASD risk genes are likely to exhibit distinct properties from multiple angles, we reason that integrating multiple levels of genomic data is a powerful approach to pinpoint genuine ASD risk genes. RESULTS: We present BNScore, a Bayesian model selection framework to probabilistically prioritize ASD risk genes through explicitly integrating evidence from sequencing-identified ASD genes, biological annotations, and gene functional network. We demonstrate the validity of our approach and its improved performance over existing methods by examining the resulting top candidate ASD risk genes against sets of high-confidence benchmark genes and large-scale ASD genome-wide association studies. We assess the tissue-, cell type- and development stage-specific expression properties of top prioritized genes, and find strong expression specificity in brain tissues, striatal medium spiny neurons, and fetal developmental stages. CONCLUSIONS: In summary, we show that by integrating sequencing findings, functional annotation profiles, and gene-gene functional network, our proposed BNScore provides competitive performance compared to current state-of-the-art methods in prioritizing ASD genes. Our method offers a general and flexible strategy to risk gene prioritization that can potentially be applied to other complex traits as well.
Ying Ji 0002, Rui Chen 0021, Quan Wang 0004, Bingshan Li
BMC Bioinform.3
2020 DRAMS: A tool to detect and re-align mixed-up samples for integrative studies of multi-omics data
abstract
Studies of complex disorders benefit from integrative analyses of multiple omics data. Yet, sample mix-ups frequently occur in multi-omics studies, weakening statistical power and risking false findings. Accurately aligning sample information, genotype, and corresponding omics data is critical for integrative analyses. We developed DRAMS (https://github.com/Yi-Jiang/DRAMS) to Detect and Re-Align Mixed-up Samples to address the sample mix-up problem. It uses a logistic regression model followed by a modified topological sorting algorithm to identify the potential true IDs based on data relationships of multi-omics. According to tests using simulated data, the more types of omics data used or the smaller the proportion of mix-ups, the better that DRAMS performs. Applying DRAMS to real data from the PsychENCODE BrainGVEX project, we detected and corrected 201 (12.5% of total data generated) mix-ups. Of the 21 mix-ups involving errors of racial identity, DRAMS re-assigned all data to the correct racial group in the 1000 Genomes project. In doing so, quantitative trait loci (QTL) (FDR<0.01) increased by an average of 1.62-fold. The use of DRAMS in multi-omics studies will strengthen statistical power of the study and improve quality of the results. Even though very limited studies have multi-omics data in place, we expect such data will increase quickly with the needs of DRAMS.
Gina Giase, Kay Grennan, Annie W. Shieh, Lide Han, Quan Wang 0004, Rui Chen 0021, Kevin P. White, Chao Chen 0041, Bingshan Li, Chunyu Liu 0001
PLoS Comput. Biol.7
2019 De novo pattern discovery enables robust assessment of functional consequences of non-coding variants
abstract
MOTIVATION: Given the complexity of genome regions, prioritize the functional effects of non-coding variants remains a challenge. Although several frameworks have been proposed for the evaluation of the functionality of non-coding variants, most of them used 'black boxes' methods that simplify the task as the pathogenicity/benign classification problem, which ignores the distinct regulatory mechanisms of variants and leads to less desirable performance. In this study, we developed DVAR, an unsupervised framework that leverage various biochemical and evolutionary evidence to distinguish the gene regulatory categories of variants and assess their comprehensive functional impact simultaneously. RESULTS: DVAR performed de novo pattern discovery in high-dimensional data and identified five regulatory clusters of non-coding variants. Leveraging the new insights into the multiple functional patterns, it measures both the between-class and the within-class functional implication of the variants to achieve accurate prioritization. Compared to other two-class learning methods, it showed improved performance in identification of clinically significant variants, fine-mapped GWAS variants, eQTLs and expression-modulating variants. Moreover, it has superior performance on disease causal variants verified by genome-editing (like CRISPR-Cas9), which could provide a pre-selection strategy for genome-editing technologies across the whole genome. Finally, evaluated in BioVU and UK Biobank, two large-scale DNA biobanks linked to complete electronic health records, DVAR demonstrated its effectiveness in prioritizing non-coding variants associated with medical phenotypes. AVAILABILITY AND IMPLEMENTATION: The C++ and Python source codes, the pre-computed DVAR-cluster labels and DVAR-scores across the whole genome are available at https://www.vumc.org/cgg/dvar. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hai Yang 0002, Rui Chen 0021, Quan Wang 0004, Ying Ji 0002, Guangze Zheng 0002, Xue Zhong, Nancy J. Cox, Bingshan Li
Bioinform.3
2016 Systematic dissection of dysregulated transcription factor-miRNA feed-forward loops across tumor types
abstract
Transcription factor and microRNA (miRNA) can mutually regulate each other and jointly regulate their shared target genes to form feed-forward loops (FFLs). While there are many studies of dysregulated FFLs in a specific cancer, a systematic investigation of dysregulated FFLs across multiple tumor types (pan-cancer FFLs) has not been performed yet. In this study, using The Cancer Genome Atlas data, we identified 26 pan-cancer FFLs, which were dysregulated in at least five tumor types. These pan-cancer FFLs could communicate with each other and form functionally consistent subnetworks, such as epithelial to mesenchymal transition-related subnetwork. Many proteins and miRNAs in each subnetwork belong to the same protein and miRNA family, respectively. Importantly, cancer-associated genes and drug targets were enriched in these pan-cancer FFLs, in which the genes and miRNAs also tended to be hubs and bottlenecks. Finally, we identified potential anticancer indications for existing drugs with novel mechanism of action. Collectively, this study highlights the potential of pan-cancer FFLs as a novel paradigm in elucidating pathogenesis of cancer and developing anticancer drugs.
Wei Jiang 0023, Ramkrishna Mitra, Chen-Ching Lin, Quan Wang 0004, Feixiong Cheng, Zhongming Zhao
Briefings Bioinform.4
2015 EW_dmGWAS: edge-weighted dense module search for genome-wide association studies and gene expression profiles
abstract
Abstract Summary: We previously developed dmGWAS to search for dense modules in a human protein–protein interaction (PPI) network; it has since become a popular tool for network-assisted analysis of genome-wide association studies (GWAS). dmGWAS weights nodes by using GWAS signals. Here, we introduce an upgraded algorithm, EW_dmGWAS, to boost GWAS signals in a node- and edge-weighted PPI network. In EW_dmGWAS, we utilize condition-specific gene expression profiles for edge weights. Specifically, differential gene co-expression is used to infer the edge weights. We applied EW_dmGWAS to two diseases and compared it with other relevant methods. The results suggest that EW_dmGWAS is more powerful in detecting disease-associated signals. Availability and implementation: The algorithm of EW_dmGWAS is implemented in the R package dmGWAS_3.0 and is available at http://bioinfo.mc.vanderbilt.edu/dmGWAS. Contact: [email protected] or [email protected] Supplementary information: Supplementary materials are available at Bioinformatics online.
Quan Wang 0004, Zhongming Zhao, Peilin Jia
Bioinform.1
2013 Computational tools for copy number variation (CNV) detection using next-generation sequencing data: features and perspectives
abstract
Copy number variation (CNV) is a prevalent form of critical genetic variation that leads to an abnormal number of copies of large genomic regions in a cell. Microarray-based comparative genome hybridization (arrayCGH) or genotyping arrays have been standard technologies to detect large regions subject to copy number changes in genomes until most recently high-resolution sequence data can be analyzed by next-generation sequencing (NGS). During the last several years, NGS-based analysis has been widely applied to identify CNVs in both healthy and diseased individuals. Correspondingly, the strong demand for NGS-based CNV analyses has fuelled development of numerous computational methods and tools for CNV detection. In this article, we review the recent advances in computational methods pertaining to CNV detection using whole genome and whole exome sequencing data. Additionally, we discuss their strengths and weaknesses and suggest directions for future development.
Min Zhao 0006, Qingguo Wang, Quan Wang 0004, Peilin Jia, Zhongming Zhao
BMC Bioinform.3
2009 Searching for bidirectional promoters in Arabidopsis thaliana
abstract
BACKGROUND: A "bidirectional gene pair" is defined as two adjacent genes which are located on opposite strands of DNA with transcription start sites (TSSs) not more than 1000 base pairs apart and the intergenic region between two TSSs is commonly designated as a putative "bidirectional promoter". Individual examples of bidirectional gene pairs have been reported for years, as well as a few genome-wide analyses have been studied in mammalian and human genomes. However, no genome-wide analysis of bidirectional genes for plants has been done. Furthermore, the exact mechanism of this gene organization is still less understood. RESULTS: We conducted comprehensive analysis of bidirectional gene pairs through the whole Arabidopsis thaliana genome and identified 2471 bidirectional gene pairs. The analysis shows that bidirectional genes are often coexpressed and tend to be involved in the same biological function. Furthermore, bidirectional gene pairs associated with similar functions seem to have stronger expression correlation. We pay more attention to the regulatory analysis on the intergenic regions between bidirectional genes. Using a hierarchical stochastic language model (HSL) (which is developed by ourselves), we can identify intergenic regions enriched of regulatory elements which are essential for the initiation of transcription. Finally, we picked 27 functionally associated bidirectional gene pairs with their intergenic regions enriched of regulatory elements and hypothesized them to be regulated by bidirectional promoters, some of which have the same orthologs in ancient organisms. More than half of these bidirectional gene pairs are further supported by sharing similar functional categories as these of handful experimental verified bidirectional genes. CONCLUSION: Bidirectional gene pairs are concluded also prevalent in plant genome. Promoter analyses of the intergenic regions between bidirectional genes could be a new way to study the bidirectional gene structure, which may provide a important clue for further analysis. Such a method could be applied to other genomes.
Quan Wang 0004, Dayong Li, Lihuang Zhu, Minping Qian, Minghua Deng
BMC Bioinform.1