Tao Wang 0082

dblp:12/5838-82 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-5728-6463ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 11 first-author · 21 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MFCLDTA: Multi-scale feature contrastive learning for predicting drug-target binding affinity
Zhen Tian 0004, Saisai Zhu, Zhixia Teng, Tao Wang 0082
Expert Syst. Appl.5
2026 DiffuST: A Latent Diffusion Model for Spatial Transcriptomics Denoising
abstract
Spatial transcriptomics technologies have enabled comprehensive measurements of gene expression profiles while retaining spatial information, with most platforms also providing matched pathology images. However, noise resulting from low RNA capture efficiency and experimental steps needed to keep spatial information may corrupt the biological signals and obstruct analyses. Here, we develop a latent diffusion model DiffuST to denoise spatial transcriptomics. DiffuST employs a graph autoencoder and a pre-trained model to extract different-scale features from spatial information and pathology images. Then, a latent diffusion model is leveraged to map different scales of features to the same space for denoising. The evaluation based on various spatial transcriptomics datasets showed the superiority of DiffuST over existing denoising methods. Furthermore, the results demonstrated that DiffuST can enhance downstream analysis of spatial transcriptomics and yield significant biological insights.
Shaoqing Jiao, Dazhi Lu, Tao Wang 0082, Yongtian Wang, Yunwei Dong, Jiajie Peng
IEEE Trans. Comput. Biol. Bioinform.4
2025 BioRAGent: natural language biomedical querying with retrieval-augmented multiagent systems
abstract
Understanding the roles of genes, phenotypes, and diseases is crucial for advancing biomedical research. However, efficient and accessible retrieval of biomedical knowledge remains a challenge due to the complexity of the relevant data. We introduce BioRAGent, an intelligent biomedical assistant that combines Tool-augmented retrieval-augmented generation (RAG) with a multiagent system. Leveraging the ability of large language models, BioRAGent facilitates natural language queries about genes, phenotypes, diseases, and their interrelationships. BioRAGent employs three specialized agents: Guide (query optimization), Retriever (data retrieval), and Reviewer (answer validation) to access authoritative biomedical databases and to generate accurate responses. We evaluate the performance of BioRAGent on a benchmark of eleven single-hop and three multi-hop tasks, demonstrating superior results compared with state-of-the-art models. User evaluations highlight the practicality and robust user experience of BioRAGent, particularly in handling complex multi-hop queries. Moreover, ablation experiments validate the contribution of each agent in improving retrieval accuracy.
Manlian Bi, Zhijie Bao, Dongna Xie, Xiaohan Xie, Changxiao Yang, Tao Wang 0082, Yongtian Wang, Jiajie Peng
Briefings Bioinform.6
2025 Inference of gene coexpression networks from single-cell transcriptome data based on variance decomposition analysis
abstract
Gene regulation varies across different cell types and developmental stages, leading to distinct cellular roles across cellular populations. Investigating cell type-specific gene coexpression is therefore crucial for understanding gene functions and disease pathology. However, reconstructing gene coexpression networks from single-cell transcriptome data is challenging due to artifacts, noise, and data sparsity. Here, we present an efficient method for inference of gene coexpression networks via variance decomposition analysis (GCNVDA) to explore the underlying gene regulatory mechanisms from single-cell transcriptome data. Our model incorporates multiple sources of variability, including a random effect term $G$ to capture gene-level variance and a random effect term $E$ to account for residual errors. We applied GCNVDA to three real-world single-cell datasets, demonstrating that our method outperforms existing state-of-the-art algorithms in both sensitivity and specificity for identifying tissue- or state-specific gene regulations. Furthermore, GCNVDA facilitates the discovery of functional modules that play critical roles in key biological processes such as embryonic development. These findings provide new insights into cell-specific regulatory mechanisms and have the potential to significantly advance research in developmental biology and disease pathology.
Bin Lian, Haohui Zhang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, N. Ahmad Aziz, Jialu Hu
Briefings Bioinform.3
2025 cfMethylPre: deep transfer learning enhances cancer detection based on circulating cell-free DNA methylation profiling
abstract
Cancer remains a significant global health burden, underscoring the need for innovative diagnostic tools to enable early detection and improve patient outcomes. While circulating cell-free DNA (cfDNA) methylation has emerged as a promising biomarker for noninvasive cancer diagnostics, existing methods often face limitations in handling the high-dimensionality of methylation data, small sample sizes, and a lack of biological interpretability. To address these challenges, we propose cfMethylPre, a novel deep transfer learning framework tailored for cancer detection using cfDNA methylation data. cfMethylPre leverages large language model pretrained embeddings from DNA sequence information and integrates them with methylation profiles to enhance feature representation. The deep transfer learning process involves pretraining on bulk DNA methylation data encompassing 2801 samples across 82 cancer types and normal controls, followed by fine-tuning with cfDNA methylation data. This approach ensures robust adaptation to cfDNA's unique characteristics while improving predictive accuracy. Our model achieved superior predictive accuracy compared with state-of-the-art methods, with a weighted Matthews Correlation Coefficient of 0.926 and a weighted F1-score of 0.942. Through model interpretation and biological experimental validation, we identified three novel breast cancer genes-PCDHA10, PRICKLE2, and PRTG-demonstrating their inhibitory effects on cell proliferation and migration in breast cancer cell lines. These findings establish cfMethylPre as a powerful and interpretable tool for cancer diagnostics and biological discovery, paving the way for its application in precision oncology.
Xuchao Zhang, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Yanpu Wang, Tao Wang 0082
Briefings Bioinform.9
2025 VGAE-CCI: variational graph autoencoder-based construction of 3D spatial cell-cell communication network
abstract
Cell-cell communication plays a critical role in maintaining normal biological functions, regulating development and differentiation, and controlling immune responses. The rapid development of single-cell RNA sequencing and spatial transcriptomics sequencing (ST-seq) technologies provides essential data support for in-depth and comprehensive analysis of cell-cell communication. However, ST-seq data often contain incomplete data and systematic biases, which may reduce the accuracy and reliability of predicting cell-cell communication. Furthermore, other methods for analyzing cell-cell communication mainly focus on individual tissue sections, neglecting cell-cell communication across multiple tissue layers, and fail to comprehensively elucidate cell-cell communication networks within three-dimensional tissues. To address the aforementioned issues, we propose VGAE-CCI, a deep learning framework based on the Variational Graph Autoencoder, capable of identifying cell-cell communication across multiple tissue layers. Additionally, this model can be applied to spatial transcriptomics data with missing or partially incomplete data and can clustered cells at single-cell resolution based on spatial encoding information within complex tissues, thereby enabling more accurate inference of cell-cell communication. Finally, we tested our method on six datasets and compared it with other state of art methods for predicting cell-cell communication. Our method outperformed other methods across multiple metrics, demonstrating its efficiency and reliability in predicting cell-cell communication.
Zhenao Wu, Jixiang Ren, Zhongqian Zhao, Guohua Wang 0001, Tao Wang 0082
Briefings Bioinform.8
2025 MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoder
abstract
Spatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST.
Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082
Briefings Bioinform.11
2025 Integrative Graph-Based Framework for Predicting circRNA Drug Resistance Using Disease Contextualization and Deep Learning
abstract
Circular RNAs (circRNAs) play a crucial role in gene regulation and have been implicated in the development of drug resistance in cancer, representing a significant challenge in oncological therapeutics. Despite advancements in computational models predicting RNA-drug interactions, existing frameworks often overlook the complex interplay between circRNAs, drug mechanisms, and disease contexts. This study aims to bridge this gap by introducing a novel computational model, circRDRP, that enhances prediction accuracy by integrating disease-specific contexts into the analysis of circRNA-drug interactions. It employs a hybrid graph neural network that combines features from Graph Attention Networks (GAT) and Graph Convolutional Networks (GCN) in a two-layer structure, with further enhancement through convolutional neural networks. This approach allows for sophisticated feature extraction from integrated networks of circRNAs, drugs, and diseases. Our results demonstrate that the circRDRP model outperforms existing models in predicting drug resistance, showing significant improvements in accuracy, precision, and recall. Specifically, the model shows robust predictive capability in case studies involving major anticancer drugs such as Cisplatin and Methotrexate, indicating its potential utility in precision medicine. In conclusion, circRDRP offers a powerful tool for understanding and predicting drug resistance mediated by circRNAs, with implications for designing more effective cancer therapies.
Yongtian Wang, Wenkai Shen, Yewei Shen, Shang Feng, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
IEEE J. Biomed. Health Informatics5
2024 A Scenario Model-driven Task Planning Method for Unmanned Aerial Vehicle Swarm
abstract
As the demand for smart city services grows, unmanned aerial vehicle (UAV) swarm have achieved tremendous success in industries such as traffic management, logistics transportation, and road inspection. Despite their promising potential, a critical gap exists in the domain of drone swarm mission planning-a lack of a universal task planning method that can effectively address the complexities of diverse mission scenarios. To address this challenge, this paper introduces a novel scenario model-driven task planning method for UAV swarm. This method leverages scenario models as input, enabling the parsing of scenario tasks, UAV swarm resources, and scenario constraints. It subsequently facilitates multi-constraint task allocation through auction mechanisms and path planning via reinforcement learning. Through simulation experiments conducted in scenarios such as highway inspection and campus logistics, we validate the efficacy and versatility of the proposed method across different contexts.
Yunwei Dong, Zeshan Li, Ruiheng Zhang 0002, Rubing Huang, Tao Wang 0082
Internetware5
2024 Accurately deciphering spatial domains for spatially resolved transcriptomics with stCluster
abstract
Spatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster.
Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001
Briefings Bioinform.1
2024 Scbean: a python library for single-cell multi-omics data analysis
abstract
SUMMARY: Single-cell multi-omics technologies provide a unique platform for characterizing cell states and reconstructing developmental process by simultaneously quantifying and integrating molecular signatures across various modalities, including genome, transcriptome, epigenome, and other omics layers. However, there is still an urgent unmet need for novel computational tools in this nascent field, which are critical for both effective and efficient interrogation of functionality across different omics modalities. Scbean represents a user-friendly Python library, designed to seamlessly incorporate a diverse array of models for the examination of single-cell data, encompassing both paired and unpaired multi-omics data. The library offers uniform and straightforward interfaces for tasks, such as dimensionality reduction, batch effect elimination, cell label transfer from well-annotated scRNA-seq data to scATAC-seq data, and the identification of spatially variable genes. Moreover, Scbean's models are engineered to harness the computational power of GPU acceleration through Tensorflow, rendering them capable of effortlessly handling datasets comprising millions of cells. AVAILABILITY AND IMPLEMENTATION: Scbean is released on the Python Package Index (PyPI) (https://pypi.org/project/scbean/) and GitHub (https://github.com/jhu99/scbean) under the MIT license. The documentation and example code can be found at https://scbean.readthedocs.io/en/latest/.
Haohui Zhang, Bin Lian, Xingyi Li 0003, Tao Wang 0082, Xuequn Shang 0001, Ahmad Aziz, Jialu Hu
Bioinform.6
2023 Collaborative deep learning improves disease-related circRNA prediction based on multi-source functional information
abstract
Emerging studies have shown that circular RNAs (circRNAs) are involved in a variety of biological processes and play a key role in disease diagnosing, treating and inferring. Although many methods, including traditional machine learning and deep learning, have been developed to predict associations between circRNAs and diseases, the biological function of circRNAs has not been fully exploited. Some methods have explored disease-related circRNAs based on different views, but how to efficiently use the multi-view data about circRNA is still not well studied. Therefore, we propose a computational model to predict potential circRNA-disease associations based on collaborative learning with circRNA multi-view functional annotations. First, we extract circRNA multi-view functional annotations and build circRNA association networks, respectively, to enable effective network fusion. Then, a collaborative deep learning framework for multi-view information is designed to get circRNA multi-source information features, which can make full use of the internal relationship among circRNA multi-view information. We build a network consisting of circRNAs and diseases by their functional similarity and extract the consistency description information of circRNAs and diseases. Last, we predict potential associations between circRNAs and diseases based on graph auto encoder. Our computational model has better performance in predicting candidate disease-related circRNAs than the existing ones. Furthermore, it shows the high practicability of the method that we use several common diseases as case studies to find some unknown circRNAs related to them. The experiments show that CLCDA can efficiently predict disease-related circRNAs and are helpful for the diagnosis and treatment of human disease.
Yongtian Wang, Xinmeng Liu, Yewei Shen, Xuerui Song, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.5
2023 scMultiGAN: cell-specific imputation for single-cell transcriptomes with multiple deep generative adversarial networks
abstract
The emergence of single-cell RNA sequencing (scRNA-seq) technology has revolutionized the identification of cell types and the study of cellular states at a single-cell level. Despite its significant potential, scRNA-seq data analysis is plagued by the issue of missing values. Many existing imputation methods rely on simplistic data distribution assumptions while ignoring the intrinsic gene expression distribution specific to cells. This work presents a novel deep-learning model, named scMultiGAN, for scRNA-seq imputation, which utilizes multiple collaborative generative adversarial networks (GAN). Unlike traditional GAN-based imputation methods that generate missing values based on random noises, scMultiGAN employs a two-stage training process and utilizes multiple GANs to achieve cell-specific imputation. Experimental results show the efficacy of scMultiGAN in imputation accuracy, cell clustering, differential gene expression analysis and trajectory analysis, significantly outperforming existing state-of-the-art techniques. Additionally, scMultiGAN is scalable to large scRNA-seq datasets and consistently performs well across sequencing platforms. The scMultiGAN code is freely available at https://github.com/Galaxy8172/scMultiGAN.
Tao Wang 0082, Yungang Xu, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng, Bing Xiao 0001
Briefings Bioinform.1
2023 DFinder: a novel end-to-end graph embedding-based method to identify drug-food interactions
abstract
MOTIVATION: Drug-food interactions (DFIs) occur when some constituents of food affect the bioaccessibility or efficacy of the drug by involving in drug pharmacodynamic and/or pharmacokinetic processes. Many computational methods have achieved remarkable results in link prediction tasks between biological entities, which show the potential of computational methods in discovering novel DFIs. However, there are few computational approaches that pay attention to DFI identification. This is mainly due to the lack of DFI data. In addition, food is generally made up of a variety of chemical substances. The complexity of food makes it difficult to generate accurate feature representations for food. Therefore, it is urgent to develop effective computational approaches for learning the food feature representation and predicting DFIs. RESULTS: In this article, we first collect DFI data from DrugBank and PubMed, respectively, to construct two datasets, named DrugBank-DFI and PubMed-DFI. Based on these two datasets, two DFI networks are constructed. Then, we propose a novel end-to-end graph embedding-based method named DFinder to identify DFIs. DFinder combines node attribute features and topological structure features to learn the representations of drugs and food constituents. In topology space, we adopt a simplified graph convolution network-based method to learn the topological structure features. In feature space, we use a deep neural network to extract attribute features from the original node attributes. The evaluation results indicate that DFinder performs better than other baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/23AIBox/23AIBox-DFinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Wang 0082, Jinjin Yang, Yifu Xiao, Yuxian Wang, Yongtian Wang, Jiajie Peng
Bioinform.1
2022 Exploring genetic mechanisms underlying EEG endophenotypes via summary-data-based Mendelian randomization
abstract
Oscillations in neuronal activity play a crucial role in neural processing. Electroencephalogram (EEG) records the amplitudes of oscillations at different frequencies. The robust interindividual variation and the high intraindividual stability of EEG oscillations can be explained by genetic variants. However, our knowledge of how genetic variants affect EEG endophenotypes is still limited. In this paper, we aim to identify the linkage between genes and EEG oscillations at different frequencies. We integrate the summary-level datasets from the genome-wide association studies (GWASs) of seven EEG endophenotypes and expression quantitative trait loci (eQTL) study of the human brain, to explore the genetic mechanisms underlying EEG using summary-data-based two-sample Mendelian randomization. As a result, twelve candidate causal genes are prioritized across four EEG endophenotypes. These genes offer opportunities for a deeper understanding of the molecular mechanisms whereby genomic variation leads to variable brain activities.
Jinghui Pan, Tao Wang 0082
BIBM4
2022 Hypergraph-based Gene Ontology Embedding for Disease Gene Prediction
abstract
Disease gene identification has provided valuable insights into illuminating the molecular mechanisms underlying complex diseases. And it has been shown that novel drugs with genetically supported targets were more likely to be successful in clinical trials. In recent years, multiple graph machine learning-based methods for this purpose have been proposed. However, those methods were mainly based on various well-established biological molecular networks, while seldomly considering the curated biological annotations of genes. To fill this gap, we aim to integrate the gene ontology annotations (GOA), including the biological process (BP), the cellular component (CC), and the molecular function (MF), into the process of disease gene prediction. Our method treated the GOA as a hypergraph and used the hypergraph-based embedding technique to extract the deep features underlying gene annotations. Besides, we also extracted gene features from the protein-protein interaction (PPI) network using graph representation learning methods. The convolutional neural network (CNN) framework was followed to fuse the features extracted from the two networks and make the final prediction. Experiments on a range of diseases have demonstrated the accuracy and robustness of our method. The average area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC) reached 0.85 and 0.79, respectively. Besides, the hypergraph-based gene ontology embedding can be generalized to other bioinformatics applications.
Tao Wang 0082, Hengbo Xu, Ranye Zhang, Yifu Xiao, Jiajie Peng, Xuequn Shang 0001
BIBM1
2022 Discovering eQTL Regulatory Patterns Through eQTLMotif
abstract
The expression quantitative trait loci (eQTL) analysis has become important for understanding the regulatory function of genomic variants on gene expression in a tissuespecific manner and has been widely applied across species from microbes to mammals. Current eQTL studies mainly focus on the simple one-to-one regulation between variant and gene. Recent research have demonstrated there are also more complex regulatory patterns between eQTLs and genes. However, there is a lack of studies and relevant methods to systematically discover the regulatory patterns between multiple eQTLs and multiple genes. In this regard, this study has proposed a novel computational framework, called eQTLMotif, to discover regulation patterns of eQTLs in a many-to-many manner. This framework mainly consists of two steps: (1) construct a novel eQTL regulatory network by integrating bipartite eQTL network, eQTL mediation effects, and gene regulatory network; (2) perform motif mining through exactly enumerating frequently appeared eQTL regulatory structures. Based on this framework, we for the first time systematically investigated the eQTL regulatory patterns in the human frontal cortex based on a large cohort of postmortem human brains. Experiments have demonstrated that our framework can effectively reveal novel eQTL regulatory patterns. And some are in similar structure to the existing gene regulation patterns, such as feed-forward loop (FFL)-like motif, single input module (SIM)-like motif, and dense overlapping regulons (DOR)- like motif. Our method and findings will further enhance the understanding of regulatory mechanisms of eQTLs in multiple tissues and species.
Tao Wang 0082, Yifu Xiao, Hanzi Yang, Xipeng Yin, Yongtian Wang, Bing Xiao 0001, Xuequn Shang 0001, Jiajie Peng
BIBM1
2022 A network-based method for brain disease gene prediction by integrating brain connectome and molecular network
abstract
Brain disease gene identification is critical for revealing the biological mechanism and developing drugs for brain diseases. To enhance the identification of brain disease genes, similarity-based computational methods, especially network-based methods, have been adopted for narrowing down the searching space. However, these network-based methods only use molecular networks, ignoring brain connectome data, which have been widely used in many brain-related studies. In our study, we propose a novel framework, named brainMI, for integrating brain connectome data and molecular-based gene association networks to predict brain disease genes. For the consistent representation of molecular-based network data and brain connectome data, brainMI first constructs a novel gene network, called brain functional connectivity (BFC)-based gene network, based on resting-state functional magnetic resonance imaging data and brain region-specific gene expression data. Then, a multiple network integration method is proposed to learn low-dimensional features of genes by integrating the BFC-based gene network and existing protein-protein interaction networks. Finally, these features are utilized to predict brain disease genes based on a support vector machine-based model. We evaluate brainMI on four brain diseases, including Alzheimer's disease, Parkinson's disease, major depressive disorder and autism. brainMI achieves of 0.761, 0.729, 0.728 and 0.744 using the BFC-based gene network alone and enhances the molecular network-based performance by 6.3% on average. In addition, the results show that brainMI achieves higher performance in predicting brain disease genes compared to the existing three state-of-the-art methods.
Ruijiang Han, Menghan Zhang, Yuxian Wang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.5
2022 Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
Quantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.1
2022 Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
In the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.1
2022 A review and performance evaluation of clustering frameworks for single-cell Hi-C data
abstract
The three-dimensional genome structure plays a key role in cellular function and gene regulation. Single-cell Hi-C (high-resolution chromosome conformation capture) technology can capture genome structure information at the cell level, which provides the opportunity to study how genome structure varies among different cell types. Recently, a few methods are well designed for single-cell Hi-C clustering. In this manuscript, we perform an in-depth benchmark study of available single-cell Hi-C data clustering methods to implement an evaluation system for multiple clustering frameworks based on both human and mouse datasets. We compare eight methods in terms of visualization and clustering performance. Performance is evaluated using four benchmark metrics including adjusted rand index, normalized mutual information, homogeneity and Fowlkes-Mallows index. Furthermore, we also evaluate the eight methods for the task of separating cells at different stages of the cell cycle based on single-cell Hi-C data.
Caiwei Zhen, Yuxian Wang, Jiaquan Geng, Jinghao Peng, Tao Wang 0082, Jianye Hao, Xuequn Shang 0001, Zhongyu Wei, Peican Zhu, Jiajie Peng
Briefings Bioinform.7
2021 Predicting Hepatoma-Related Genes Based on Representation Learning of PPI network and Gene Ontology Annotations
abstract
Hepatoma is the most common type of primary liver cancer with a high mortality rate in the world. The genetic causes of the disease pathology remain largely unknown. Effective discovery of the genes associated with hepatoma has become important in disease prevention, early diagnosis, and therapeutic treatments. With the developments of molecular networks, graph-based methods have been tremendously successful in predicting disease genes based on the hypothesis of guilt-by-association. Network representation learning (NRL) techniques have accelerated disease gene discovery in recent years because of their powerful network feature extraction ability. However, the current network representation learning-based methods for disease gene discovery did not consider the gene features derived from gene ontology annotations, which apriori group genes with similar functions. To fill this gap, here we propose a novel framework to predict hepatoma-related genes based on representation learning from both protein-protein interactions (PPI) network and gene ontology annotations. Our framework has three steps: learning features from PPI network and gene ontologies using NRL techniques, integrating different features based on autoencoder, predicting hepatoma-related genes using machine learning classifiers. Experiments have demonstrated that our framework could accurately predict hepatoma-related genes with AUROC and AUPRC reaching 0.93 and 0.94, respectively. Compared with other methods using only single representation features, our framework also shows superior performance on hepatoma gene prediction.
Tao Wang 0082, Zhiyuan Shao, Yifu Xiao, Xuchao Zhang, Binze Shi, Siyu Chen 0024, Yuxian Wang, Jiajie Peng, Xuequn Shang 0001
BIBM1
2021 A pipeline for RNA-seq based eQTL analysis with automated quality control procedures
abstract
BACKGROUND: Advances in the expression quantitative trait loci (eQTL) studies have provided valuable insights into the mechanism of diseases and traits-associated genetic variants. However, it remains challenging to evaluate and control the quality of multi-source heterogeneous eQTL raw data for researchers with limited computational background. There is an urgent need to develop a powerful and user-friendly tool to automatically process the raw datasets in various formats and perform the eQTL mapping afterward. RESULTS: In this work, we present a pipeline for eQTL analysis, termed eQTLQC, featured with automated data preprocessing for both genotype data and gene expression data. Our pipeline provides a set of quality control and normalization approaches, and utilizes automated techniques to reduce manual intervention. We demonstrate the utility and robustness of this pipeline by performing eQTL case studies using multiple independent real-world datasets with RNA-seq data and whole genome sequencing (WGS) based genotype data. CONCLUSIONS: eQTLQC provides a reliable computational workflow for eQTL analysis. It provides standard quality control and normalization as well as eQTL mapping procedures for eQTL raw data in multiple formats. The source code, demo data, and instructions are freely available at https://github.com/stormlovetao/eQTLQC .
Tao Wang 0082, Yongzhuang Liu, Junpeng Ruan, Xianjun Dong, Yadong Wang 0001, Jiajie Peng
BMC Bioinform.1
2019 An automated quality control pipeline for eQTL analysis with RNA-seq data
abstract
Expression quantitative trait loci (eQTL) analysis is of critical importance to understand the mechanism underlying trait associated variants. Evaluating and controlling the data quality of transcripts and genotypes, which are basis of eQTL analysis, remains challenging for researchers with limited computational backgrounds. There is a strong need for a user-friendly and comprehensive tool to pre-process those data sets automatically. Here we propose such a solution, eQTLQC, an automated quality control pipeline for preprocessing both RNA-seq and genotype data. The eQTLQC pipeline provides multiple informative quality control measurements and data normalization approaches. And it provides a easy-to-use configuration file for users to flexibly set up the parameters and control the pipeline. We demonstrate its utility by performing RNA-seq and genotype preprocessing on real data sets. eQTLQC is open source and freely available at https://github.com/ruanjunpeng/eQTLQC.
Tao Wang 0082, Junpeng Ruan, Quanwei Yin, Xianjun Dong, Yadong Wang 0001
BIBM1
2019 TS-GOEA: a web tool for tissue-specific gene set enrichment analysis based on gene ontology
abstract
BACKGROUND: The Gene Ontology (GO) knowledgebase is the world's largest source of information on the functions of genes. Since the beginning of GO project, various tools have been developed to perform GO enrichment analysis experiments. GO enrichment analysis has become a commonly used method of gene function analysis. Existing GO enrichment analysis tools do not consider tissue-specific information, although this information is very important to current research. RESULTS: In this paper, we built an easy-to-use web tool called TS-GOEA that allows users to easily perform experiments based on tissue-specific GO enrichment analysis. TS-GOEA uses strict threshold statistical method for GO enrichment analysis, and provides statistical tests to improve the reliability of the analysis results. Meanwhile, TS-GOEA provides tools to compare different experimental results, which is convenient for users to compare the experimental results. To evaluate its performance, we tested the genes associated with platelet disease with TS-GOEA. CONCLUSIONS: TS-GOEA is an effective GO analysis tool with unique features. The experimental results show that our method has better performance and provides a useful supplement for the existing GO enrichment analysis tools. TS-GOEA is available at http://120.77.47.2:5678.
Jiajie Peng, Guilin Lu, Hansheng Xue, Tao Wang 0082, Xuequn Shang 0001
BMC Bioinform.4
2018 TSGOE: A web tool for tissue-specific gene ontology enrichment
Jiajie Peng, Guilin Lu, Hansheng Xue, Tao Wang 0082, Xuequn Shang 0001
BIBM4
2018 Identifying Representative Network Motifs for Inferring Higher-order Structure of Biological Networks
Tao Wang 0082, Jiajie Peng, Yadong Wang 0001, Jin Chen 0004
BIBM1
2016 Extending gene ontology with gene association networks
abstract
MOTIVATION: Gene ontology (GO) is a widely used resource to describe the attributes for gene products. However, automatic GO maintenance remains to be difficult because of the complex logical reasoning and the need of biological knowledge that are not explicitly represented in the GO. The existing studies either construct whole GO based on network data or only infer the relations between existing GO terms. None is purposed to add new terms automatically to the existing GO. RESULTS: We proposed a new algorithm 'GOExtender' to efficiently identify all the connected gene pairs labeled by the same parent GO terms. GOExtender is used to predict new GO terms with biological network data, and connect them to the existing GO. Evaluation tests on biological process and cellular component categories of different GO releases showed that GOExtender can extend new GO terms automatically based on the biological network. Furthermore, we applied GOExtender to the recent release of GO and discovered new GO terms with strong support from literature. AVAILABILITY AND IMPLEMENTATION: Software and supplementary document are available at www.msu.edu/%7Ejinchen/GOExtender CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiajie Peng, Tao Wang 0082, Jixuan Wang, Yadong Wang 0001, Jin Chen 0004
Bioinform.2