VLDB 2026 Research / reviewers in the wild / expert
Yadong Wang 0001
dblp:68/782-1
· DBLP profile ↗
121ranked-venue papers
1as first author
61since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 118 · 1 first-author · 60 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MZSGO: multimodal zero-shot protein function annotation via evolutionary signals and textual semanticsabstractMOTIVATION: Although deep learning has significantly advanced the field of protein function prediction, current approaches are limited by their reliance on a narrow set of modalities. Specifically, they primarily rely on sequence patterns and treat protein domain data and functional labels merely as categorical tags. Consequently, they fail to capitalize on the semantic richness embedded within their textual definitions. These constraints hinder their ability to generalize to novel labels. To tackle this issue, we present MZSGO, a multimodal zero-shot framework that fuses evolutionary signals from protein language models with semantic features derived from large language models (LLMs). By employing an adaptive gated fusion mechanism, MZSGO effectively aligns sequence-based and text-based modalities to enable robust predictions for unseen labels. RESULTS: By unifying protein representations and functional annotations, we bridge the semantic gap that limits current approaches. Results indicate that while our model remains competitive on supervised benchmarks, it demonstrates a marked advantage over existing methods in zero-shot tasks. It specifically excels at recognizing previously unseen long-tail and novel Gene Ontology (GO) terms. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/toxic-byte/MZSGO. Boyue Cui, Yujuan Li, Shiqu Chen, Jiaming Wei, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 6 |
| 2026 | cuteSV-OL: a real-time structural variation detection framework for nanopore sequencing devicesabstractSUMMARY: Nanopore sequencing technology enables real-time sequencing and is widely used in rapid detection applications. However, in clinical scenarios, existing structural variant (SV) detection tools typically separate sequencing from computation, limiting their timeliness for clinical applications. To address this, we introduce cuteSV-OL, a novel framework designed for real-time SV discovery, which can be embedded within nanopore sequencing instruments to analyze data concurrently with its generation. Additionally, cuteSV-OL features a real-time SV detection rate evaluation module, allowing users to terminate sequencing early when appropriate, thereby reducing time and cost. Experimental results show that on a standard desktop computer, cuteSV-OL can perform real-time analysis during sequencing and complete SV calling within min after sequencing ends, achieving performance comparable to offline methods. This approach has the potential to enhance rapid clinical diagnostics. AVAILABILITY AND IMPLEMENTATION: cuteSV-OL is released under the MIT license and is available at https://github.com/gwmHIT/cuteSV-OL. It can also be installed via Bioconda or accessed through https://doi.org/10.5281/zenodo.17777436. Weimin Guo, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
Bioinform. | 3 |
| 2026 | GR2ST: spatial transcriptomics prediction based on graph-enhanced multimodal contrastive learningabstractMOTIVATION: Spatial transcriptomics techniques capture gene expression data and spatial coordinates, while simultaneously correlating them with tissue section images. This advantage makes Spatial transcriptomics data highly valuable for research, such as investigating disease mechanisms and cancer prognosis. However, the extended time and high cost of spatial transcriptomic sequencing currently limit further advancements in this field. The development of numerous deep learning methods aimed at predicting spatial transcriptomics from histology images has advanced significantly. However, these approaches often lack the ability to effectively integrate histology images with spatial transcriptomic data. Here, we propose GR2ST, a deep learning model that learns the underlying connections between image features and gene expression to predict spatial transcriptomics. RESULTS: GR2ST leverages a large pre-trained pathology model to extract high-level histological features. We designed a dual-branch graph architecture, consisting of a dynamic threshold-based functional graph and a radius-constrained spatial graph, to capture complex spot interactions within heterogeneous tissues. The model aligns histology images with gene expression representations through a multimodal contrastive learning framework. It achieves adaptive gene expression generation via a Cell-Type Guided Multi-Branch Regression Head supervised by a context-aware weighting network, which is further integrated with cross-sample retrieval to construct an ensemble prediction. The performance of the model is evaluated on three cancer-related spatial transcriptomics datasets, including cutaneous squamous cell carcinoma and two human breast cancer cohorts, to demonstrate its effectiveness and robustness. AVAILABILITY: https://github.com/zjl1109294570/GR2ST. Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 5 |
| 2025 | RVC: A Real-Time Variant Calling Framework for Short-Read Sequencing DataabstractAccurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics. Miao Cui 0005, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023, Guohua Wang 0001, Yadong Liu 0001 |
BIBM | 4 |
| 2025 | MethSV: A Long-Read-Based Framework for Profiling DNA Methylation in Structural VariationabstractDNA methylation plays a crucial role in regulating diverse molecular processes in living organisms. However, methylation patterns associated with structural variations (SVs) remain poorly understood due to technical limitations. with the advent of long-read sequencing technologies, it is now feasible to simultaneously detect SVs and methylation at single-molecule resolution. Here, we propose methSV, a novel framework to profile DNA methylation within SV regions, revealing potential interactions between genetic and epigenetic regulation. Using methSV, we achieved a$\sim 3.7$-fold increase in the detection of SV-associated population-level differentially methylated regions (pDMRs) and \~{}4.2-fold increase in high-confidence haplotype-resolved differentially methylated regions (hDMRs). Notably, distinct methylation signatures in high-type deletions (DELs) and low-type insertions (INSs) suggested regulatory potential and were enriched in disease-associated genes such as GP1BA, CHRNE, and HCN2. This framework extends the analytical capabilities of long-read epigenomic studies and offers new insights into the functional consequences of SV. Weize Kong, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 4 |
| 2025 | Bameth: A Bilstm-Cross Attention Network for Methylation-Based Progression-Level Classification of Colon AdenocarcinomaabstractColon adenocarcinoma (COAD) is a prevalent malignancy with high morbidity and mortality, largely due to its asymptomatic onset and late diagnosis. DNA methylation has emerged as a promising biomarker for cancer progression and molecular classification. To capture the stage-based methylation patterns, we propose BAMeth, a novel deep learning framework combining bidirectional Long ShortTerm Memory (BiLSTM) and cross-attention mechanisms. In this study, COAD samples were categorized into four types based on clinical stages and pathological characteristics, representing early (Type I), intermediate (Type II), late (Type III), and normal (Type IV) conditions. BAMeth performs multi-class classification across these stage-based categories using genome-wide methylation profiles. The architecture integrates dimensionality reduction, sequential modeling, and inter-feature dependency learning, enabling extraction of both sequential and contextual dependencies from highdimensional methylation data. Experimental results demonstrate that BAMeth achieves superior performance compared with some existing methods, particularly in terms of accuracy (ACC) by leveraging differentially methylated positions. Functional enrichment analysis of the genes annotated to the key CpG sites further reveals biological relevance in Wnt signaling and ubiquitin-mediated proteolysis pathways, highlighting its interpretability. Overall, BAMeth provides a robust and biologically interpretable framework for stage-based molecular classification of colon adenocarcinoma, offering potential for early detection and progression monitoring in epigenomics-driven oncology research. Yue Liu 0034, Zhongyu Liu, Tao Jiang 0014, Yadong Wang 0001, Yadong Liu 0001 |
BIBM | 4 |
| 2025 | SimPG: A Pangenome-Guided Population-Specific Human Genome Simulation ToolabstractThe rapid advancement of high-throughput genome sequencing has enabled large-scale reconstruction of genome sequences at both individual and population levels. However, accurately evaluating genome assemblies and associated variants remains a significant challenge due to sequencing errors, assembly artifacts, and limited experimental validation, which hinder the development of novel algorithms, benchmarking of genomic pipelines, and validation of biological hypotheses. Existing linear genome simulation tools, typically based on a single reference genome, offer limited biological realism and fail to capture the genomic diversity and structural complexity needed for comprehensive evaluation. To address these limitations, we introduce SimPG, a novel simulation framework that generates individual genomes with populationlevel characteristics by leveraging the rich variant and structural information embedded in pangenomes. SimPG produces realistic, high-quality simulated genomes that support diverse applications such as structural variant detection and population genetics research. By providing a reproducible, controllable, and standardized environment, SimPG facilitates the development, testing, and benchmarking of genome analysis tools, enabling robust performance assessment and error analysis, and ultimately advancing computational genomics and genome biology. Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 4 |
| 2025 | GDTGO: Advancing Protein Function Prediction via Graph Convolutional Network and Iterative OptimizationabstractFunction annotation of proteins is fundamental for revealing the nature of life phenomena and understanding the mechanisms of disease. Several computational methods have been developed to predict protein functions. However, existing methods ignore the prior knowledge among Gene ontology (GO) terms and simply regard the prediction problem as an independent multilabel classification task. To mitigate these issues, we propose a novel predictor based on graph representation learning and iterative optimization, named GDTGO, to explore the potential protein function on the GO. GDTGO employs a graph convolutional network to learn semantically rich and topologically aware knowledge from GO terms, which serves as a strong prior to guide prediction. Furthermore, we model the function prediction task as an iterative optimization problem for set prediction. A DETR-based decoder dynamically refines predictions through a series of layers, where each layer provides corrective feedback to progressively enhance the final output. Experimental results show that GDTGO achieves state-of-the-art performance on the PDB dataset. Yuzhi Sun, Yadong Wang 0001, Tianyi Zhao 0001 |
BIBM | 4 |
| 2025 | DNAMatch: An Ultra-Fast and Memory-Efficient Deep Learning Framework for Aligning Ultra-Long DNA FragmentsabstractEfficient and accurate alignment of DNA fragments is fundamental to genomics research. With the advent of advanced sequencing technologies, the length of sequencing reads and assembled contigs has increased significantly, posing substantial challenges for existing alignment algorithms. These methods often struggle with megabase-scale DNA fragments due to the computational burden of global searches and exhaustive chromosomal queries. To address this, we propose DNAMatch, a novel alignment framework that integrates the DNABERT2 pre-trained model for feature extraction with a deep residual network for chromosome identification. DNAMatch introduces a chromosome pre-localization strategy, which effectively narrows the search space and significantly reduces the memory footprint required by the downstream aligner, minimap2. This design enables fast and precise alignment of ultra-long DNA fragments. Benchmarking on simulated ultra-long reads from multiple model organisms-including human, Drosophila melanogaster, and Arabidopsis thaliana-demonstrates that DNAMatch achieves 9 8-9 9% accuracy in chromosome identification, accelerates the alignment process by 52.7%, and reduces memory usage by 73%. Importantly, when applied to downstream structural variation detection, DNAMatch maintains high accuracy, with only a marginal 0.05% decrease in F1-score compared to the standard minimap2 pipeline. These results highlight DNAMatch as a powerful and efficient tool for aligning ultra-long reads and analyzing complex genomes. Yuansong Zhu, Chuanmin Wu, Yadong Wang 0001, Tao Jiang 0021, Yadong Liu 0001 |
BIBM | 3 |
| 2025 | DMGAT: predicting ncRNA-drug resistance associations based on diffusion map and heterogeneous graph attention networkabstractNon-coding RNAs (ncRNAs) play crucial roles in drug resistance and sensitivity, making them important biomarkers and therapeutic targets. However, predicting ncRNA-drug associations is challenging due to issues such as dataset imbalance and sparsity, limiting the identification of robust biomarkers. Existing models often fall short in capturing local and global sequence information, limiting the reliability of predictions. This study introduces DMGAT (diffusion map and heterogeneous graph attention network), a novel deep learning model designed to predict ncRNA-drug associations. DMGAT integrates diffusion maps for sequence embedding, graph convolutional networks for feature extraction, and GAT for heterogeneous information fusion. To address dataset imbalance, the model incorporates sensitivity associations and employs a random forest classifier to select reliable negative samples. DMGAT embeds ncRNA sequences and drug SMILES using the word2vec technique, capturing local and global sequence information. The model constructs a heterogeneous network by combining sequence similarity and Gaussian Interaction Profile kernel similarity, providing a comprehensive representation of ncRNA-drug interactions. Evaluated through five-fold cross-validation on a curated dataset from NoncoRNA and ncDR, DMGAT outperforms seven state-of-the-art methods, achieving the highest area under the receiver operating characteristic curve (0.8964), area under the precision-recall curve (0.8984), recall (0.9576), and F1-score (0.8285). The raw data are released to Zenodo with identifier 13929676. The source code of DMGAT is available at https://github.com/liutingyu0616/DMGAT/tree/main. Tingyu Liu, Qiuhao Chen, Yuzhi Sun, Yadong Wang 0001, Tianyi Zhao 0001 |
Briefings Bioinform. | 5 |
| 2025 | CKG-TPI: integrating collaborative knowledge graph with sequence interactions for TCR-peptide binding specificityabstractAccurately identifying interactions between T-cell receptors (TCRs) and peptides is a fundamental challenge in immunology, with significant implications for vaccine design and immunotherapy. While computational methods offer efficient alternatives to labor-intensive experimental screening, achieving robust and accurate TCR-peptide binding prediction remains a challenging task. To address this, we propose collaborative knowledge graph (CKG-TPI), a novel prediction framework based on graph neural networks that integrates both interaction patterns between TCR and peptide sequences and their higher-order biological context through a constructed collaborative knowledge graph. Experimental results on multiple publicly available independent datasets demonstrate that CKG-TPI consistently outperforms state-of-the-art models. Specifically, it achieves a 9.89% improvement in area under the ROC curve compared to the strongest baseline model UnifyImmun, and a 23.93% increase in area under the precision-recall curve over the leading baseline method. Moreover, attention weight visualization and peptide-specific TCR screening validate the model's effectiveness, underscoring its potential as a powerful tool for immunological research and therapeutic discovery. Yue Liu 0034, Haoyan Wang, Guohua Wang 0001, Yadong Liu 0001, Tao Jiang 0021, Yadong Wang 0001 |
Briefings Bioinform. | 6 |
| 2025 | MetImputBERT: a pretrained BERT framework for missing value imputation in NMR metabolomics dataabstractMissing values in nuclear magnetic resonance metabolomics data compromise downstream clinical interpretation. Here, we present MetImputBERT, an imputation method based on a pretrained BERT framework. MetImputBERT uses the masks in the masked language model to simulate missing values and leverages predictions and reconstructions to these positions to simulate the imputation process. The learning of MetImputBERT is driven by minimizing the reconstruction error. MetImputBERT was pretrained on the largest metabolomics dataset to date, comprising data from over 230 000 individuals in the UK Biobank. When new datasets with missing values were encountered, MetImputBERT loaded the pretrained parameters and directly imputed the missing values by inferring their reconstructed estimates. MetImputBERT outperformed commonly used methods-K-nearest neighbors, multiple imputation by chained equations, and singular value decomposition-in imputation performance on two independent test sets. We provide an open-source Python tool that allows users to quickly impute missing values in their own NMR metabolomics data without any additional training. Shizheng Qiu, Yang Hu 0008, Guiyou Liu, Yadong Wang 0001 |
Briefings Bioinform. | 4 |
| 2025 | Proformer: a multimodal proteomics transformer model for multidisease early risk assessmentabstractEarly identification of individuals at high risk for chronic diseases is crucial for prevention and intervention, yet current risk assessment tools are disease-specific, require extensive clinical data collection, and cannot provide multidisease risk profiles from a single measurement. Several protein large language models have been developed for tasks such as protein structure prediction, function prediction, and sequence design. However, none of these models can be directly applied in clinical settings to predict an individual's future disease risk. Here, we present a multimodal proteomics Transformer (Proformer) model that integrates protein expression, sequence, and function information for multidisease risk assessment. We trained Proformer using real proteomics data from 47 124 individuals from the UK Biobank to evaluate its performance in discriminating the risk of 20 common chronic diseases. Proformer achieved state-of-the-art (SOTA) performance in all 20 diseases compared with five common machine learning and deep learning models. Compared to three common clinical predictors, Proformer's 10-year discriminative performance outperforms Age + Sex model for 19 diseases, outperforms the ASCVD risk score for 16 diseases, and outperforms the panel composed of 35 clinical variables for 11 diseases. These results were replicated in the Scotland and Wales cohort from UK Biobank. In conclusion, Proformer enabled users to directly obtain a 10-year risk report for common chronic diseases by inputting their individual proteomics data. Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001 |
Briefings Bioinform. | 4 |
| 2025 | MSNGO: multi-species protein function annotation based on 3D protein structure and network propagationabstractMOTIVATION: In recent years, protein function prediction has broken through the bottleneck of sequence features, significantly improving prediction accuracy using high-precision protein structures predicted by AlphaFold2. While single-species protein function prediction methods have achieved remarkable success, multi-species approaches still face challenges such as difficulties in multi-source data integration and insufficient knowledge transfer between distantly-related species. How to integrate large-scale data and provide effective cross-species label propagation for species with sparse protein annotations remains a critical and unresolved challenge. To address this problem, we propose the MSNGO (Multi-species protein Structures and Network to predict GO terms) model, which integrates structural features and network propagation methods. Our validation shows that using structural features can significantly improve the accuracy of multi-species protein function prediction. RESULTS: We employ graph representation learning techniques to extract amino acid representations from protein structure contact maps and train a structural model using a graph convolution pooling module to derive protein-level structural features. After incorporating the sequence features from ESM-2, we apply a network propagation algorithm to aggregate information and update node representations within a heterogeneous network. The results demonstrate that MSNGO outperforms previous multi-species protein function prediction methods that rely on sequence features and protein-protein networks. AVAILABILITY AND IMPLEMENTATION: https://github.com/blingbell/MSNGO. Boyue Cui, Shiqu Chen, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 5 |
| 2024 | Comprehensive Benchmarking of Genotype Imputation Tools Using a Large-Scale Chinese Reference PanelabstractGenome-wide association studies (GWAS) remain as one of the most essential and potent strategies for identifying genetic markers associated with common human diseases. Gene arrays, which are widely used in GWAS, target predesigned known genomic mutations to enhance clinical diagnostics. The approach of genotype imputation which leverages reference panels has shown great promise in extrapolating additional meaningful variants from existing data. Therefore, accurately broadening the spectrum of mutations beyond those included in gene arrays can unlock significant potential for cost-effective and advanced GWAS analysis. In this study, we focus on the four most popular imputation tools currently to assess their performances based on a large-scale Chinese reference panel. Benchmarking results demonstrate that Minimac4 is the best imputation method, which achieved higher accuracy with more well-imputed loci and maintained lower computational consumption. Moreover, we further revealed the minimum sample size required in the reference panel and the mutations involved in arrays that ensure the acquisition of acceptable imputed call sets. We believe this study can provide practical significance guidelines for the research community to select the suitable imputation tool via an existing reference panel, or construct their own reference panel and genotype array in the most effective way. Yadong Liu 0001, Zhongbo Yang, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 4 |
| 2024 | scEAGC: an efficient anchor graph clustering for single-cell transcriptomics and proteomics dataabstractSingle-cell multi-omics sequencing allows researchers to simultaneously sequence multiple types of molecular information from the same individual cell, like transcriptomics and proteomics. However, the research of identifying cell types from single-cell multi-omics data is still challenging. In this article, we proposed scEAGC, an efficient anchor graph clustering for single-cell transcriptomics and proteomics data. It first constructs anchor cell graphs for every omics and then integrates separated omics-specific anchor cell graphs on a weighted multi-view clustering model with F-norm and the Orthogonal constraint, finally through the divided iterative optimization method to obtain cluster partition without any extra post-processing. Since scEAGC combines the high-efficiency property of anchor graph clustering, its efficiency is substantially higher than widely used algorithms. Extensive experiments demonstrate that, compared to other state-of-the-art clustering algorithms, scEAGC can boost clustering accuracy and robustness and detect the new cell subtypes in CITE-seq and scRNA-seq data. Qiaoming Liu, Yadong Wang 0001, Guohua Wang 0001 |
BIBM | 2 |
| 2024 | Comprehensive evaluation of haplotype phasing tools with different strategies across diverse sequence technologiesabstractHaplotype phasing is a computational technique used to determine the allelic arrangement in a chromosome from genotype data. This process is crucial for understanding genetic diversity and inheritance patterns, which are essential for fields like genetic research, personalized medicine, and population genetic studies. In this study, we conducted a comprehensive evaluation of five state-of-the-art haplotype phasing tools under different strategies including statistical-based Beagle5, Eagle2 and Shapeit5, and read-based WhatsHap, and HapCut2, across various sequencing platforms including Illumina, PacBio, and ONT. Our findings indicate that statistical-based tools generally provide better phasing continuity and are less dependent on the sequencing technology used, whereas read-based tools offer superior performance with long-read datasets but struggle with short reads due to insufficient sequence length. In addition, statistical-based tools were found to be less resource-intensive compared to their read-based counterparts. The insights gained from this benchmarking provide a valuable guide for researchers in selecting the most appropriate phasing tools, potentially enhancing the accuracy and efficiency of genetic analysis and supporting advancements in genetic research methodologies. Zhongbo Yang, Zhenhao Lu, Tao Jiang 0021, Yadong Wang 0001, Yadong Liu 0001 |
BIBM | 5 |
| 2024 | miniSNV: accurate and fast single nucleotide variant calling from nanopore sequencing dataabstractNanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV. Miao Cui 0005, Yadong Liu 0001, Hongzhe Guo, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023 |
Briefings Bioinform. | 6 |
| 2024 | THItoGene: a deep learning method for predicting spatial transcriptomics from histological imagesabstractSpatial transcriptomics unveils the complex dynamics of cell regulation and transcriptomes, but it is typically cost-prohibitive. Predicting spatial gene expression from histological images via artificial intelligence offers a more affordable option, yet existing methods fall short in extracting deep-level information from pathological images. In this paper, we present THItoGene, a hybrid neural network that utilizes dynamic convolutional and capsule networks to adaptively sense potential molecular signals in histological images for exploring the relationship between high-resolution pathology image phenotypes and regulation of gene expression. A comprehensive benchmark evaluation using datasets from human breast cancer and cutaneous squamous cell carcinoma has demonstrated the superior performance of THItoGene in spatial gene expression prediction. Moreover, THItoGene has demonstrated its capacity to decipher both the spatial context and enrichment signals within specific tissue regions. THItoGene can be freely accessed at https://github.com/yrjia1015/THItoGene. Yuran Jia, Tianyi Zhao 0001, Yadong Wang 0001 |
Briefings Bioinform. | 5 |
| 2024 | Kled: an ultra-fast and sensitive structural variant detection tool for long-read sequencing dataabstractStructural Variants (SVs) are a crucial type of genetic variant that can significantly impact phenotypes. Therefore, the identification of SVs is an essential part of modern genomic analysis. In this article, we present kled, an ultra-fast and sensitive SV caller for long-read sequencing data given the specially designed approach with a novel signature-merging algorithm, custom refinement strategies and a high-performance program structure. The evaluation results demonstrate that kled can achieve optimal SV calling compared to several state-of-the-art methods on simulated and real long-read data for different platforms and sequencing depths. Furthermore, kled excels at rapid SV calling and can efficiently utilize multiple Central Processing Unit (CPU) cores while maintaining low memory usage. The source code for kled can be obtained from https://github.com/CoREse/kled. Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Bo Liu 0023, Yadong Wang 0001 |
Briefings Bioinform. | 7 |
| 2024 | MEHunter: transformer-based mobile element variant detection from long readsabstractSUMMARY: Mobile genetic elements (MEs) are heritable mutagens that significantly contribute to genetic diseases. The advent of long-read sequencing technologies, capable of resolving large DNA fragments, offers promising prospects for the comprehensive detection of ME variants (MEVs). However, achieving high precision while maintaining recall performance remains challenging mainly brought by the variable length and similar content of MEV signatures, which are often obscured by the noise in long reads. Here, we propose MEHunter, a high-performance MEV detection approach utilizing a fine-tuned transformer model adept at identifying potential MEVs with fragmented features. Benchmark experiments on both simulated and real datasets demonstrate that MEHunter consistently achieves higher accuracy and sensitivity than the state-of-the-art tools. Furthermore, it is capable of detecting novel potentially individual-specific MEVs that have been overlooked in published population projects. AVAILABILITY AND IMPLEMENTATION: MEHunter is available from https://github.com/120L021101/MEHunter. Tao Jiang 0021, Zuji Zhou, Shuqi Cao, Yadong Wang 0001, Yadong Liu 0001 |
Bioinform. | 5 |
| 2024 | MCDHGN: heterogeneous network-based cancer driver gene prediction and interpretability analysisabstractMOTIVATION: Accurately predicting the driver genes of cancer is of great significance for carcinogenesis progress research and cancer treatment. In recent years, more and more deep-learning-based methods have been used for predicting cancer driver genes. However, deep-learning algorithms often have black box properties and cannot interpret the output results. Here, we propose a novel cancer driver gene mining method based on heterogeneous network meta-paths (MCDHGN), which uses meta-path aggregation to enhance the interpretability of predictions. RESULTS: MCDHGN constructs a heterogeneous network by using several types of multi-omics data that are biologically linked to genes. And the differential probabilities of SNV, DNA methylation, and gene expression data between cancerous tissues and normal tissues are extracted as initial features of genes. Nine meta-paths are manually selected, and the representation vectors obtained by aggregating information within and across meta-path nodes are used as new features for subsequent classification and prediction tasks. By comparing with eight homogeneous and heterogeneous network models on two pan-cancer datasets, MCDHGN has better performance on AUC and AUPR values. Additionally, MCDHGN provides interpretability of predicted cancer driver genes through the varying weights of biologically meaningful meta-paths. AVAILABILITY AND IMPLEMENTATION: https://github.com/1160300611/MCDHGN. Lexiang Wang, Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 4 |
| 2024 | HBFormer: a single-stream framework based on hybrid attention mechanism for identification of human-virus protein-protein interactionsabstractMOTIVATION: Exploring human-virus protein-protein interactions (PPIs) is crucial for unraveling the underlying pathogenic mechanisms of viruses. Limitations in the coverage and scalability of high-throughput approaches have impeded the identification of certain key interactions. Current popular computational methods adopt a two-stream pipeline to identify PPIs, which can only achieve relation modeling of protein pairs at the classification phase. However, the fitting capacity of the classifier is insufficient to comprehensively mine the complex interaction patterns between protein pairs. RESULTS: In this study, we propose a pioneering single-stream framework HBFormer that combines hybrid attention mechanism and multimodal feature fusion strategy for identifying human-virus PPIs. The Transformer architecture based on hybrid attention can bridge the bidirectional information flows between human protein and viral protein, thus unifying joint feature learning and relation modeling of protein pairs. The experimental results demonstrate that HBFormer not only achieves superior performance on multiple human-virus PPI datasets but also outperforms 5 other state-of-the-art human-virus PPI identification methods. Moreover, ablation studies and scalability experiments further validate the effectiveness of our single-stream framework. AVAILABILITY AND IMPLEMENTATION: Codes and datasets are available at https://github.com/RmQ5v/HBFormer. Yadong Wang 0001, Tianyi Zhao 0001 |
Bioinform. | 3 |
| 2024 | DVA: predicting the functional impact of single nucleotide missense variantsabstractBACKGROUND: In the past decade, single nucleotide variants (SNVs) have been identified as having a significant relationship with the development and treatment of diseases. Among them, prioritizing missense variants for further functional impact investigation is an essential challenge in the study of common disease and cancer. Although several computational methods have been developed to predict the functional impacts of variants, the predictive ability of these methods is still insufficient in the Mendelian and cancer missense variants. RESULTS: We present a novel prediction method called the disease-related variant annotation (DVA) method that predicts the effect of missense variants based on a comprehensive feature set of variants, notably, the allele frequency and protein-protein interaction network feature based on graph embedding. Benchmarked against datasets of single nucleotide missense variants, the DVA method outperforms the state-of-the-art methods by up to 0.473 in the area under receiver operating characteristic curve. The results demonstrate that the proposed method can accurately predict the functional impact of single nucleotide missense variants and substantially outperforms existing methods. CONCLUSIONS: DVA is an effective framework for identifying the functional impact of disease missense variants based on a comprehensive feature set. Based on different datasets, DVA shows its generalization ability and robustness, and it also provides innovative ideas for the study of the functional mechanism and impact of SNVs. Dong Wang 0066, Jie Li 0055, Edwin Wang, Yadong Wang 0001 |
BMC Bioinform. | 4 |
| 2024 | circ2DGNN: circRNA-Disease Association Prediction via Transformer-Based Graph Neural NetworkabstractInvestigating the associations between circRNA and diseases is vital for comprehending the underlying mechanisms of diseases and formulating effective therapies. Computational prediction methods often rely solely on known circRNA-disease data, indirectly incorporating other biomolecules' effects by computing circRNA and disease similarities based on these molecules. However, this approach is limited, as other biomolecules also play significant roles in circRNA-disease interactions. To address this, we construct a comprehensive heterogeneous network incorporating data on human circRNAs, diseases, and other biomolecule interactions to develop a novel computational model, circ2DGNN, which is built upon a heterogeneous graph neural network. circ2DGNN directly takes heterogeneous networks as inputs and obtains the embedded representation of each node for downstream link prediction through graph representation learning. circ2DGNN employs a Transformer-like architecture, which can compute heterogeneous attention score for each edge, and perform message propagation and aggregation, using a residual connection to enhance the representation vector. It uniquely applies the same parameter matrix only to identical meta-relationships, reflecting diverse parameter spaces for different relationship types. After fine-tuning hyperparameters via five-fold cross-validation, evaluation conducted on a test dataset shows circ2DGNN outperforms existing state-of-the-art(SOTA) methods. Keliang Cen, Zheming Xing, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Automatically Detecting Anchor Cells and Clustering for scRNA-Seq Data Using scTSNNabstractAdvancing in single-cell RNA sequencing techniques enhances the resolution of cell heterogeneity study. Density-based unsupervised clustering has the potential to detect the representative anchor points and the number of clusters automatically. Meanwhile, discovering the true cell type of scRNA-seq data in the unsupervised scenario is still challenging. To this end, we proposed a tensor shared nearest neighbor anchor clustering for scRNA-seq data, named scTSNN, which first makes use of the tensor affinity learning module to mine the local-global balanced topological structures among cells, next designs density-based shared nearest neighbor measurement method to automatically detect anchor cells, finally partitions the non-anchor cells to obtain the clustering results. Validated on synthetic datasets and scRNA-seq datasets, scTSNN not only exactly detects the complicated structures but also has better performance in accuracy and robustness compared with the state-of-the-art methods. Moreover, case studies on mammalian cells and cervical cancer tumor cells demonstrate the selected anchor cells of scTSNN benefit the cell pseudotime inference and rare cell identification, which show good application and research value of scTSNN. Qiaoming Liu, Dong Wang 0066, Guohua Wang 0001, Yadong Wang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Comprehensive evaluation of RNA-seq alignment methods based on long-read sequencing dataabstractLong-read RNA sequencing (RNA-seq) has revolutionized our ability to comprehensively study transcriptomes, enabling the detection of full-length transcripts. The accurate alignment of long-read RNA-seq is a critical step in downstream analysis. However, with the development of alignment tools and each tool declaring the ability to handle long-read RNA-seq alignment on their own, researchers face the challenge of distinguishing and selecting the most suitable tool for their tasks. But to the best of our knowledge, there is still a lack of comprehensive evaluations on the aligners designed for long-read sequencing data. Here, we conducted a benchmark on five state-of-the-art tools on simulated long-read RNA-seq data from three levels to comprehensively assess the performance of each aligner, and provide valuable insights into the strengths and limitations of each aligner, aiding researchers in selecting the most appropriate tool for their long-read RNA-seq studies. We expected our study to advance the field of transcriptomics and promote the accurate analysis of long-read RNA-seq data in diverse biological contexts. Yadong Liu 0001, Hongzhe Guo, Zhenhao Lu, Yadong Wang 0001, Zhongyu Liu, Tao Jiang 0021 |
BIBM | 4 |
| 2023 | Identification of risk genes and biological pathways influencing myopia via transcriptome association study and biomedical ontology methodsabstractRefractive error remains one of the most common eye diseases in the world, which causes great burden on public health and economy every year. Although genome-wide association studies (GWAS) have identified susceptibility loci for refractive error, the credible causal genes associated with gene expression in brain are still unclear. In addition, the influence of RNA splicing events on refractive error in brain remains unknown. Here, we identified candidate refractive error gene targets and biological pathways. In stage 1, we carried out gene-based association analysis to identify refractive error risk genes. In stage 2, by using transcriptome-wide association studies (TWAS) and colocalization analysis, we identified cis-regulated genes in brain tissues that were positively related to refractive error from the whole gene region and the exon region after RNA splicing, respectively. In stage 3, we reveal biological pathways influencing refractive error using biological ontology and knowledge base methods. We identified 367 risk genes passed gene-based association test. A total of 43 genes in brain could be considered as causal genes using TWAS and colocalization analysis, and 27 were novel. After RNA splicing, RCBTB1 significantly affected refractive error. Risk genes for refractive error significantly enriched in biological processes related to synaptic transmission and acetylcholine. In conclusion, our analysis provides a better understanding of potential therapeutic targets and biological pathways for refractive error. Shizheng Qiu, Yadong Wang 0001, Yang Hu 0008 |
BIBM | 2 |
| 2023 | HNetGO: protein function prediction via heterogeneous network transformerabstractProtein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches. Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004 |
Briefings Bioinform. | 8 |
| 2023 | Struct2GO: protein function prediction based on graph pooling algorithm and AlphaFold2 structure informationabstractMOTIVATION: In recent years, there has been a breakthrough in protein structure prediction, and the AlphaFold2 model of the DeepMind team has improved the accuracy of protein structure prediction to the atomic level. Currently, deep learning-based protein function prediction models usually extract features from protein sequences and combine them with protein-protein interaction networks to achieve good results. However, for newly sequenced proteins that are not in the protein-protein interaction network, such models cannot make effective predictions. To address this, this article proposes the Struct2GO model, which combines protein structure and sequence data to enhance the precision of protein function prediction and the generality of the model. RESULTS: We obtain amino acid residue embeddings in protein structure through graph representation learning, utilize the graph pooling algorithm based on a self-attention mechanism to obtain the whole graph structure features, and fuse them with sequence features obtained from the protein language model. The results demonstrate that compared with the traditional protein sequence-based function prediction model, the Struct2GO model achieves better results. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available at https://github.com/lyjps/Struct2GO. Peishun Jiao, Xuan Wang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 5 |
| 2023 | ScCCL: Single-Cell Data Clustering Based on Self-Supervised Contrastive LearningabstractThe growing maturity of single-cell RNA-sequencing (scRNA-seq) technology allows us to explore the heterogeneity of tissues, organisms, and complex diseases at cellular level. In single-cell data analysis, clustering calculation is very important. However, the high dimensionality of scRNA-seq data, the ever-increasing number of cells, and the unavoidable technical noise bring great challenges to clustering calculations. Motivated by the good performance of contrastive learning in multiple domains, we propose ScCCL, a novel self-supervised contrastive learning method for clustering of scRNA-seq data. ScCCL first randomly masks the gene expression of each cell twice and adds a small amount of Gaussian noise, and then uses the momentum encoder structure to extract features from the enhanced data. Contrastive learning is then applied in the instance-level contrastive learning module and the cluster-level contrastive learning module, respectively. After training, a representation model that can efficiently extract high-order embeddings of single cells is obtained. We selected two evaluation metrics, ARI and NMI, to conduct experiments on multiple public datasets. The results show that ScCCL improves the clustering effect compared with the benchmark algorithms. Notably, since ScCCL does not depend on a specific type of data, it can also be helpful in clustering analysis of single-cell multi-omics data. Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | PSPGO: Cross-Species Heterogeneous Network Propagation for Protein Function PredictionabstractHow to use computational methods to effectively predict the function of proteins remains a challenge. Most prediction methods based on single species or single data source have some limitations: the former need to train different models for different species, the latter only to infer protein function from a single perspective, such as the method only using Protein-Protein Interaction (PPI) network just considers the protein environment but ignore the intrinsic characteristics of protein sequences. We found that in some network-based multi-species methods the networks of each species are isolated, which means there is no communication between networks of different species. To solve these problems, we propose a cross-species heterogeneous network propagation method based on graph attention mechanism, PSPGO, which can propagate feature and label information on sequence similarity (SS) network and PPI network for predicting gene ontology terms. Our model is evaluated on a large multi-species dataset split based on time and is compared with several state-of-the-art methods. The results show that our method has good performance. We also explore the predictive performance of PSPGO for a single species. The results illustrate that PSPGO also performs well in prediction for single species. Kaitao Wu, Lexiang Wang, Bo Liu 0023, Yang Liu 0039, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Prot2GO: Predicting GO Annotations From Protein Sequences and InteractionsabstractProtein is the main material basis of living organisms and plays crucial role in life activities. Understanding the function of protein is of great significance for new drug discovery, disease treatment and vaccine development. In recent years, with the widespread application of deep learning in bioinformatics, researchers have proposed many deep learning models to predict protein functions. However, the existing deep learning methods usually only consider protein sequences, and thus cannot effectively integrate multi-source data to annotate protein functions. In this article, we propose the Prot2GO model, which can integrate protein sequence and PPI network data to predict protein functions. We utilize an improved biased random walk algorithm to extract the features of PPI network. For sequence data, we use a convolutional neural network to obtain the local features of the sequence and a recurrent neural network to capture the long-range associations between amino acid residues in protein sequence. Moreover, Prot2GO adopts the attention mechanism to identify protein motifs and structural domains. Experiments show that Prot2GO model achieves the state-of-the-art performance on multiple metrics. Xiaoshuai Zhang, Hucheng Liu, Xiaofeng Zhang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | Predicting Drug-Disease Associations Through Similarity Network Fusion and Multi-View Feature Projection RepresentationabstractPredicting drug-disease associations (DDAs) through computational methods has become a prevalent trend in drug development because of their high efficiency and low cost. Existing methods usually focus on constructing heterogeneous networks by collecting multiple data resources to improve prediction ability. However, potential association possibilities of numerous unconfirmed drug-related or disease-related pairs are not sufficiently considered. In this article, we propose a novel computational model to predict new DDAs. First, a heterogeneous network is constructed, including four types of nodes (drugs, targets, cell lines, diseases) and three types of edges (associations, association scores, similarities). Second, an updating and merging-based similarity network fusion method, termed UM-SF, is presented to fuse various similarity networks with diverse weights. Finally, an intermediate layer-mediated multi-view feature projection representation method, termed IM-FP, is proposed to calculate the predicted DDA scores. This method uses multiple association scores to construct multi-view drug features, then projects them into disease space through the intermediate layer, where an intermediate layer similarity constraint is designed to learn the projection matrices. Results of comparative experiments reveal the effectiveness of our innovations. Comparisons with other state-of-the-art models by the 10-fold cross-validation experiment indicate our model's advantage on AUROC and AUPR metrics. Moreover, our proposed model successfully predicted 107 novel high-ranked DDAs. Jie Li 0055, Dong Wang 0066, Dechen Xu, Jiahuan Jin, Yadong Wang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Real-time Image Enhancement with Attention AggregationabstractImage enhancement has stimulated significant research works over the past years for its great application potential in video conferencing scenarios. Nevertheless, most existing image enhancement approaches are still struggling to find a good tradeoff that reduces the computational cost as much as possible while maintaining plausible result quality. Recently, curve-based mapping methods are proposed and have shown great potential for real-time and high-quality image enhancement of arbitrary resolutions. In this article, we take advantage of the curve-based mapping representation and focus on further improving the enhancement quality and robustness, while minimizing additional computational costs. Specifically, we (1) carefully re-formulate the curve function to improve learning stability, and (2) aggregate different semantic attention into the curve regression process, which can overcome the major problems of curve-based methods that generate moderate results with low contrast. The semantic attention is jointly learned with the supervision from class activation mapping of pre-trained feature extractors, thus reducing the manual annotation cost of semantic labels. Experiments have shown that our proposed method significantly improves curve-based methods both qualitatively and quantitatively, achieving visually plausible results compared with other deep neural network-based enhancement methods, and maintains a very low computational cost, i.e., taking 18.7 ms for a 360p image on a single P40 GPU. Extensive experiments demonstrate that our method is also capable of video enhancement tasks. Jie Li 0055, Tiejun Zhao, Yadong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Comparison of the Nanopore and PacBio sequencing technologies for DNA 5-methylcytosine detectionabstractDNA methylation provides a pivotal layer of epigenetic regulation in eukaryotes that has significant involvement for numerous biological processes in health and disease. Recent long-read sequencing technology including Oxford Nanopore sequencing and PacBio HiFi sequencing greatly expands the capacity of long-range, single-molecule, and direct DNA modification detection from reads without extra laboratory techniques. A growing number of analytical pipelines including base-calling and 5mC methylation detection have been developed, but there is still a lack of comprehensive evaluations of the two sequencing technologies. Here, we assess the performance of different methylation-calling pipelines based on Nanopore and HiFi sequencing datasets to provide a systematic evaluation to guide researchers on how to select the long-read sequencing technologies in performing human epigenome-wide studies. Yadong Liu 0001, Zhongyu Liu, Tao Jiang 0021, Tianyi Zang, Yadong Wang 0001 |
BIBM | 5 |
| 2022 | NSAP: A Neighborhood Subgraph Aggregation Method for Drug-Disease Association Prediction
Qiqi Jiao, Yu Jiang 0002, Yang Zhang 0057, Yadong Wang 0001, Junyi Li 0004 |
ICIC (2) | 4 |
| 2022 | Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractQuantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 10 |
| 2022 | Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputationabstractIn the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’. Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng |
Briefings Bioinform. | 10 |
| 2022 | Evaluation of classification in single cell atac-seq data with machine learning methodsabstractBACKGROUND: The technologies advances of single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) allowed to generate thousands of single cells in a relatively easy and economic manner and it is rapidly advancing the understanding of the cellular composition of complex organisms and tissues. The data structure and feature in scRNA-seq is similar to that in scATAC-seq, therefore, it's encouraged to identify and classify the cell types in scATAC-seq through traditional supervised machine learning methods, which are proved reliable in scRNA-seq datasets. RESULTS: In this study, we evaluated the classification performance of 6 well-known machine learning methods on scATAC-seq. A total of 4 public scATAC-seq datasets vary in tissues, sizes and technologies were applied to the evaluation of the performance of the methods. We assessed these methods using a 5-folds cross validation experiment, called intra-dataset experiment, based on recall, precision and the percentage of correctly predicted cells. The results show that these methods performed well in some specific types of the cell in a specific scATAC-seq dataset, while the overall performance is not as well as that in scRNA-seq analysis. In addition, we evaluated the classification performance of these methods by training and predicting in different datasets generated from same sample, called inter-datasets experiments, which may help us to assess the performance of these methods in more realistic scenarios. CONCLUSIONS: Both in intra-dataset and in inter-dataset experiment, SVM and NMC are overall outperformed others across all 4 datasets. Thus, we recommend researchers to use SVM and NMC as the underlying classifier when developing an automatic cell-type classification method for scATAC-seq. Hongzhe Guo, Zhongbo Yang, Tao Jiang 0021, Yadong Wang 0001 |
BMC Bioinform. | 5 |
| 2022 | M2PP: a novel computational model for predicting drug-targeted pathogenic proteinsabstractBACKGROUND: Detecting pathogenic proteins is the origin way to understand the mechanism and resist the invasion of diseases, making pathogenic protein prediction develop into an urgent problem to be solved. Prediction for genome-wide proteins may be not necessarily conducive to rapidly cure diseases as developing new drugs specifically for the predicted pathogenic protein always need major expenditures on time and cost. In order to facilitate disease treatment, computational method to predict pathogenic proteins which are targeted by existing drugs should be exploited. RESULTS: In this study, we proposed a novel computational model to predict drug-targeted pathogenic proteins, named as M2PP. Three types of features were presented on our constructed heterogeneous network (including target proteins, diseases and drugs), which were based on the neighborhood similarity information, drug-inferred information and path information. Then, a random forest regression model was trained to score unconfirmed target-disease pairs. Five-fold cross-validation experiment was implemented to evaluate model's prediction performance, where M2PP achieved advantageous results compared with other state-of-the-art methods. In addition, M2PP accurately predicted high ranked pathogenic proteins for common diseases with public biomedical literature as supporting evidence, indicating its excellent ability. CONCLUSIONS: M2PP is an effective and accurate model to predict drug-targeted pathogenic proteins, which could provide convenience for the future biological researches. Jie Li 0055, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2022 | A Neighborhood-Based Global Network Model to Predict Drug-Target InteractionsabstractThe detection of drug-target interactions (DTIs) plays an important role in drug discovery and development, making DTI prediction urgent to be solved. Existing computational methods usually utilize drug similarity, target similarity and DTI information to make prediction, providing the convenience of fast time and low cost. However, they usually learn features for drugs and targets separately, lacking of a global consideration. In this study, we proposed a novel neighborhood-based global network model, named as NGN, to accurately predict DTIs from the global perspective. We designed a distance constraint for features of all entities (drugs and targets) in the latent space to ensure the close distance between adjacent entities, and defined a global probability matrix to compute the predicted DTI scores on our constructed neighborhood-based global network. Results showed that NGN obtained advantageous performance compared with other state-of-the-art methods, especially surpassing them by 4.2-9.1 percent on AUPR values in the biggest dataset. Furthermore, several novel high-ranked DTIs were successfully predicted with confirmations by public sources, demonstrating the effectiveness of our method. Jie Li 0055, Yadong Wang 0001, Liran Juan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Deepgmd: A Graph-Neural-Network-Based Method to Detect Gene Regulator ModuleabstractRegulatory module mining methods divide genes into multiple gene subgroups and explore potential biological mechanisms from omics data. By transforming gene expression profile data into gene co-expression network, we transform the task of gene module detection into the problem of finding community structure in the graph, and introduce the latest network representation learning method-graph neural network to optimize this problem. In order to systematically evaluate whether the algorithm allows overlap to affect such problems, we make two variants of the output of the algorithm, Deepgmd_cluster and Deepgmd. The difference between them is whether overlap is allowed. By comparing the known modules and the modules generated by the algorithm, we can evaluate the quality of the algorithm. We use this method to compare our algorithm with some current mainstream methods. The results show that our method has greater advantages. In the end, we analyze some typical modules from the modules found by the algorithm for visualization, and use the GO database and KEGG database to perform enrichment analysis and pathway analysis on these modules. Yulin Wu 0001, Jiangsheng Pi, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | Differentially Expressed Mutant Genes Reveal Potential Prognostic Markers For Lung AdenocarcinomaabstractLung adenocarcinoma is a serious lung cancer, belonging to the category of non-small cell lung cancer, accounting for 30 to 35 percent of the total, and lung adenocarcinoma is more common in women and non-smoking patients. In this study, LUAD samples of mutation genes and RNA-seq expression dataset were retrieved from TCGA database. WGCNA, Cox regression analysis were used to classify melanoma prognosis. It revealed that eleven mutant and differentially expressed prognosis biomarkers were significantly associated with the prognosis of patients. Therefore, detecting these gene mutations and exploring their corresponding expression could be valuable in predicting the prognosis of patients. The results of the high-throughput data mining provide important fundamental bioinformatics information and a relevant theoretical basis for further exploring the molecular pathogenesis of LUAD and assessing the prognosis of patients. Yue Liu 0034, Shizheng Qiu, Yang Hu 0008, Yadong Wang 0001 |
BIBM | 4 |
| 2021 | Novel Multikernel Trick for Predicting Pan-CancerDistant Metastatic Sites Using a Feature Extraction StrategyabstractDistant metastasis is the leading cause of cancer death. Identifying the tendency of a given cancer to metastasize could be conducive to cancer diagnosis and therapeutic schedules. In cancer studies, mRNA gene expression data have been widely used to predict cancer metastasis due to the ease with which they can be obtained. Moreover, mRNA gene expression data represent cancer progression directly and in detail. In these studies, feature extraction followed by a prediction model has been a commonly used solution to predict pan-cancer prognosis and tumor stage. Limitations of these studies include a lack of comprehensive feature extraction, relatively low prediction accuracy of cancer outcomes and a lack of precise pan-cancer metastasis site prediction.To address the questions mentioned above, we designed an innovative pipeline to determine the heterogeneity of pan-cancer distant metastatic sites using mRNA gene expression data. We used a directed relational graph convolutional network (DRGCN) for feature extraction and a multikernel support vector machine (SVM) for pan-cancer distant metastasis site prediction. DR-GCN successfully excavated hidden features from relational networks and effectively extracted features from gene-cancer relations, cancer-disease relations and gene-gene relational networks. DR-GCN was demonstrably able to deal with complex prior knowledge-based feature extraction tasks. A dynamic weight multikernel SVM was then applied to predict pan-cancer distant metastasis sites. By this method, the AUROC (0.7542) of the multikernel SVM outperformed that of the single kernel SVM (polykernel: 0.7346, RBF kernel: 0.72, linear kernel: 0. 725S). We last applied our pipeline to an extremely unbalanced small sample dataset and obtained a higher AUPRC (0.2606) than other semisupervised learning methods (Laplacian SVM: 0. 1S, TSVM: 0.21, SSL-EM: 0. 1S, RRLSL: 0.22) while predicting TCGA glioblastoma (GBM) patient prognosis. Xinran Cui, Tianyi Zhao 0001, Yadong Wang 0001 |
BIBM | 5 |
| 2021 | PocaCNV: A Tool to Detect Copy Number Variants from Population-Scale Genome Sequencing DataabstractAmong the procedures of modern genome analysis, the variation calling is a crucial part. Copy Number Variation(CNV) is an important variation type, many algorithms are developed to detect them. As the scale of samples in genome projects growing, it’s wise to use data from multiple samples jointly to call variants. We present PocaCNV, a population caller of CNVs that can call CNVs from multiple samples of the same population. PocaCNV can produce many unique results with good precision and sensitivity, and can be a good supplement to other modern SV calling procedures. PocaCNV can be openly accessed from https://github.com/CoREse/PocaCNV. Yongzhuang Liu, Yadong Wang 0001 |
BIBM | 4 |
| 2021 | ScSSC: Semi-supervised Single Cell Clustering Based on 2D Embedding
Naile Shi, Yulin Wu 0001, Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
ICIC (3) | 5 |
| 2021 | A deep learning approach for filtering structural variants in short read sequencing dataabstractShort read whole genome sequencing has become widely used to detect structural variants in human genetic studies and clinical practices. However, accurate detection of structural variants is a challenging task. Especially existing structural variant detection approaches produce a large proportion of incorrect calls, so effective structural variant filtering approaches are urgently needed. In this study, we propose a novel deep learning-based approach, DeepSVFilter, for filtering structural variants in short read whole genome sequencing data. DeepSVFilter encodes structural variant signals in the read alignments as images and adopts the transfer learning with pre-trained convolutional neural networks as the classification models, which are trained on the well-characterized samples with known high confidence structural variants. We use two well-characterized samples to demonstrate DeepSVFilter's performance and its filtering effect coupled with commonly used structural variant detection approaches. The software DeepSVFilter is implemented using Python and freely available from the website at https://github.com/yongzhuang/DeepSVFilter. Yongzhuang Liu, Yalin Huang, Guohua Wang 0001, Yadong Wang 0001 |
Briefings Bioinform. | 4 |
| 2021 | An integrated approach for copy number variation discovery in parent-offspring triosabstractWhole-genome sequencing (WGS) of parent-offspring trios has become widely used to identify causal copy number variations (CNVs) in rare and complex diseases. Existing CNV detection approaches usually do not make effective use of Mendelian inheritance in parent-offspring trios and yield low accuracy. In this study, we propose a novel integrated approach, TrioCNV2, for jointly detecting CNVs from WGS data of the parent-offspring trio. TrioCNV2 first makes use of the read depth and discordant read pairs to infer approximate locations of CNVs and then employs the split read and local de novo assembly approaches to refine the breakpoints. We use the real WGS data of two parent-offspring trios to demonstrate TrioCNV2's performance and compare it with other CNV detection approaches. The software TrioCNV2 is implemented using a combination of Java and R and is freely available from the website at https://github.com/yongzhuang/TrioCNV2. Yongzhuang Liu, Yadong Wang 0001 |
Briefings Bioinform. | 3 |
| 2021 | abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractSUMMARY: Partial order alignment, which aligns a sequence to a directed acyclic graph, is now frequently used as a key component in long-read error correction and assembly. We present abPOA (adaptive banded Partial Order Alignment), a Single Instruction Multiple Data (SIMD)-based C library for fast partial order alignment using adaptive banded dynamic programming. It can work as a stand-alone multiple sequence alignment and consensus calling tool or be easily integrated into any long-read error correction and assembly workflow. Compared to a state-of-the-art tool (SPOA), abPOA is up to 10 times faster with a comparable alignment accuracy. AVAILABILITY AND IMPLEMENTATION: abPOA is implemented in C. A stand-alone tool and a C/Python software interface are freely available at https://github.com/yangao07/abPOA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 5 |
| 2021 | Erratum to: abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractBioinformatics, 2021, 37(15), 2209–2211. doi: https://doi.org/10.1093/bioinformatics/btaa963 Upon the original publication of this article, the affiliation: “Center for Bioinformatics, Department of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang 150001, China” was inadvertently omitted prior to publication. The affiliation has been reinstated. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 5 |
| 2021 | SKSV: ultrafast structural variation detection from circular consensus sequencing readsabstractSUMMARY: Circular consensus sequencing reads are promising for the comprehensive detection of structural variants (SVs). However, alignment-based SV calling pipelines are computationally intensive due to the generation of complete read-alignments and its post-processing. Herein, we propose a SKeleton-based analysis toolkit for Structural Variation detection (SKSV). Benchmarks on real and simulated datasets demonstrate that SKSV has an order of magnitude of faster speed than state-of-the-art SV calling approaches; moreover, it achieves higher F1 scores for various types of SVs. AVAILABILITY AND IMPLEMENTATION: SKSV is available from https://github.com/ydLiu-HIT/SKSV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yadong Liu 0001, Tao Jiang 0021, Junhao Su, Bo Liu 0023, Tianyi Zang, Yadong Wang 0001 |
Bioinform. | 6 |
| 2021 | Efficient iterative Hi-C scaffolder based on N-best neighborsabstractBACKGROUND: Efficient and effective genome scaffolding tools are still in high demand for generating reference-quality assemblies. While long read data itself is unlikely to create a chromosome-scale assembly for most eukaryotic species, the inexpensive Hi-C sequencing technology, capable of capturing the chromosomal profile of a genome, is now widely used to complete the task. However, the existing Hi-C based scaffolding tools either require a priori chromosome number as input, or lack the ability to build highly continuous scaffolds. RESULTS: We design and develop a novel Hi-C based scaffolding tool, pin_hic, which takes advantage of contact information from Hi-C reads to construct a scaffolding graph iteratively based on N-best neighbors of contigs. Subsequent to scaffolding, it identifies potential misjoins and breaks them to keep the scaffolding accuracy. Through our tests on three long read based de novo assemblies from three different species, we demonstrate that pin_hic is more efficient than current standard state-of-art tools, and it can generate much more continuous scaffolds, while achieving a higher or comparable accuracy. CONCLUSIONS: Pin_hic is an efficient Hi-C based scaffolding tool, which can be useful for building chromosome-scale assemblies. As many sequencing projects have been launched in the recent years, we believe pin_hic has potential to be applied in these projects and makes a meaningful contribution. Dengfeng Guan, Shane A. McCarthy, Zemin Ning, Guohua Wang 0001, Yadong Wang 0001, Richard Durbin |
BMC Bioinform. | 5 |
| 2021 | Correction to: Efficient iterative Hi-C scaffolder based on N-best neighbors
Dengfeng Guan, Shane A. McCarthy, Zemin Ning, Guohua Wang 0001, Yadong Wang 0001, Richard Durbin |
BMC Bioinform. | 5 |
| 2021 | Factor graph-aggregated heterogeneous network embedding for disease-gene association predictionabstractBACKGROUND: Exploring the relationship between disease and gene is of great significance for understanding the pathogenesis of disease and developing corresponding therapeutic measures. The prediction of disease-gene association by computational methods accelerates the process. RESULTS: Many existing methods cannot fully utilize the multi-dimensional biological entity relationship to predict disease-gene association due to multi-source heterogeneous data. This paper proposes FactorHNE, a factor graph-aggregated heterogeneous network embedding method for disease-gene association prediction, which captures a variety of semantic relationships between the heterogeneous nodes by factorization. It produces different semantic factor graphs and effectively aggregates a variety of semantic relationships, by using end-to-end multi-perspectives loss function to optimize model. Then it produces good nodes embedding to prediction disease-gene association. CONCLUSIONS: Experimental verification and analysis show FactorHNE has better performance and scalability than the existing models. It also has good interpretability and can be extended to large-scale biomedical network data analysis. Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
BMC Bioinform. | 4 |
| 2021 | Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluationabstractBACKGROUND: With the rapid development of long-read sequencing technologies, it is possible to reveal the full spectrum of genetic structural variation (SV). However, the expensive cost, finite read length and high sequencing error for long-read data greatly limit the widespread adoption of SV calling. Therefore, it is urgent to establish guidance concerning sequencing coverage, read length, and error rate to maintain high SV yields and to achieve the lowest cost simultaneously. RESULTS: In this study, we generated a full range of simulated error-prone long-read datasets containing various sequencing settings and comprehensively evaluated the performance of SV calling with state-of-the-art long-read SV detection methods. The benchmark results demonstrate that almost all SV callers perform better when the long-read data reach 20× coverage, 20 kbp average read length, and approximately 10-7.5% or below 1% error rates. Furthermore, high sequencing coverage is the most influential factor in promoting SV calling, while it also directly determines the expensive costs. CONCLUSIONS: Based on the comprehensive evaluation results, we provide important guidelines for selecting long-read sequencing settings for efficient SV calling. We believe these recommended settings of long-read sequencing will have extraordinary guiding significance in cutting-edge genomic studies and clinical practices. Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Yadong Wang 0001, Hongzhe Guo |
BMC Bioinform. | 6 |
| 2021 | IIMLP: integrated information-entropy-based method for LncRNA predictionabstractBACKGROUND: The prediction of long non-coding RNA (lncRNA) has attracted great attention from researchers, as more and more evidence indicate that various complex human diseases are closely related to lncRNAs. In the era of bio-med big data, in addition to the prediction of lncRNAs by biological experimental methods, many computational methods based on machine learning have been proposed to make better use of the sequence resources of lncRNAs. RESULTS: We developed the lncRNA prediction method by integrating information-entropy-based features and machine learning algorithms. We calculate generalized topological entropy and generate 6 novel features for lncRNA sequences. By employing these 6 features and other features such as open reading frame, we apply supporting vector machine, XGBoost and random forest algorithms to distinguish human lncRNAs. We compare our method with the one which has more K-mer features and results show that our method has higher area under the curve up to 99.7905%. CONCLUSIONS: We develop an accurate and efficient method which has novel information entropy features to analyze and classify lncRNAs. Our method is also extendable for research on the other functional elements in DNA sequences. Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Qing Liao 0001, Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 11 |
| 2021 | Fast and SNP-aware short read alignment with SALTabstractBACKGROUND: DNA sequence alignment is a common first step in most applications of high-throughput sequencing technologies. The accuracy of sequence alignments directly affects the accuracy of downstream analyses, such as variant calling and quantitative analysis of transcriptome; therefore, rapidly and accurately mapping reads to a reference genome is a significant topic in bioinformatics. Conventional DNA read aligners map reads to a linear reference genome (such as the GRCh38 primary assembly). However, such a linear reference genome represents the genome of only one or a few individuals and thus lacks information on variations in the population. This limitation can introduce bias and impact the sensitivity and accuracy of mapping. Recently, a number of aligners have begun to map reads to populations of genomes, which can be represented by a reference genome and a large number of genetic variants. However, compared to linear reference aligners, an aligner that can store and index all genetic variants has a high cost in memory (RAM) space and leads to extremely long run time. Aligning reads to a graph-model-based index that includes all types of variants is ultimately an NP-hard problem in theory. By contrast, considering only single nucleotide polymorphism (SNP) information will reduce the complexity of the index and improve the speed of sequence alignment. RESULTS: The SNP-aware alignment tool (SALT) is a fast, memory-efficient, and SNP-aware short read alignment tool. SALT uses 5.8 GB of RAM to index a human reference genome (GRCh38) and incorporates 12.8M UCSC common SNPs. Compared with a state-of-the-art aligner, SALT has a similar speed but higher accuracy. CONCLUSIONS: Herein, we present an SNP-aware alignment tool (SALT) that aligns reads to a reference genome that incorporates an SNP database. We benchmarked SALT using simulated and real datasets. The results demonstrate that SALT can efficiently map reads to the reference genome with significantly improved accuracy. Incorporating SNP information can improve the accuracy of read alignment and can reveal novel variants. The source code is freely available at https://github.com/weiquan/SALT . Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2021 | A pipeline for RNA-seq based eQTL analysis with automated quality control proceduresabstractBACKGROUND: Advances in the expression quantitative trait loci (eQTL) studies have provided valuable insights into the mechanism of diseases and traits-associated genetic variants. However, it remains challenging to evaluate and control the quality of multi-source heterogeneous eQTL raw data for researchers with limited computational background. There is an urgent need to develop a powerful and user-friendly tool to automatically process the raw datasets in various formats and perform the eQTL mapping afterward. RESULTS: In this work, we present a pipeline for eQTL analysis, termed eQTLQC, featured with automated data preprocessing for both genotype data and gene expression data. Our pipeline provides a set of quality control and normalization approaches, and utilizes automated techniques to reduce manual intervention. We demonstrate the utility and robustness of this pipeline by performing eQTL case studies using multiple independent real-world datasets with RNA-seq data and whole genome sequencing (WGS) based genotype data. CONCLUSIONS: eQTLQC provides a reliable computational workflow for eQTL analysis. It provides standard quality control and normalization as well as eQTL mapping procedures for eQTL raw data in multiple formats. The source code, demo data, and instructions are freely available at https://github.com/stormlovetao/eQTLQC . Tao Wang 0082, Yongzhuang Liu, Junpeng Ruan, Xianjun Dong, Yadong Wang 0001, Jiajie Peng |
BMC Bioinform. | 5 |
| 2021 | deGSM: Memory Scalable Construction Of Large Scale de Bruijn GraphabstractThe de Bruijn graph, a fundamental data structure to represent and organize genome sequence, plays important roles in various kinds of sequence analysis tasks. With the rapid development of HTS data and ever-increasing number of assembled genomes, there is a high demand to construct the very large de Bruijn graph for sequences up to Tera-base-pair level. Current approaches may have unaffordable memory footprints to handle such a large de Bruijn graph. We propose a lightweight parallel de Bruijn graph construction approach: de Bruijn Graph Constructor in Scalable Memory (deGSM). The main idea of deGSM is to efficiently construct the Burrows-Wheeler Transformation (BWT) of the unipaths of the de Bruijn graph in constant RAM space and transform the BWT into the original unitigs. The experimental results demonstrate that, just with a commonly available machine, deGSM is able to handle very large genome sequence(s), e.g., the contigs (305 Gbp) and scaffolds (1.1 Tbp) recorded in GenBank database and Picea abies HTS dataset (9.7 Tbp). Moreover, deGSM also has faster or comparable construction speed compared with state-of-the-art approaches. With its high scalability and efficiency, deGSM has enormous potential in many large scale genomics studies. The deGSM is publicly available at: https://github.com/hitbc/deGSM. Hongzhe Guo, Yilei Fu, Yan Gao 0031, Junyi Li 0004, Yadong Wang 0001, Bo Liu 0023 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | WMMDCA: Prediction of Drug Responses by Weight-Based Modular Mapping in Cancer Cell LinesabstractDue to the high consumption of cost and time for experimental verification in clinical trials, drug response prediction by computational models have become important challenges. The existing drug response data in diverse cell lines enable prediction of potential sensitive associations. Here, we propose a weight-based modular mapping method, named as WMMDCA, to predict drug-cell line associations. The method fully considers the effects of drugs' chemical structural feature, and adds modular information into the network projection. Leave-one-out cross-validation was used to evaluate the predictive ability of WMMDCA, which showed the best performance among several state-of-the-art methods in not only the whole dataset but also the major tissue types of cell lines. Literature support of highly ranked potential associations was found manually, demonstrating the effectiveness of WMMDCA on drug response prediction. Jie Li 0055, Yadong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Assessment of Machine Learning Methods for Classification in Single Cell ATAC-seqabstractSingle-cell assay for transposase accessible chromatin using sequencing(scATAC-seq) is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. The similarity of data structure and feature between scRNA-seq and scATAC-seq makes it feasible to identify the cell types in scATAC-seq through traditional supervised machine learning methods. Here, we evaluated 6 popular machine learning methods for classification in scATAC-seq. The performance of the methods is evaluated using 4 public single cell ATAC-seq datasets of different tissues, sizes and technologies. We evaluated these methods using intradatasets experiments of 5-folds cross validation based on accuracy, recall and percentage of correctly predicted cells. We found that these methods may perform well in some types of cells in a single dataset, but the overall results are not as well as in scRNA-seq analysis. For testing the classification ability of machine learning methods across datasets, we applied inter-dataset experiments to test the performance of machine learning methods in realistic scenarios. SVM and NMC are overall the top 2 best-performing methods across all experiments. We recommend researchers to apply SVM and NMC as the underlying classifier when developing an automatic classification method in scATAC-seq. Bo Liu 0023, Liran Juan, Tianyi Zang, Tao Jiang 0021, Yadong Wang 0001 |
BIBM | 6 |
| 2020 | GONET: A Deep Network to Annotate Proteins via Recurrent Convolution NetworksabstractFinding out the functions of protein in life activities precisely is nontrivial, which is the core of current proteomics research. Gene Ontology standardizes the function of protein into a series of GO terms, each of which belongs to exactly one of the three subontologies: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF). The prediction of protein function can be considered as a multi-label classification problem. Traditional methods often spend a lot of costs to extract handcrafted features and plenty of domain knowledge is needed when solving these tasks, while using deep learning technology can overcome these shortcomings. Here, we propose a deep model GONET based on recurrent convolutional neural networks, which annotates protein in an end-to-end manner. Our model combines protein sequences and protein-protein interaction (PPI) network data, and utilizes representation learning to learn distributed representation of proteins to overcome the sparse nature and semantic independence problem. Moreover, we adopt a quite deep CNNRNN-Attention model, which is able to effectively extract high-order features of protein sequences. We have carried out experiments on several datasets, which achieve the state-of-the-art in some metrics compared with the existing competitive methods. Junyi Li 0004, Xiaoshuai Zhang, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 5 |
| 2020 | An Approximate Multiplane Network-on-ChipabstractThe increasing communication demands in chip multiprocessors (CMPs) and many error-tolerant applications are driving the approximate design of the network-on-chip (NoC) for power-efficient packet delivery. However, current approximate NoC designs achieve improvements in network performance or dynamic power savings at the cost of additional circuit design and increased area overhead. In this paper, we propose a novel approximate multiplane NoC (AMNoC) that provides low-latency transfer for latency-sensitive packets and minimizes the power consumption of approximable packets through a lossy bufferless subnetwork. The AMNoC also includes a regular buffered subnetwork to guarantee the lossless delivery of nonapproximable packets. Evaluations show that, compared with a single-plane buffered NoC, the AMNoC reduces the average latency by 41.9%. In addition, the AMNoC achieves 48.6% and 53.4% savings in power consumption and area overhead, respectively. Ling Wang 0005, Yadong Wang 0001, Xiaohang Wang 0001 |
DATE | 2 |
| 2020 | HGAlinker: Drug-Disease Association Prediction Based on Attention Mechanism of Heterogeneous Graph
Xiaozhu Jing, Wei Jiang 0023, Zhongqing Zhang, Yadong Wang 0001, Junyi Li 0004 |
ICIC (2) | 4 |
| 2020 | Identifying and removing haplotypic duplication in primary genome assembliesabstractMOTIVATION: Rapid development in long-read sequencing and scaffolding technologies is accelerating the production of reference-quality assemblies for large eukaryotic genomes. However, haplotype divergence in regions of high heterozygosity often results in assemblers creating two copies rather than one copy of a region, leading to breaks in contiguity and compromising downstream steps such as gene annotation. Several tools have been developed to resolve this problem. However, they either focus only on removing contained duplicate regions, also known as haplotigs, or fail to use all the relevant information and hence make errors. RESULTS: Here we present a novel tool, purge_dups, that uses sequence similarity and read depth to automatically identify and remove both haplotigs and heterozygous overlaps. In comparison with current tools, we demonstrate that purge_dups can reduce heterozygous duplication and increase assembly continuity while maintaining completeness of the primary assembly. Moreover, purge_dups is fully automatic and can easily be integrated into assembly pipelines. AVAILABILITY AND IMPLEMENTATION: The source code is written in C and is available at https://github.com/dfguan/purge_dups. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dengfeng Guan, Shane A. McCarthy, Jonathan Wood, Kerstin Howe, Yadong Wang 0001, Richard Durbin |
Bioinform. | 5 |
| 2020 | Evaluating individual genome similarity with a topic modelabstractMOTIVATION: Evaluating genome similarity among individuals is an essential step in data analysis. Advanced sequencing technology detects more and rarer variants for massive individual genomes, thus enabling individual-level genome similarity evaluation. However, the current methodologies, such as the principal component analysis (PCA), lack the capability to fully leverage rare variants and are also difficult to interpret in terms of population genetics. RESULTS: Here, we introduce a probabilistic topic model, latent Dirichlet allocation, to evaluate individual genome similarity. A total of 2535 individuals from the 1000 Genomes Project (KGP) were used to demonstrate our method. Various aspects of variant choice and model parameter selection were studied. We found that relatively rare (0.001 20 000 bp) variants are more efficient for genome similarity evaluation. At least 100 000 such variants are necessary. In our results, the populations show significantly less mixed and more cohesive visualization than the PCA results. The global similarities among the KGP genomes are consistent with known geographical, historical and cultural factors. AVAILABILITY AND IMPLEMENTATION: The source code and data access are available at: https://github.com/lrjuan/LDA_genome. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Liran Juan, Yongtian Wang, Jingyi Jiang, Guohua Wang 0001, Yadong Wang 0001 |
Bioinform. | 6 |
| 2020 | Filtering de novo indels in parent-offspring triosabstractBACKGROUND: Identification of de novo indels from whole genome or exome sequencing data of parent-offspring trios is a challenging task in human disease studies and clinical practices. Existing computational approaches usually yield high false positive rate. RESULTS: In this study, we developed a gradient boosting approach for filtering de novo indels obtained by any computational approaches. Through application on the real genome sequencing data, our approach showed it could significantly reduce the false positive rate of de novo indels without a significant compromise on sensitivity. CONCLUSIONS: The software DNMFilter_Indel was written in a combination of Java and R and freely available from the website at https://github.com/yongzhuang/DNMFilter_Indel . Yongzhuang Liu, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2020 | Mining Relationships among Multiple Entities in Biological NetworksabstractIdentifying topological relationships among multiple entities in biological networks is critical towards the understanding of the organizational principles of network functionality. Theoretically, this problem can be solved using minimum Steiner tree (MSTT) algorithms. However, due to large network size, it remains to be computationally challenging, and the predictive value of multi-entity topological relationships is still unclear. We present a novel solution called Cluster-based Steiner Tree Miner (CST-Miner) to instantly identify multi-entity topological relationships in biological networks. Given a list of user-specific entities, CST-Miner decomposes a biological network into nested cluster-based subgraphs, on which multiple minimum Steiner trees are identified. By merging all of them into a minimum cost tree, the optimal topological relationships among all the user-specific entities are revealed. Experimental results showed that CST-Miner can finish in nearly log-linear time and the tree constructed by CST-Miner is close to the global minimum. Jiajie Peng, Linjiao Zhu, Yadong Wang 0001, Jin Chen 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Y-SPCR: A new dimensionality reduction method for gene expression data classificationabstractWith the increase of the scale and complexity of massive data, data dimensionality reduction technologies, such as principal component analysis, have developed rapidly. The performance of dimension reduction technologies still needs to be further improved. In the paper we proposed a new dimensionality reduction method (Y-SPCR) based Supervised Principal Component Regression (SPCR) and Y-aware Principal Component Regression (Y-aware PCR). Experimental results on four gene expression data sets show that Y-SPCR effectively overcomes the shortcomings of SPCR and Y-aware PCR and improves the accuracy and stability of on gene expression data classification. Jie Li 0055, Zhun Zhao, Yadong Wang 0001 |
BIBM | 4 |
| 2019 | DNMFilter_Indel: Filtering de novo Indels in Parent-Offspring TriosabstractIdentification of de novo indels from whole genome or exome sequencing data of parent-offspring trios is a challenging task in human disease studies and clinical practices. Existing computational approaches usually yield high false positive rate. In this study, we developed a gradient boosting approach for filtering de novo indels obtained by any computational approaches. Through application on the real genome sequencing data, our approach showed it could significantly reduce the false positive rate of de novo indels without a significant compromise on sensitivity. The software DNMFilter_Indel was written in a combination of Java and R and freely available from the website at https://github.com/yongzhuang/DNMFilter_Indel. Yongzhuang Liu, Yadong Wang 0001 |
BIBM | 3 |
| 2019 | SALT: a fast, memory-efficient and SNP-aware short read alignment toolabstractDNA sequence alignment tools play an essential role in genomics and genetics. The accuracy of the alignment directly affects the accuracy of downstream analysis, such as variant calling, so it is essential to map reads to the reference genome rapidly and accurately. It has become an essential topic in the field of bioinformatics. Conventional read aligners map reads to a linear reference genome (such as GRCh38 primary). However, the linear reference genome only represents one or a few individuals of genomes, which lacks the variation information in population. It can introduce bias and impact sensitivity and accuracy of mapping. Recently, a few aligners are beginning to map reads to a graph that captures the entire human genome along with a large number of variants. However, compared to linear reference aligners, storing and indexing all genetic variants require costly memory(RAM) space and make extremely long runtime. Aligning reads to a graph model-based index, including the whole set of variants, is ultimately an NP-hard problem in theory. Considering only SNPs information will reduce the complexity of index and improve the speed of alignments. Herein, we present an SNP-aware alignment tool (SALT) that aligns reads to a reference genome that incorporates the SNP database. The SALT is benchmarked both on simulated reads and the real dataset. The results demonstrate that SALT can efficiently map reads to the reference genome, and significantly improve accuracy and sensitivity. Read alignment incorporating SNPs information can improve the sensitivity and accuracy of the read alignment. Moreover, it helps to discover novel variants. SALT is distributed under the GNU General Public License (GPL). Source code is freely available at https://github.com/weiquan/SALT. Bo Liu 0023, Yadong Wang 0001 |
BIBM | 3 |
| 2019 | A Bidirectional Fuzzy Index and Approximate Search Algorithm for Next Generation SequencingabstractSequence alignment is one of the most important problems in bioinformatics. However, the existing alignment tools may result in a large number of candidate locations which degradate alignment performance. Recent researches discard the high repetitive seeds for improving the alignment speed, which influences alignment accuracy. To this end, we propose a novel fuzzy index ( fBWT), which is available at https://github.com/weiquan/appr_bwt. It allows approximate search and extending the length of seeds to reduce the candidate locations and accelerate the sequence alignment. The performance of our tool was compared with BWA using 150bp and 250bp length datasets. The result shows the number of misaligned reads (the correct position is not included in the high-score candidate position set) of the current mainstream tool is 5-10 times higher than it. The efficiency between the presented and the existing tools are also compared. Under the above conditions, the alignment time of the presented tool is very close. However, the alignment speed of fBWT is much faster than BWA under the requirement of similar alignment accuracy. Guangri Quan, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 4 |
| 2019 | An automated quality control pipeline for eQTL analysis with RNA-seq dataabstractExpression quantitative trait loci (eQTL) analysis is of critical importance to understand the mechanism underlying trait associated variants. Evaluating and controlling the data quality of transcripts and genotypes, which are basis of eQTL analysis, remains challenging for researchers with limited computational backgrounds. There is a strong need for a user-friendly and comprehensive tool to pre-process those data sets automatically. Here we propose such a solution, eQTLQC, an automated quality control pipeline for preprocessing both RNA-seq and genotype data. The eQTLQC pipeline provides multiple informative quality control measurements and data normalization approaches. And it provides a easy-to-use configuration file for users to flexibly set up the parameters and control the pipeline. We demonstrate its utility by performing RNA-seq and genotype preprocessing on real data sets. eQTLQC is open source and freely available at https://github.com/ruanjunpeng/eQTLQC. Tao Wang 0082, Junpeng Ruan, Quanwei Yin, Xianjun Dong, Yadong Wang 0001 |
BIBM | 5 |
| 2019 | Prediction of Human LncRNAs Based on Integrated Information Entropy Features
Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Bo Liu 0023, Yadong Wang 0001 |
ICIC (2) | 9 |
| 2019 | TideHunter: efficient and sensitive tandem repeat detection from noisy long-reads using seed-and-chainabstractMOTIVATION: Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) sequencing technologies can produce long-reads up to tens of kilobases, but with high error rates. In order to reduce sequencing error, Rolling Circle Amplification (RCA) has been used to improve library preparation by amplifying circularized template molecules. Linear products of the RCA contain multiple tandem copies of the template molecule. By integrating additional in silico processing steps, these tandem sequences can be collapsed into a consensus sequence with a higher accuracy than the original raw reads. Existing pipelines using alignment-based methods to discover the tandem repeat patterns from the long-reads are either inefficient or lack sensitivity. RESULTS: We present a novel tandem repeat detection and consensus calling tool, TideHunter, to efficiently discover tandem repeat patterns and generate high-quality consensus sequences from amplified tandemly repeated long-read sequencing data. TideHunter works with noisy long-reads (PacBio and ONT) at error rates of up to 20% and does not have any limitation of the maximal repeat pattern size. We benchmarked TideHunter using simulated and real datasets with varying error rates and repeat pattern sizes. TideHunter is tens of times faster than state-of-the-art methods and has a higher sensitivity and accuracy. AVAILABILITY AND IMPLEMENTATION: TideHunter is written in C, it is open source and is available at https://github.com/yangao07/TideHunter. Yan Gao 0031, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 3 |
| 2019 | rMETL: sensitive mobile element insertion detection with long read realignmentabstractSUMMARY: Mobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing technologies provides the opportunity to detect MEIs sensitively. However, the signals of MEI implied by noisy long reads are highly complex due to the repetitiveness of mobile elements as well as the high sequencing error rates. Herein, we propose the Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). Benchmarking results of simulated and real datasets demonstrate that rMETL enables to handle the complex signals to discover MEIs sensitively. It is suited to produce high-quality MEI callsets in many genomics studies. AVAILABILITY AND IMPLEMENTATION: rMETL is available from https://github.com/hitbc/rMETL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Jiang 0021, Bo Liu 0023, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 4 |
| 2019 | Joint detection of germline and somatic copy number events in matched tumor-normal sample pairsabstractMOTIVATION: Whole-genome sequencing (WGS) of tumor-normal sample pairs is a powerful approach for comprehensively characterizing germline copy number variations (CNVs) and somatic copy number alterations (SCNAs) in cancer research and clinical practice. Existing computational approaches for detecting copy number events cannot detect germline CNVs and SCNAs simultaneously, and yield low accuracy for SCNAs. RESULTS: In this study, we developed TumorCNV, a novel approach for jointly detecting germline CNVs and SCNAs from WGS data of the matched tumor-normal sample pair. We compared TumorCNV with existing copy number event detection approaches using the simulated data and real data for the COLO-829 melanoma cell line. The experimental results showed that TumorCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TumorCNV is implemented using a combination of Java and R, and it is freely available from the website at https://github.com/yongzhuang/TumorCNV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongzhuang Liu, Yadong Wang 0001 |
Bioinform. | 3 |
| 2019 | Integrated entropy-based approach for analyzing exons and introns in DNA sequencesabstractBACKGROUND: Numerous essential algorithms and methods, including entropy-based quantitative methods, have been developed to analyze complex DNA sequences since the last decade. Exons and introns are the most notable components of DNA and their identification and prediction are always the focus of state-of-the-art research. RESULTS: In this study, we designed an integrated entropy-based analysis approach, which involves modified topological entropy calculation, genomic signal processing (GSP) method and singular value decomposition (SVD), to investigate exons and introns in DNA sequences. We optimized and implemented the topological entropy and the generalized topological entropy to calculate the complexity of DNA sequences, highlighting the characteristics of repetition sequences. By comparing digitalizing entropy values of exons and introns, we observed that they are significantly different. After we converted DNA data to numerical topological entropy value, we applied SVD method to effectively investigate exon and intron regions on a single gene sequence. Additionally, several genes across five species are used for exon predictions. CONCLUSIONS: Our approach not only helps to explore the complexity of DNA sequence and its functional elements, but also provides an entropy-based GSP method to analyze exon and intron regions. Our work is feasible across different species and extendable to analyze other components in both coding and noncoding region of DNA sequences. Junyi Li 0004, Huinian Li, Qingzhe Xu, Renjie Tan, Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 10 |
| 2019 | Prioritizing candidate diseases-related metabolites based on literature and functional similarityabstractBACKGROUND: As the terminal products of cellular regulatory process, functional related metabolites have a close relationship with complex diseases, and are often associated with the same or similar diseases. Therefore, identification of disease related metabolites play a critical role in understanding comprehensively pathogenesis of disease, aiming at improving the clinical medicine. Considering that a large number of metabolic markers of diseases need to be explored, we propose a computational model to identify potential disease-related metabolites based on functional relationships and scores of referred literatures between metabolites. First, obtaining associations between metabolites and diseases from the Human Metabolome database, we calculate the similarities of metabolites based on modified recommendation strategy of collaborative filtering utilizing the similarities between diseases. Next, a disease-associated metabolite network (DMN) is built with similarities between metabolites as weight. To improve the ability of identifying disease-related metabolites, we introduce scores of text mining from the existing database of chemicals and proteins into DMN and build a new disease-associated metabolite network (FLDMN) by fusing functional associations and scores of literatures. Finally, we utilize random walking with restart (RWR) in this network to predict candidate metabolites related to diseases. RESULTS: We construct the disease-associated metabolite network and its improved network (FLDMN) with 245 diseases, 587 metabolites and 28,715 disease-metabolite associations. Subsequently, we extract training sets and testing sets from two different versions of the Human Metabolome database and assess the performance of DMN and FLDMN on 19 diseases, respectively. As a result, the average AUC (area under the receiver operating characteristic curve) of DMN is 64.35%. As a further improved network, FLDMN is proven to be successful in predicting potential metabolic signatures for 19 diseases with an average AUC value of 76.03%. CONCLUSION: In this paper, a computational model is proposed for exploring metabolite-disease pairs and has good performance in predicting potential metabolites related to diseases through adequate validation. This result suggests that integrating literature and functional associations can be an effective way to construct disease associated metabolite network for prioritizing candidate diseases-related metabolites. Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001 |
BMC Bioinform. | 5 |
| 2019 | LncDisAP: a computation model for LncRNA-disease association prediction based on multiple biological datasetsabstractBACKGROUND: Over the past decades, a large number of long non-coding RNAs (lncRNAs) have been identified. Growing evidence has indicated that the mutation and dysregulation of lncRNAs play a critical role in the development of many complex human diseases. Consequently, identifying potential disease-related lncRNAs is an effective means to improve the quality of disease diagnostics and treatment, which is the motivation of this work. Here, we propose a computational model (LncDisAP) for potential disease-related lncRNA identification based on multiple biological datasets. First, the associations between lncRNA and different data sources are collected from different databases. With these data sources as dimensions, we calculate the functional associations between lncRNAs by the recommendation strategy of collaborative filtering. Subsequently, a disease-associated lncRNA functional network is built with functional similarities between lncRNAs as the weight. Ultimately, potential disease-related lncRNAs can be identified based on ranked scores derived by random walking with restart (RWR). Then, training sets and testing sets are extracted from two different versions of a disease-lncRNA dataset to assess the performance of LncDisAP on 54 diseases. RESULTS: A lncRNA functional network is built based on the proposed computational model, and it contains 66,060 associations among 364 lncRNAs associated with 182 diseases in total. We extract 218 known disease-lncRNA pairs associated with 54 diseases to assess the network. As a result, the average AUC (area under the receiver operating characteristic curve) of LncDisAP is 78.08%. CONCLUSION: In this article, a computational model integrating multiple lncRNA-related biological datasets is proposed for identifying potential disease-related lncRNAs. The result shows that LncDisAP is successful in predicting novel disease-related lncRNA signatures. In addition, with several common cancers taken as case studies, we found some unknown lncRNAs that could be associated with these diseases through our network. These results suggest that this method can be helpful in improving the quality for disease diagnostics and treatment. Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001 |
BMC Bioinform. | 5 |
| 2018 | A Network-on-Chip Accelerator for Genome Variant Analysis
Ling Wang 0005, Yadong Wang 0001 |
BIBM | 2 |
| 2018 | rCANID: read Clustering and Assembly-based Novel Insertion Detection tool
Yilei Fu, Tao Jiang 0021, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 4 |
| 2018 | deSPI: efficient classification of metagenomics reads with lightweight de Bruijn graph-based reference indexing
Dengfeng Guan, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 3 |
| 2018 | Fast variation-aware read alignment with deBGA-VARA
Hongzhe Guo, Bo Liu 0023, Dengfeng Guan, Yilei Fu, Yadong Wang 0001 |
BIBM | 5 |
| 2018 | DeepDNA: a hybrid convolutional and recurrent neural network for compressing human mitochondrial genomes
Yan-Shuo Chu, Yongtian Wang, Mingrui Sun, Junyi Li 0004, Tianyi Zang, Yadong Wang 0001 |
BIBM | 9 |
| 2018 | Predicting candidate disease-related lncRNAs based on network random walk
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001 |
BIBM | 5 |
| 2018 | Identifying Representative Network Motifs for Inferring Higher-order Structure of Biological Networks
Tao Wang 0082, Jiajie Peng, Yadong Wang 0001, Jin Chen 0004 |
BIBM | 3 |
| 2018 | Modeling and correct the GC bias of tumor and normal WGS data for SCNA based tumor subclonal population inferringabstractBACKGROUND: Somatic copy number alternations (SCNAs) can be utilized to infer tumor subclonal populations in whole genome seuqncing studies, where usually their read count ratios between tumor-normal paired samples serve as the inferring proxy. Existing SCNA based subclonal population inferring tools consider the GC bias of tumor and normal sample is of the same fature, and could be fully offset by read count ratio. However, we found that, the read count ratio on SCNA segments presents a Log linear biased pattern, which influence existing read count ratios based subclonal inferring tools performance. Currently no correction tools take into account the read ratio bias. RESULTS: We present Pre-SCNAClonal, a tool that improving tumor subclonal population inferring by correcting GC-bias at SCNAs level. Pre-SCNAClonal first corrects GC bias using Markov chain Monte Carlo probability model, then accurately locates baseline DNA segments (not containing any SCNAs) with a hierarchy clustering model. We show Pre-SCNAClonal's superiority to exsiting GC-bias correction methods at any level of subclonal population. CONCLUSIONS: Pre-SCNAClonal could be run independently as well as serving as pre-processing/gc-correction step in conjuntion with exsiting SCNA-based subclonal inferring tools. Yan-Shuo Chu, Mingxiang Teng, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2017 | Pre-SCNAClonal: Efficient GC bias correction for SCNA based tumor subclonal populations inferringabstractSomatic copy number alternations (SCNAs) can be utilized to infer tumor subclonal populations in whole genome seuqncing studies, where usually their read count ratios between tumor-normal paired samples serve as the inferring proxy. We found that, in a GC study, the GC contents and read count ratios on SCNA segments present a Log linear biased pattern. However, currently no subclonal inferring tools take into account this information. We provide Pre-SCNAClonal, a comprehensive GC bias correction tool for inferring tumor subclonal populations based on SCNAs. Results show that Pre-SCNAClonal could effectively and robustly correct the GC bias and improve the performance of the SCNAs based tumor subclonal population inferring tools. Pre-SCNAClonal could be strung together with the SCNAs based subclonal population inferring tool as a pipeline or run individually as needed. Yan-Shuo Chu, Mingxiang Teng, Yongtian Wang, Yadong Wang 0001 |
BIBM | 5 |
| 2017 | Pysubsim-tree: A package for simulating tumor genomes according to tumor evolution historyabstractNext-generation sequencing (NGS) and the third generation sequencing (TGS) have recently allowed us to develop algorithms to quantitatively dissect the extent of heterogeneity within a tumour, resolve cancer evolution history and identify the somatic variations and aneuploidy events. Simulation of tumor NGS and TGS data serves as a powerful and cost-effective approach for benchmarking these algorithms, however, there is no available tool that could simulate all the distinct subclonal genomes with diverse aneuploidy events and somatic variations according to the given tumor evolution history. We provide a simulation package, Pysubsim-tree, which could simulate the tumor genomes according to their evolution history defined by the somatic variations and aneuploidy events. Pysubsim-tree is free, open source, available at: https://github.com/dustincys/pysubsimtree. Yan-Shuo Chu, Ling Wang 0005, Mingxiang Teng, Yadong Wang 0001 |
BIBM | 5 |
| 2017 | ABDTR: Approximation-Based Dynamic Traffic Regulation for Networks-on-Chip SystemsabstractTraffic regulation is an essential technology of networks-on-chip (NoC) to achieve communication performance guarantees with effective use of the system interconnect and low traffic delay. This paper presents approximation-based dynamic traffic regulation (ABDTR), which approximates part of traffic data instead of network transmission for mitigating network congestion based on the inherent error resilience of some applications. ABDTR drops a fraction of safe-to-approximate packet data that may aggravate network congestion before they are injected into network, and predicts the lost data in packet after being received in destination node. Traffic drop rate is fully adaptive to the traffic and network states. The regulation range of drop rate is a knob to control the trade-off between performance/energy efficiency and output quality. The proposed method is effective and can be implemented in hardware with small area. Experimental results have confirmed that the ABDTR can help reduce up to 44.4 percent average network delay with less than 10% loss in quality. The results also show ABDTR yields significant application speedup and NoC energy reduction for a wide range of quality-loss levels. Ling Wang 0005, Xiaohang Wang 0001, Yadong Wang 0001 |
ICCD | 3 |
| 2017 | LAMSA: fast split read alignment with long approximate matchesabstractMOTIVATION: Read length is continuously increasing with the development of novel high-throughput sequencing technologies, which has enormous potentials on cutting-edge genomic studies. However, longer reads could more frequently span the breakpoints of structural variants (SVs) than that of shorter reads. This may greatly influence read alignment, since most state-of-the-art aligners are designed for handling relatively small variants in a co-linear alignment framework. Meanwhile, long read alignment is still not as efficient as that of short reads, which could be also a bottleneck for the upcoming wide application. RESULTS: We propose long approximate matches-based split aligner (LAMSA), a novel split read alignment approach. It takes the advantage of the rareness of SVs to implement a specifically designed two-step strategy. That is, LAMSA initially splits the read into relatively long fragments and co-linearly align them to solve the small variations or sequencing errors, and mitigate the effect of repeats. The alignments of the fragments are then used for implementing a sparse dynamic programming-based split alignment approach to handle the large or non-co-linear variants. We benchmarked LAMSA with simulated and real datasets having various read lengths and sequencing error rates, the results demonstrate that it is substantially faster than the state-of-the-art long read aligners; meanwhile, it also has good ability to handle various categories of SVs. AVAILABILITY AND IMPLEMENTATION: LAMSA is available at https://github.com/hitbc/LAMSA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Yan Gao 0031, Yadong Wang 0001 |
Bioinform. | 3 |
| 2017 | rMFilter: acceleration of long read-based structure variation calling by chimeric read filteringabstractMOTIVATION: Long read sequencing technologies provide new opportunities to investigate genome structural variations (SVs) more accurately. However, the state-of-the-art SV calling pipelines are computational intensive and the applications of long reads are restricted. RESULTS: We propose a local region match-based filter (rMFilter) to efficiently nail down chimeric noisy long reads based on short token matches within local genomic regions. rMFilter is able to substantially accelerate long read-based SV calling pipelines without loss of effectiveness. It can be easily integrated into current long read-based pipelines to facilitate SV studies. AVAILABILITY AND IMPLEMENTATION: The C ++ source code of rMFilter is available at https://github.com/hitbc/rMFilter . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Tao Jiang 0021, Siu-Ming Yiu, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 5 |
| 2017 | Identifying term relations cross different gene ontology categoriesabstractBACKGROUND: The Gene Ontology (GO) is a community-based bioinformatics resource that employs ontologies to represent biological knowledge and describes information about gene and gene product function. GO includes three independent categories: molecular function, biological process and cellular component. For better biological reasoning, identifying the biological relationships between terms in different categories are important. However, the existing measurements to calculate similarity between terms in different categories are either developed by using the GO data only or only take part of combined gene co-function network information. RESULTS: We propose an iterative ranking-based method called C r o G O2 to measure the cross-categories GO term similarities by incorporating level information of GO terms with both direct and indirect interactions in the gene co-function network. CONCLUSIONS: The evaluation test shows that C r o G O2 performs better than the existing methods. A genome-specific term association network for yeast is also generated by connecting terms with the high confidence score. The linkages in the term association network could be supported by the literature. Given a gene set, the related terms identified by using the association network have overlap with the related terms identified by GO enrichment analysis. Jiajie Peng, Junya Lu, Weiwei Hui, Yadong Wang 0001, Xuequn Shang 0001 |
BMC Bioinform. | 5 |
| 2017 | DTWscore: differential expression and cell clustering analysis for time-series single-cell RNA-seq dataabstractBACKGROUND: The development of single-cell RNA sequencing has enabled profound discoveries in biology, ranging from the dissection of the composition of complex tissues to the identification of novel cell types and dynamics in some specialized cellular environments. However, the large-scale generation of single-cell RNA-seq (scRNA-seq) data collected at multiple time points remains a challenge to effective measurement gene expression patterns in transcriptome analysis. RESULTS: We present an algorithm based on the Dynamic Time Warping score (DTWscore) combined with time-series data, that enables the detection of gene expression changes across scRNA-seq samples and recovery of potential cell types from complex mixtures of multiple cell types. CONCLUSIONS: The DTWscore successfully classify cells of different types with the most highly variable genes from time-series scRNA-seq data. The study was confined to methods that are implemented and available within the R framework. Sample datasets and R packages are available at https://github.com/xiaoxiaoxier/DTWscore . Shuilin Jin, Guiyou Liu, Xiurui Zhang, Deliang Wu, Yang Hu 0008, Chiping Zhang, Qinghua Jiang, Yadong Wang 0001 |
BMC Bioinform. | 11 |
| 2016 | DisSetSim: An online system for calculating similarity between disease setsabstractFunctional similarity between molecules results in similar phenotypes, such as diseases. Therefore, it is an effective way to reveal the function of molecules based on their induced diseases. However, the lack of a tool for obtaining the similarity score of pair-wise disease sets (SSDS) limits this type of application. Here, we introduce DisSetSim, an online system to solve this problem in this article. Five state-of-the-art methods involving Resnik's, Lin's, Wang's, PSB, and SemFunSim methods were implemented to measure the similarity score of pair-wise diseases (SSD) first. And then “pair-wise-best pairs-average” (PWBPA) method was implemented to calculated the SSDS by the SSD. The system was applied for calculating the functional similarity of miRNAs based on their induced disease sets. The results were further used to predict potential disease-miRNA relationships. The high area under the receiver operating characteristic curve AUC (0.9296) based on leave-one-out cross validation shows that the PWBPA method achieves a high true positive rate and a low false positive rate. The system can be accessed from http://bio-annotation.cn/DisSetSim. Yang Hu 0008, Lingling Zhao, Zhiyan Liu, Hong Ju, Peigang Xu, Yadong Wang 0001, Liang Cheng 0006 |
BIBM | 7 |
| 2016 | Measuring phenotype semantic similarity using Human Phenotype OntologyabstractIt is critical yet remains to be challenging to make right disease diagnosis based on complex clinical characteristic and heterogeneous genetic background. Recently, Human Phenotype Ontology (HPO)-based phenotype similarity has been widely used to aid disease diagnosis. However, the existing measurements are revised based on the Gene Ontology-based term similarity models, which are not optimized for human phenotype ontologies. We propose a new similarity measure called PhenoSim. Our model includes a noise reduction component to model the noisy patient phenotype data, and a path-constrained Information Content-based method for measuring phenotype semantics similarity. Evaluation tests showed that PhenoSim could improve the performance of HPO-based phenotype similarity measurement. Jiajie Peng, Hansheng Xue, Yukai Shao, Xuequn Shang 0001, Yadong Wang 0001, Jin Chen 0004 |
BIBM | 5 |
| 2016 | ERDS-pe: A paired hidden Markov model for copy number variant detection from whole-exome sequencing dataabstractDetecting copy number variants (CNVs) is an essential part in variant calling process. Here, we describe a novel method ERDS-pe to detect CNVs from whole-exome sequencing (WES) data. ERDS-pe first employs principal component analysis to normalize WES data. Then, ERDS-pe incorporates read depth signal and single-nucleotide variation information together as a hybrid signal into a paired hidden Markov model to infer CNVs from WES data. Experimental results on real human WES data show that ERDS-pe demonstrates higher sensitivity and provides comparable or even better specificity than other tools. ERDS-pe is publicly available at: https://github.com/microtan0902/erds-pe. Renjie Tan, Jixuan Wang, Guoqiang Wan, Zhijie Han 0002, Wenyang Zhou, Shuilin Jin, Qinghua Jiang, Yadong Wang 0001 |
BIBM | 11 |
| 2016 | PrefaceabstractWelcome to the 2016 IEEE International Conference on Bioinformatics and Biomedicine (IEEE BIBM 2016) being held in the Shenzhen, China from December 15–18, 2016. On behalf of the IEEE BIBM 2016 Organizing Team, we would like to thank you for your participation and hope you enjoy the conference. Yadong Wang 0001, Kevin Burrage, Shinichi Morishita, Tianhai Tian, Qinghua Jiang, Jiangning Song, Guohua Wang 0001, Xiaohua Hu 0001 |
BIBM | 1 |
| 2016 | DMcompress: Dynamic Markov models for bacterial genome compressionabstractGenome data increasing exponentially since the last decade, compressing genome with Markov models has been proposed as an effective statistical method. However, existing methods set a static order-k Markov models to compress various genomes. Employing static order-k Markov model could result in a sub-optimal orders on some genomes. In this paper, we propose a compression method that relies on a pre-analysis of the data before compression, with the aim of estimating Markov models order k, yielding improvements over static Markov models. Experimental results on the latest complete bacterial genome data show that our method could effectively compress genome with a better performance than the state-of-the-art method. The codes of DMcompress are available at https://rongjiewang.github.io/DMcompress. Mingxiang Teng, Tianyi Zang, Yadong Wang 0001 |
BIBM | 5 |
| 2016 | Accurate annotation of metagenomic data without species-level referencesabstractTaxonomic annotation is a critical first step for analysis of metagenomic data. Despite a lot of tools being developed, the accuracy is still not satisfactory, in particular, when a close species-level reference does not exist in the database. In this paper, we propose a novel annotation tool, MetaAnnotator, to annotate metagenomic reads, which outperforms all existing tools significantly when only genus-level references exist in the database. From our experiments, MetaAnnotator can assign 87.5% reads correctly (67.5% reads are assigned to the exact genus) with only 8.5% reads wrongly assigned. The best existing tool (MetaCluster-TA) can only achieve 73.4% correct read assignment (with only 50.9% reads assigned to the exact genus and 22.6% reads wrongly assigned). The speed of MetaAnnotator is also the second faster (1 hour for 20 million reads). The core concepts behind MetaAnnotator includes: (i) we only consider exact k-mers in coding regions of the references as they should be more significant and accurate; (ii) to assign reads to taxonomy nodes, we construct genome and taxonomy specific probabilistic models from the reference database; and (iii) using the BWT data structure to speed up the k-mer matching process. Haobin Yao, Tak Wah Lam, Hing-Fung Ting, Siu-Ming Yiu, Yadong Wang 0001, Bo Liu 0023 |
BIBM | 5 |
| 2016 | deBGA: read alignment with de Bruijn graph-based seed and extensionabstractMOTIVATION: As high-throughput sequencing (HTS) technology becomes ubiquitous and the volume of data continues to rise, HTS read alignment is becoming increasingly rate-limiting, which keeps pressing the development of novel read alignment approaches. Moreover, promising novel applications of HTS technology require aligning reads to multiple genomes instead of a single reference; however, it is still not viable for the state-of-the-art aligners to align large numbers of reads to multiple genomes. RESULTS: We propose de Bruijn Graph-based Aligner (deBGA), an innovative graph-based seed-and-extension algorithm to align HTS reads to a reference genome that is organized and indexed using a de Bruijn graph. With its well-handling of repeats, deBGA is substantially faster than state-of-the-art approaches while maintaining similar or higher sensitivity and accuracy. This makes it particularly well-suited to handle the rapidly growing volumes of sequencing data. Furthermore, it provides a promising solution for aligning reads to multiple genomes and graph-based references in HTS applications. AVAILABILITY AND IMPLEMENTATION: deBGA is available at: https://github.com/hitbc/deBGA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Hongzhe Guo, Michael Brudno, Yadong Wang 0001 |
Bioinform. | 4 |
| 2016 | rHAT: fast alignment of noisy long reads with regional hashingabstractMOTIVATION: Single Molecule Real-Time (SMRT) sequencing has been widely applied in cutting-edge genomic studies. However, it is still an expensive task to align the noisy long SMRT reads to reference genome by state-of-the-art aligners, which is becoming a bottleneck in applications with SMRT sequencing. Novel approach is on demand for improving the efficiency and effectiveness of SMRT read alignment. RESULTS: We propose Regional Hashing-based Alignment Tool (rHAT), a seed-and-extension-based read alignment approach specifically designed for noisy long reads. rHAT indexes reference genome by regional hash table (RHT), a hash table-based index which describes the short tokens within local windows of reference genome. In the seeding phase, rHAT utilizes RHT for efficiently calculating the occurrences of short token matches between partial read and local genomic windows to find highly possible candidate sites. In the extension phase, a sparse dynamic programming-based heuristic approach is used for reducing the cost of aligning read to the candidate sites. By benchmarking on the real and simulated datasets from various prokaryote and eukaryote genomes, we demonstrated that rHAT can effectively align SMRT reads with outstanding throughput. AVAILABILITY AND IMPLEMENTATION: rHAT is implemented in C++; the source code is available at https://github.com/HIT-Bioinformatics/rHAT CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Dengfeng Guan, Mingxiang Teng, Yadong Wang 0001 |
Bioinform. | 4 |
| 2016 | Joint detection of copy number variations in parent-offspring triosabstractMOTIVATION: Whole genome sequencing (WGS) of parent-offspring trios is a powerful approach for identifying disease-associated genes via detecting copy number variations (CNVs). Existing approaches, which detect CNVs for each individual in a trio independently, usually yield low-detection accuracy. Joint modeling approaches leveraging Mendelian transmission within the parent-offspring trio can be an efficient strategy to improve CNV detection accuracy. RESULTS: In this study, we developed TrioCNV, a novel approach for jointly detecting CNVs in parent-offspring trios from WGS data. Using negative binomial regression, we modeled the read depth signal while considering both GC content bias and mappability bias. Moreover, we incorporated the family relationship and used a hidden Markov model to jointly infer CNVs for three samples of a parent-offspring trio. Through application to both simulated data and a trio from 1000 Genomes Project, we showed that TrioCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TrioCNV implemented using a combination of Java and R is freely available from the website at https://github.com/yongzhuang/TrioCNV CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongzhuang Liu, Jianguo Lu, Jiajie Peng, Liran Juan, Xiaolin Zhu 0002, Bingshan Li, Yadong Wang 0001 |
Bioinform. | 8 |
| 2016 | deBWT: parallel construction of Burrows-Wheeler Transform for large collection of genomes with de Bruijn-branch encodingabstractMOTIVATION: With the development of high-throughput sequencing, the number of assembled genomes continues to rise. It is critical to well organize and index many assembled genomes to promote future genomics studies. Burrows-Wheeler Transform (BWT) is an important data structure of genome indexing, which has many fundamental applications; however, it is still non-trivial to construct BWT for large collection of genomes, especially for highly similar or repetitive genomes. Moreover, the state-of-the-art approaches cannot well support scalable parallel computing owing to their incremental nature, which is a bottleneck to use modern computers to accelerate BWT construction. RESULTS: We propose de Bruijn branch-based BWT constructor (deBWT), a novel parallel BWT construction approach. DeBWT innovatively represents and organizes the suffixes of input sequence with a novel data structure, de Bruijn branch encoding. This data structure takes the advantage of de Bruijn graph to facilitate the comparison between the suffixes with long common prefix, which breaks the bottleneck of the BWT construction of repetitive genomic sequences. Meanwhile, deBWT also uses the structure of de Bruijn graph for reducing unnecessary comparisons between suffixes. The benchmarking suggests that, deBWT is efficient and scalable to construct BWT for large dataset by parallel computing. It is well-suited to index many genomes, such as a collection of individual human genomes, with multiple-core servers or clusters. AVAILABILITY AND IMPLEMENTATION: deBWT is implemented in C language, the source code is available at https://github.com/hitbc/deBWT or https://github.com/DixianZhu/deBWTContact: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Dixian Zhu, Yadong Wang 0001 |
Bioinform. | 3 |
| 2016 | Extending gene ontology with gene association networksabstractMOTIVATION: Gene ontology (GO) is a widely used resource to describe the attributes for gene products. However, automatic GO maintenance remains to be difficult because of the complex logical reasoning and the need of biological knowledge that are not explicitly represented in the GO. The existing studies either construct whole GO based on network data or only infer the relations between existing GO terms. None is purposed to add new terms automatically to the existing GO. RESULTS: We proposed a new algorithm 'GOExtender' to efficiently identify all the connected gene pairs labeled by the same parent GO terms. GOExtender is used to predict new GO terms with biological network data, and connect them to the existing GO. Evaluation tests on biological process and cellular component categories of different GO releases showed that GOExtender can extend new GO terms automatically based on the biological network. Furthermore, we applied GOExtender to the recent release of GO and discovered new GO terms with strong support from literature. AVAILABILITY AND IMPLEMENTATION: Software and supplementary document are available at www.msu.edu/%7Ejinchen/GOExtender CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiajie Peng, Tao Wang 0082, Jixuan Wang, Yadong Wang 0001, Jin Chen 0004 |
Bioinform. | 4 |
| 2016 | A network-based pathway-expanding approach for pathway analysisabstractBACKGROUND: Pathway analysis combining multiple types of high-throughput data, such as genomics and proteomics, has become the first choice to gain insights into the pathogenesis of complex diseases. Currently, several pathway analysis methods have been developed to study complex diseases. However, these methods did not take into account the interaction between internal and external genes of the pathway and between pathways. Hence, these approaches still face some challenges. Here, we propose a network-based pathway-expanding approach that takes the topological structures of biological networks into account. RESULTS: First, two weighted gene-gene interaction networks (tumor and normal) are constructed integrating protein-protein interaction(PPI) information, gene expression data and pathway databases. Then, they are used to identify significant pathways through testing the difference of topological structures of expanded pathways in the two weighted networks. The proposed method is employed to analyze two breast cancer data. As a result, the top 15 pathways identified using the proposed method are supported by biological knowledge from the published literatures and other methods. In addition, the proposed method is also compared with other methods, such as GSEA and SPIA, and estimated using the classification performance of the top 15 expanded pathways. CONCLUSIONS: A novel network-based pathway-expanding approach is proposed to avoid the limitations of existing pathway analysis approaches. Experimental results indicate that the proposed method can accurately and reliably identify significant pathways which are related to the corresponding disease. Jie Li 0055, Haozhe Xie, Hanqing Xue, Yadong Wang 0001 |
BMC Bioinform. | 5 |
| 2015 | Family genome browser: visualizing genomes with pedigree informationabstractMOTIVATION: Families with inherited diseases are widely used in Mendelian/complex disease studies. Owing to the advances in high-throughput sequencing technologies, family genome sequencing becomes more and more prevalent. Visualizing family genomes can greatly facilitate human genetics studies and personalized medicine. However, due to the complex genetic relationships and high similarities among genomes of consanguineous family members, family genomes are difficult to be visualized in traditional genome visualization framework. How to visualize the family genome variants and their functions with integrated pedigree information remains a critical challenge. RESULTS: We developed the Family Genome Browser (FGB) to provide comprehensive analysis and visualization for family genomes. The FGB can visualize family genomes in both individual level and variant level effectively, through integrating genome data with pedigree information. Family genome analysis, including determination of parental origin of the variants, detection of de novo mutations, identification of potential recombination events and identical-by-decent segments, etc., can be performed flexibly. Diverse annotations for the family genome variants, such as dbSNP memberships, linkage disequilibriums, genes, variant effects, potential phenotypes, etc., are illustrated as well. Moreover, the FGB can automatically search de novo mutations and compound heterozygous variants for a selected individual, and guide investigators to find high-risk genes with flexible navigation options. These features enable users to investigate and understand family genomes intuitively and systematically. AVAILABILITY AND IMPLEMENTATION: The FGB is available at http://mlg.hit.edu.cn/FGB/. Liran Juan, Yongzhuang Liu, Yongtian Wang, Mingxiang Teng, Tianyi Zang, Yadong Wang 0001 |
Bioinform. | 6 |
| 2015 | Measuring semantic similarities by combining gene ontology annotations and gene co-function networksabstractBACKGROUND: Gene Ontology (GO) has been used widely to study functional relationships between genes. The current semantic similarity measures rely only on GO annotations and GO structure. This limits the power of GO-based similarity because of the limited proportion of genes that are annotated to GO in most organisms. RESULTS: We introduce a novel approach called NETSIM (network-based similarity measure) that incorporates information from gene co-function networks in addition to using the GO structure and annotations. Using metabolic reaction maps of yeast, Arabidopsis, and human, we demonstrate that NETSIM can improve the accuracy of GO term similarities. We also demonstrate that NETSIM works well even for genomes with sparser gene annotation data. We applied NETSIM on large Arabidopsis gene families such as cytochrome P450 monooxygenases to group the members functionally and show that this grouping could facilitate functional characterization of genes in these families. CONCLUSIONS: Using NETSIM as an example, we demonstrated that the performance of a semantic similarity measure could be significantly improved after incorporating genome-specific information. NETSIM incorporates both GO annotations and gene co-function network data as a priori knowledge in the model. Therefore, functional similarities of GO terms that are not explicitly encoded in GO but are relevant in a taxon-specific manner become measurable when GO annotations are limited. Supplementary information and software are available at http://www.msu.edu/~jinchen/NETSIM . Jiajie Peng, Sahra Uygun, Taehyong Kim, Yadong Wang 0001, Seung Y. Rhee, Jin Chen 0004 |
BMC Bioinform. | 4 |
| 2015 | Improving multiple sequence alignment by using better guide treesabstractProgressive sequence alignment is one of the most commonly used method for multiple sequence alignment. Roughly speaking, the method first builds a guide tree, and then aligns the sequences progressively according to the topology of the tree. It is believed that guide trees are very important to progressive alignment; a better guide tree will give an alignment with higher accuracy. Recently, we have proposed an adaptive method for constructing guide trees. This paper studies the quality of the guide trees constructed by such method. Our study showed that our adaptive method can be used to improve the accuracy of many different progressive MSA tools. In fact, we give evidences showing that the guide trees constructed by the adaptive method are among the best. Qing Zhan, Yongtao Ye, Tak Wah Lam, Siu-Ming Yiu, Yadong Wang 0001, Hing-Fung Ting |
BMC Bioinform. | 5 |
| 2015 | misFinder: identify mis-assemblies in an unbiased manner using reference and paired-end readsabstractBACKGROUND: Because of the short read length of high throughput sequencing data, assembly errors are introduced in genome assembly, which may have adverse impact to the downstream data analysis. Several tools have been developed to eliminate these errors by either 1) comparing the assembled sequences with some similar reference genome, or 2) analyzing paired-end reads aligned to the assembled sequences and determining inconsistent features alone mis-assembled sequences. However, the former approach cannot distinguish real structural variations between the target genome and the reference genome while the latter approach could have many false positive detections (correctly assembled sequence being considered as mis-assembled sequence). RESULTS: We present misFinder, a tool that aims to identify the assembly errors with high accuracy in an unbiased way and correct these errors at their mis-assembled positions to improve the assembly accuracy for downstream analysis. It combines the information of reference (or close related reference) genome and aligned paired-end reads to the assembled sequence. Assembly errors and correct assemblies corresponding to structural variations can be detected by comparing the genome reference and assembled sequence. Different types of assembly errors can then be distinguished from the mis-assembled sequence by analyzing the aligned paired-end reads using multiple features derived from coverage and consistence of insert distance to obtain high confident error calls. CONCLUSIONS: We tested the performance of misFinder on both simulated and real paired-end reads data, and misFinder gave accurate error calls with only very few miscalls. And, we further compared misFinder with QUAST and REAPR. misFinder outperformed QUAST and REAPR by 1) identified more true positive mis-assemblies with very few false positives and false negatives, and 2) distinguished the correct assemblies corresponding to structural variations from mis-assembled sequence. misFinder can be freely downloaded from https://github.com/hitbio/misFinder. Henry C. M. Leung, Francis Y. L. Chin, Siu-Ming Yiu, Guangri Quan, Qinghua Jiang, Bo Liu 0023, Yucui Dong, Yadong Wang 0001 |
BMC Bioinform. | 13 |
| 2015 | Using Semantic Association to Extend and Infer Literature-Oriented Relativity Between TermsabstractRelative terms often appear together in the literature. Methods have been presented for weighting relativity of pairwise terms by their co-occurring literature and inferring new relationship. Terms in the literature are also in the directed acyclic graph of ontologies, such as Gene Ontology and Disease Ontology. Therefore, semantic association between terms may help for establishing relativities between terms in literature. However, current methods do not use these associations. In this paper, an adjusted R-scaled score (ARSS) based on information content (ARSSIC) method is introduced to infer new relationship between terms. First, set inclusion relationship between terms of ontology was exploited to extend relationships between these terms and literature. Next, the ARSS method was presented to measure relativity between terms across ontologies according to these extensional relationships. Then, the ARSSIC method using ratios of information shared of term's ancestors was designed to infer new relationship between terms across ontologies. The result of the experiment shows that ARSS identified more pairs of statistically significant terms based on corresponding gene sets than other methods. And the high average area under the receiver operating characteristic curve (0.9293) shows that ARSSIC achieved a high true positive rate and a low false positive rate. Data is available at http://mlg.hit.edu.cn/ARSSIC/. Liang Cheng 0006, Jie Li 0055, Yang Hu 0008, Yongzhuang Liu, Yan-Shuo Chu, Yadong Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2015 | GLProbs: Aligning Multiple Sequences AdaptivelyabstractThis paper introduces a simple and effective approach to improve the accuracy of multiple sequence alignment. We use a natural measure to estimate the similarity of the input sequences, and based on this measure, we align the input sequences differently. For example, for inputs with high similarity, we consider the whole sequences and align them globally, while for those with moderately low similarity, we may ignore the flank regions and align them locally. To test the effectiveness of this approach, we have implemented a multiple sequence alignment tool called GLProbs and compared its performance with about one dozen leading alignment tools on three benchmark alignment databases, and GLProbs's alignments have the best scores in almost all testings. We have also evaluated the practicability of the alignments of GLProbs by applying the tool to three biological applications, namely phylogenetic trees construction, protein secondary structure prediction and the detection of high risk members for cervical cancer in the HPV-E6 family, and the results are very encouraging. Yongtao Ye, David Wai-Lok Cheung, Yadong Wang 0001, Siu-Ming Yiu, Qing Zhan, Tak Wah Lam, Hing-Fung Ting |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | A gradient-boosting approach for filtering de novo mutations in parent-offspring triosabstractMOTIVATION: Whole-genome and -exome sequencing on parent-offspring trios is a powerful approach to identifying disease-associated genes by detecting de novo mutations in patients. Accurate detection of de novo mutations from sequencing data is a critical step in trio-based genetic studies. Existing bioinformatic approaches usually yield high error rates due to sequencing artifacts and alignment issues, which may either miss true de novo mutations or call too many false ones, making downstream validation and analysis difficult. In particular, current approaches have much worse specificity than sensitivity, and developing effective filters to discriminate genuine from spurious de novo mutations remains an unsolved challenge. RESULTS: In this article, we curated 59 sequence features in whole genome and exome alignment context which are considered to be relevant to discriminating true de novo mutations from artifacts, and then employed a machine-learning approach to classify candidates as true or false de novo mutations. Specifically, we built a classifier, named De Novo Mutation Filter (DNMFilter), using gradient boosting as the classification algorithm. We built the training set using experimentally validated true and false de novo mutations as well as collected false de novo mutations from an in-house large-scale exome-sequencing project. We evaluated DNMFilter's theoretical performance and investigated relative importance of different sequence features on the classification accuracy. Finally, we applied DNMFilter on our in-house whole exome trios and one CEU trio from the 1000 Genomes Project and found that DNMFilter could be coupled with commonly used de novo mutation detection approaches as an effective filtering approach to significantly reduce false discovery rate without sacrificing sensitivity. AVAILABILITY: The software DNMFilter implemented using a combination of Java and R is freely available from the website at http://humangenome.duke.edu/software. Yongzhuang Liu, Bingshan Li, Renjie Tan, Xiaolin Zhu 0002, Yadong Wang 0001 |
Bioinform. | 5 |
| 2014 | Towards integrative gene functional similarity measurementabstractBACKGROUND: In Gene Ontology, the "Molecular Function" (MF) categorization is a widely used knowledge framework for gene function comparison and prediction. Its structure and annotation provide a convenient way to compare gene functional similarities at the molecular level. The existing gene similarity measures, however, solely rely on one or few aspects of MF without utilizing all the rich information available including structure, annotation, common terms, lowest common parents. RESULTS: We introduce a rank-based gene semantic similarity measure called InteGO by synergistically integrating the state-of-the-art gene-to-gene similarity measures. By integrating three GO based seed measures, InteGO significantly improves the performance by about two-fold in all the three species studied (yeast, Arabidopsis and human). CONCLUSIONS: InteGO is a systematic and novel method to study gene functional associations. The software and description are available at http://www.msu.edu/~jinchen/InteGO. Jiajie Peng, Yadong Wang 0001, Jin Chen 0004 |
BMC Bioinform. | 2 |
| 2014 | Transcriptional regulation prediction of antiestrogen resistance in breast cancer based on RNA polymerase II binding dataabstractBACKGROUND: Although endocrine therapy impedes estrogen-ER signaling pathway and thus reduces breast cancer mortality, patients remain at continued risk of relapse after tamoxifen or other endocrine therapies. Understanding the mechanisms of endocrine resistance, particularly the role of transcriptional regulation is very important and necessary. METHODS: We propose a two-step workflow based on linear model to investigate the significant differences between MCF7 and OHT cells stimulated by 17β-estradiol (E2) respect to regulatory transcription factors (TFs) and their interactions. We additionally compared predicted regulatory TFs based on RNA polymerase II (PolII) binding quantity data and gene expression data, which were taken from MCF7/MCF7+E2 and OHT/OHT+E2 cell lines following the same analysis workflow. Enrichment analysis concerning diseases and cell functions and regulatory pattern analysis of different motifs of the same TF also were performed. RESULTS: The results showed PolII data could provide more information and predict more recognizably important regulatory TFs. Large differences in TF regulatory mode were found between two cell lines. Through verified through GO annotation, enrichment analysis and related literature regarding these TFs, we found some regulatory TFs such as AP-1, C/EBP, FoxA1, GATA1, Oct-1 and NF-κB, maintained OHT cells through molecular interactions or signaling pathways that were different from the surviving MCF7 cells. From TF regulatory interaction network, we identified E2F, E2F-1 and AP-2 as hub-TFs in MCF7 cells; whereas, in addition to E2F and E2F-1, we identified C/EBP and Oct-1 as hub-TFs in OHT cells. Notably, we found the regulatory patterns of different motifs of the same TF were very different from one another sometimes. CONCLUSIONS: We inferred some regulatory TFs, such as AP-1 and NF-κB, cooperated with ER through both genomic action and non-genomic action. The TFs that were involved in both protein-protein interactions and signaling pathways could be one of the key resistant mechanisms of endocrine therapy and thus also could be new treatment targets for endocrine resistance. Our flexible workflow could be integrated into an existing analytical framework and guide biologists to further determine underlying mechanisms in human diseases. Denan Zhang, Guohua Wang 0001, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2013 | Identifying cross-category relations in gene ontology and constructing genome-specific term association networksabstractBACKGROUND: Gene Ontology (GO) has been widely used in biological databases, annotation projects, and computational analyses. Although the three GO categories are structured as independent ontologies, the biological relationships across the categories are not negligible for biological reasoning and knowledge integration. However, the existing cross-category ontology term similarity measures are either developed by utilizing the GO data only or based on manually curated term name similarities, ignoring the fact that GO is evolving quickly and the gene annotations are far from complete. RESULTS: In this paper we introduce a new cross-category similarity measurement called CroGO by incorporating genome-specific gene co-function network data. The performance study showed that our measurement outperforms the existing algorithms. We also generated genome-specific term association networks for yeast and human. An enrichment based test showed our networks are better than those generated by the other measures. CONCLUSIONS: The genome-specific term association networks constructed using CroGO provided a platform to enable a more consistent use of GO. In the networks, the frequently occurred MF-centered hub indicates that a molecular function may be shared by different genes in multiple biological processes, or a set of genes with the same functions may participate in distinct biological processes. And common subgraphs in multiple organisms also revealed conserved GO term relationships. Software and data are available online at http://www.msu.edu/~jinchen/CroGO. Jiajie Peng, Jin Chen 0004, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2012 | PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variantsabstractMOTIVATION: The development of high-throughput sequencing technologies has enabled novel methods for detecting structural variants (SVs). Current methods are typically based on depth of coverage or pair-end mapping clusters. However, most of these only report an approximate location for each SV, rather than exact breakpoints. RESULTS: We have developed pair-read informed split mapping (PRISM), a method that identifies SVs and their precise breakpoints from whole-genome resequencing data. PRISM uses a split-alignment approach informed by the mapping of paired-end reads, hence enabling breakpoint identification of multiple SV types, including arbitrary-sized inversions, deletions and tandem duplications. Comparisons to previous datasets and simulation experiments illustrate PRISM's high sensitivity, while PCR validations of PRISM results, including previously uncharacterized variants, indicate an overall precision of ~90%. AVAILABILITY: PRISM is freely available at http://compbio.cs.toronto.edu/prism. Yadong Wang 0001, Michael Brudno |
Bioinform. | 2 |
| 2012 | regSNPs: a strategy for prioritizing regulatory single nucleotide substitutionsabstractMOTIVATION: One of the fundamental questions in genetics study is to identify functional DNA variants that are responsible to a disease or phenotype of interest. Results from large-scale genetics studies, such as genome-wide association studies (GWAS), and the availability of high-throughput sequencing technologies provide opportunities in identifying causal variants. Despite the technical advances, informatics methodologies need to be developed to prioritize thousands of variants for potential causative effects. RESULTS: We present regSNPs, an informatics strategy that integrates several established bioinformatics tools, for prioritizing regulatory SNPs, i.e. the SNPs in the promoter regions that potentially affect phenotype through changing transcription of downstream genes. Comparing to existing tools, regSNPs has two distinct features. It considers degenerative features of binding motifs by calculating the differences on the binding affinity caused by the candidate variants and integrates potential phenotypic effects of various transcription factors. When tested by using the disease-causing variants documented in the Human Gene Mutation Database, regSNPs showed mixed performance on various diseases. regSNPs predicted three SNPs that can potentially affect bone density in a region detected in an earlier linkage study. Potential effects of one of the variants were validated using luciferase reporter assay. Mingxiang Teng, Shoji Ichikawa, Leah R. Padgett, Yadong Wang 0001, Matthew E. Mort, David N. Cooper, Daniel L. Koller, Tatiana Foroud, Howard J. Edenberg, Michael J. Econs |
Bioinform. | 4 |
| 2010 | Predicting human microRNA-disease associations based on support vector machineabstractThe identification of disease-related microRNAs is vital for understanding the pathogenesis of disease at the molecular level and may lead to the design of specific molecular tools for diagnosis, treatment and prevention. Experimental identification of disease-related microRNAs poses difficulties. Computational prediction of microRNA-disease associations is one of the complementary means. However, one major issue in microRNA studies is the lack of bioinformatics programs to accurately predict microRNA-disease associations. Herein, we present a machine learning-based approach for distinguishing positive microRNA-disease associations from negative microRNA-disease associations. A set of features was extracted for each positive and negative microRNA-disease association, and a support vector machine (SVM) classifier was trained, which achieved the area under the ROC curve of up to 0.8884 in 10-fold cross-validation procedure, indicating that the SVM-based approach described here can be used to predict potential microRNA-disease associations and formulate testable hypotheses to guide future biological experiments. Qinghua Jiang, Guohua Wang 0001, Yadong Wang 0001 |
BIBM | 4 |