EDBT 2026 Demo / reviewers in the wild / expert
Bo Liu 0023
dblp:58/2670-23
· DBLP profile ↗
35ranked-venue papers
5as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 5 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RVC: A Real-Time Variant Calling Framework for Short-Read Sequencing DataabstractAccurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics. Miao Cui 0005, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023, Guohua Wang 0001, Yadong Liu 0001 |
BIBM | 5 |
| 2024 | miniSNV: accurate and fast single nucleotide variant calling from nanopore sequencing dataabstractNanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV. Miao Cui 0005, Yadong Liu 0001, Hongzhe Guo, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023 |
Briefings Bioinform. | 7 |
| 2024 | Kled: an ultra-fast and sensitive structural variant detection tool for long-read sequencing dataabstractStructural Variants (SVs) are a crucial type of genetic variant that can significantly impact phenotypes. Therefore, the identification of SVs is an essential part of modern genomic analysis. In this article, we present kled, an ultra-fast and sensitive SV caller for long-read sequencing data given the specially designed approach with a novel signature-merging algorithm, custom refinement strategies and a high-performance program structure. The evaluation results demonstrate that kled can achieve optimal SV calling compared to several state-of-the-art methods on simulated and real long-read data for different platforms and sequencing depths. Furthermore, kled excels at rapid SV calling and can efficiently utilize multiple Central Processing Unit (CPU) cores while maintaining low memory usage. The source code for kled can be obtained from https://github.com/CoREse/kled. Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Bo Liu 0023, Yadong Wang 0001 |
Briefings Bioinform. | 6 |
| 2023 | HNetGO: protein function prediction via heterogeneous network transformerabstractProtein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches. Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004 |
Briefings Bioinform. | 7 |
| 2023 | Struct2GO: protein function prediction based on graph pooling algorithm and AlphaFold2 structure informationabstractMOTIVATION: In recent years, there has been a breakthrough in protein structure prediction, and the AlphaFold2 model of the DeepMind team has improved the accuracy of protein structure prediction to the atomic level. Currently, deep learning-based protein function prediction models usually extract features from protein sequences and combine them with protein-protein interaction networks to achieve good results. However, for newly sequenced proteins that are not in the protein-protein interaction network, such models cannot make effective predictions. To address this, this article proposes the Struct2GO model, which combines protein structure and sequence data to enhance the precision of protein function prediction and the generality of the model. RESULTS: We obtain amino acid residue embeddings in protein structure through graph representation learning, utilize the graph pooling algorithm based on a self-attention mechanism to obtain the whole graph structure features, and fuse them with sequence features obtained from the protein language model. The results demonstrate that compared with the traditional protein sequence-based function prediction model, the Struct2GO model achieves better results. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available at https://github.com/lyjps/Struct2GO. Peishun Jiao, Xuan Wang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
Bioinform. | 4 |
| 2023 | ScCCL: Single-Cell Data Clustering Based on Self-Supervised Contrastive LearningabstractThe growing maturity of single-cell RNA-sequencing (scRNA-seq) technology allows us to explore the heterogeneity of tissues, organisms, and complex diseases at cellular level. In single-cell data analysis, clustering calculation is very important. However, the high dimensionality of scRNA-seq data, the ever-increasing number of cells, and the unavoidable technical noise bring great challenges to clustering calculations. Motivated by the good performance of contrastive learning in multiple domains, we propose ScCCL, a novel self-supervised contrastive learning method for clustering of scRNA-seq data. ScCCL first randomly masks the gene expression of each cell twice and adds a small amount of Gaussian noise, and then uses the momentum encoder structure to extract features from the enhanced data. Contrastive learning is then applied in the instance-level contrastive learning module and the cluster-level contrastive learning module, respectively. After training, a representation model that can efficiently extract high-order embeddings of single cells is obtained. We selected two evaluation metrics, ARI and NMI, to conduct experiments on multiple public datasets. The results show that ScCCL improves the clustering effect compared with the benchmark algorithms. Notably, since ScCCL does not depend on a specific type of data, it can also be helpful in clustering analysis of single-cell multi-omics data. Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | PSPGO: Cross-Species Heterogeneous Network Propagation for Protein Function PredictionabstractHow to use computational methods to effectively predict the function of proteins remains a challenge. Most prediction methods based on single species or single data source have some limitations: the former need to train different models for different species, the latter only to infer protein function from a single perspective, such as the method only using Protein-Protein Interaction (PPI) network just considers the protein environment but ignore the intrinsic characteristics of protein sequences. We found that in some network-based multi-species methods the networks of each species are isolated, which means there is no communication between networks of different species. To solve these problems, we propose a cross-species heterogeneous network propagation method based on graph attention mechanism, PSPGO, which can propagate feature and label information on sequence similarity (SS) network and PPI network for predicting gene ontology terms. Our model is evaluated on a large multi-species dataset split based on time and is compared with several state-of-the-art methods. The results show that our method has good performance. We also explore the predictive performance of PSPGO for a single species. The results illustrate that PSPGO also performs well in prediction for single species. Kaitao Wu, Lexiang Wang, Bo Liu 0023, Yang Liu 0039, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | Prot2GO: Predicting GO Annotations From Protein Sequences and InteractionsabstractProtein is the main material basis of living organisms and plays crucial role in life activities. Understanding the function of protein is of great significance for new drug discovery, disease treatment and vaccine development. In recent years, with the widespread application of deep learning in bioinformatics, researchers have proposed many deep learning models to predict protein functions. However, the existing deep learning methods usually only consider protein sequences, and thus cannot effectively integrate multi-source data to annotate protein functions. In this article, we propose the Prot2GO model, which can integrate protein sequence and PPI network data to predict protein functions. We utilize an improved biased random walk algorithm to extract the features of PPI network. For sequence data, we use a convolutional neural network to obtain the local features of the sequence and a recurrent neural network to capture the long-range associations between amino acid residues in protein sequence. Moreover, Prot2GO adopts the attention mechanism to identify protein motifs and structural domains. Experiments show that Prot2GO model achieves the state-of-the-art performance on multiple metrics. Xiaoshuai Zhang, Hucheng Liu, Xiaofeng Zhang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Deepgmd: A Graph-Neural-Network-Based Method to Detect Gene Regulator ModuleabstractRegulatory module mining methods divide genes into multiple gene subgroups and explore potential biological mechanisms from omics data. By transforming gene expression profile data into gene co-expression network, we transform the task of gene module detection into the problem of finding community structure in the graph, and introduce the latest network representation learning method-graph neural network to optimize this problem. In order to systematically evaluate whether the algorithm allows overlap to affect such problems, we make two variants of the output of the algorithm, Deepgmd_cluster and Deepgmd. The difference between them is whether overlap is allowed. By comparing the known modules and the modules generated by the algorithm, we can evaluate the quality of the algorithm. We use this method to compare our algorithm with some current mainstream methods. The results show that our method has greater advantages. In the end, we analyze some typical modules from the modules found by the algorithm for visualization, and use the GO database and KEGG database to perform enrichment analysis and pathway analysis on these modules. Yulin Wu 0001, Jiangsheng Pi, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | ScSSC: Semi-supervised Single Cell Clustering Based on 2D Embedding
Naile Shi, Yulin Wu 0001, Linlin Du, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
ICIC (3) | 4 |
| 2021 | abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractSUMMARY: Partial order alignment, which aligns a sequence to a directed acyclic graph, is now frequently used as a key component in long-read error correction and assembly. We present abPOA (adaptive banded Partial Order Alignment), a Single Instruction Multiple Data (SIMD)-based C library for fast partial order alignment using adaptive banded dynamic programming. It can work as a stand-alone multiple sequence alignment and consensus calling tool or be easily integrated into any long-read error correction and assembly workflow. Compared to a state-of-the-art tool (SPOA), abPOA is up to 10 times faster with a comparable alignment accuracy. AVAILABILITY AND IMPLEMENTATION: abPOA is implemented in C. A stand-alone tool and a C/Python software interface are freely available at https://github.com/yangao07/abPOA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 4 |
| 2021 | Erratum to: abPOA: an SIMD-based C library for fast partial order alignment using adaptive bandabstractBioinformatics, 2021, 37(15), 2209–2211. doi: https://doi.org/10.1093/bioinformatics/btaa963 Upon the original publication of this article, the affiliation: “Center for Bioinformatics, Department of Computer Science and Technology, Harbin Institute of Technology, Harbin, Heilongjiang 150001, China” was inadvertently omitted prior to publication. The affiliation has been reinstated. Yan Gao 0031, Yongzhuang Liu, Yanmei Ma, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 4 |
| 2021 | SKSV: ultrafast structural variation detection from circular consensus sequencing readsabstractSUMMARY: Circular consensus sequencing reads are promising for the comprehensive detection of structural variants (SVs). However, alignment-based SV calling pipelines are computationally intensive due to the generation of complete read-alignments and its post-processing. Herein, we propose a SKeleton-based analysis toolkit for Structural Variation detection (SKSV). Benchmarks on real and simulated datasets demonstrate that SKSV has an order of magnitude of faster speed than state-of-the-art SV calling approaches; moreover, it achieves higher F1 scores for various types of SVs. AVAILABILITY AND IMPLEMENTATION: SKSV is available from https://github.com/ydLiu-HIT/SKSV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yadong Liu 0001, Tao Jiang 0021, Junhao Su, Bo Liu 0023, Tianyi Zang, Yadong Wang 0001 |
Bioinform. | 4 |
| 2021 | Factor graph-aggregated heterogeneous network embedding for disease-gene association predictionabstractBACKGROUND: Exploring the relationship between disease and gene is of great significance for understanding the pathogenesis of disease and developing corresponding therapeutic measures. The prediction of disease-gene association by computational methods accelerates the process. RESULTS: Many existing methods cannot fully utilize the multi-dimensional biological entity relationship to predict disease-gene association due to multi-source heterogeneous data. This paper proposes FactorHNE, a factor graph-aggregated heterogeneous network embedding method for disease-gene association prediction, which captures a variety of semantic relationships between the heterogeneous nodes by factorization. It produces different semantic factor graphs and effectively aggregates a variety of semantic relationships, by using end-to-end multi-perspectives loss function to optimize model. Then it produces good nodes embedding to prediction disease-gene association. CONCLUSIONS: Experimental verification and analysis show FactorHNE has better performance and scalability than the existing models. It also has good interpretability and can be extended to large-scale biomedical network data analysis. Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004 |
BMC Bioinform. | 3 |
| 2021 | IIMLP: integrated information-entropy-based method for LncRNA predictionabstractBACKGROUND: The prediction of long non-coding RNA (lncRNA) has attracted great attention from researchers, as more and more evidence indicate that various complex human diseases are closely related to lncRNAs. In the era of bio-med big data, in addition to the prediction of lncRNAs by biological experimental methods, many computational methods based on machine learning have been proposed to make better use of the sequence resources of lncRNAs. RESULTS: We developed the lncRNA prediction method by integrating information-entropy-based features and machine learning algorithms. We calculate generalized topological entropy and generate 6 novel features for lncRNA sequences. By employing these 6 features and other features such as open reading frame, we apply supporting vector machine, XGBoost and random forest algorithms to distinguish human lncRNAs. We compare our method with the one which has more K-mer features and results show that our method has higher area under the curve up to 99.7905%. CONCLUSIONS: We develop an accurate and efficient method which has novel information entropy features to analyze and classify lncRNAs. Our method is also extendable for research on the other functional elements in DNA sequences. Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Qing Liao 0001, Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 10 |
| 2021 | Fast and SNP-aware short read alignment with SALTabstractBACKGROUND: DNA sequence alignment is a common first step in most applications of high-throughput sequencing technologies. The accuracy of sequence alignments directly affects the accuracy of downstream analyses, such as variant calling and quantitative analysis of transcriptome; therefore, rapidly and accurately mapping reads to a reference genome is a significant topic in bioinformatics. Conventional DNA read aligners map reads to a linear reference genome (such as the GRCh38 primary assembly). However, such a linear reference genome represents the genome of only one or a few individuals and thus lacks information on variations in the population. This limitation can introduce bias and impact the sensitivity and accuracy of mapping. Recently, a number of aligners have begun to map reads to populations of genomes, which can be represented by a reference genome and a large number of genetic variants. However, compared to linear reference aligners, an aligner that can store and index all genetic variants has a high cost in memory (RAM) space and leads to extremely long run time. Aligning reads to a graph-model-based index that includes all types of variants is ultimately an NP-hard problem in theory. By contrast, considering only single nucleotide polymorphism (SNP) information will reduce the complexity of the index and improve the speed of sequence alignment. RESULTS: The SNP-aware alignment tool (SALT) is a fast, memory-efficient, and SNP-aware short read alignment tool. SALT uses 5.8 GB of RAM to index a human reference genome (GRCh38) and incorporates 12.8M UCSC common SNPs. Compared with a state-of-the-art aligner, SALT has a similar speed but higher accuracy. CONCLUSIONS: Herein, we present an SNP-aware alignment tool (SALT) that aligns reads to a reference genome that incorporates an SNP database. We benchmarked SALT using simulated and real datasets. The results demonstrate that SALT can efficiently map reads to the reference genome with significantly improved accuracy. Incorporating SNP information can improve the accuracy of read alignment and can reveal novel variants. The source code is freely available at https://github.com/weiquan/SALT . Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 2 |
| 2021 | deGSM: Memory Scalable Construction Of Large Scale de Bruijn GraphabstractThe de Bruijn graph, a fundamental data structure to represent and organize genome sequence, plays important roles in various kinds of sequence analysis tasks. With the rapid development of HTS data and ever-increasing number of assembled genomes, there is a high demand to construct the very large de Bruijn graph for sequences up to Tera-base-pair level. Current approaches may have unaffordable memory footprints to handle such a large de Bruijn graph. We propose a lightweight parallel de Bruijn graph construction approach: de Bruijn Graph Constructor in Scalable Memory (deGSM). The main idea of deGSM is to efficiently construct the Burrows-Wheeler Transformation (BWT) of the unipaths of the de Bruijn graph in constant RAM space and transform the BWT into the original unitigs. The experimental results demonstrate that, just with a commonly available machine, deGSM is able to handle very large genome sequence(s), e.g., the contigs (305 Gbp) and scaffolds (1.1 Tbp) recorded in GenBank database and Picea abies HTS dataset (9.7 Tbp). Moreover, deGSM also has faster or comparable construction speed compared with state-of-the-art approaches. With its high scalability and efficiency, deGSM has enormous potential in many large scale genomics studies. The deGSM is publicly available at: https://github.com/hitbc/deGSM. Hongzhe Guo, Yilei Fu, Yan Gao 0031, Junyi Li 0004, Yadong Wang 0001, Bo Liu 0023 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2020 | Assessment of Machine Learning Methods for Classification in Single Cell ATAC-seqabstractSingle-cell assay for transposase accessible chromatin using sequencing(scATAC-seq) is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. The similarity of data structure and feature between scRNA-seq and scATAC-seq makes it feasible to identify the cell types in scATAC-seq through traditional supervised machine learning methods. Here, we evaluated 6 popular machine learning methods for classification in scATAC-seq. The performance of the methods is evaluated using 4 public single cell ATAC-seq datasets of different tissues, sizes and technologies. We evaluated these methods using intradatasets experiments of 5-folds cross validation based on accuracy, recall and percentage of correctly predicted cells. We found that these methods may perform well in some types of cells in a single dataset, but the overall results are not as well as in scRNA-seq analysis. For testing the classification ability of machine learning methods across datasets, we applied inter-dataset experiments to test the performance of machine learning methods in realistic scenarios. SVM and NMC are overall the top 2 best-performing methods across all experiments. We recommend researchers to apply SVM and NMC as the underlying classifier when developing an automatic classification method in scATAC-seq. Bo Liu 0023, Liran Juan, Tianyi Zang, Tao Jiang 0021, Yadong Wang 0001 |
BIBM | 2 |
| 2020 | GONET: A Deep Network to Annotate Proteins via Recurrent Convolution NetworksabstractFinding out the functions of protein in life activities precisely is nontrivial, which is the core of current proteomics research. Gene Ontology standardizes the function of protein into a series of GO terms, each of which belongs to exactly one of the three subontologies: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF). The prediction of protein function can be considered as a multi-label classification problem. Traditional methods often spend a lot of costs to extract handcrafted features and plenty of domain knowledge is needed when solving these tasks, while using deep learning technology can overcome these shortcomings. Here, we propose a deep model GONET based on recurrent convolutional neural networks, which annotates protein in an end-to-end manner. Our model combines protein sequences and protein-protein interaction (PPI) network data, and utilizes representation learning to learn distributed representation of proteins to overcome the sparse nature and semantic independence problem. Moreover, we adopt a quite deep CNNRNN-Attention model, which is able to effectively extract high-order features of protein sequences. We have carried out experiments on several datasets, which achieve the state-of-the-art in some metrics compared with the existing competitive methods. Junyi Li 0004, Xiaoshuai Zhang, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 4 |
| 2019 | SALT: a fast, memory-efficient and SNP-aware short read alignment toolabstractDNA sequence alignment tools play an essential role in genomics and genetics. The accuracy of the alignment directly affects the accuracy of downstream analysis, such as variant calling, so it is essential to map reads to the reference genome rapidly and accurately. It has become an essential topic in the field of bioinformatics. Conventional read aligners map reads to a linear reference genome (such as GRCh38 primary). However, the linear reference genome only represents one or a few individuals of genomes, which lacks the variation information in population. It can introduce bias and impact sensitivity and accuracy of mapping. Recently, a few aligners are beginning to map reads to a graph that captures the entire human genome along with a large number of variants. However, compared to linear reference aligners, storing and indexing all genetic variants require costly memory(RAM) space and make extremely long runtime. Aligning reads to a graph model-based index, including the whole set of variants, is ultimately an NP-hard problem in theory. Considering only SNPs information will reduce the complexity of index and improve the speed of alignments. Herein, we present an SNP-aware alignment tool (SALT) that aligns reads to a reference genome that incorporates the SNP database. The SALT is benchmarked both on simulated reads and the real dataset. The results demonstrate that SALT can efficiently map reads to the reference genome, and significantly improve accuracy and sensitivity. Read alignment incorporating SNPs information can improve the sensitivity and accuracy of the read alignment. Moreover, it helps to discover novel variants. SALT is distributed under the GNU General Public License (GPL). Source code is freely available at https://github.com/weiquan/SALT. Bo Liu 0023, Yadong Wang 0001 |
BIBM | 2 |
| 2019 | A Bidirectional Fuzzy Index and Approximate Search Algorithm for Next Generation SequencingabstractSequence alignment is one of the most important problems in bioinformatics. However, the existing alignment tools may result in a large number of candidate locations which degradate alignment performance. Recent researches discard the high repetitive seeds for improving the alignment speed, which influences alignment accuracy. To this end, we propose a novel fuzzy index ( fBWT), which is available at https://github.com/weiquan/appr_bwt. It allows approximate search and extending the length of seeds to reduce the candidate locations and accelerate the sequence alignment. The performance of our tool was compared with BWA using 150bp and 250bp length datasets. The result shows the number of misaligned reads (the correct position is not included in the high-score candidate position set) of the current mainstream tool is 5-10 times higher than it. The efficiency between the presented and the existing tools are also compared. Under the above conditions, the alignment time of the presented tool is very close. However, the alignment speed of fBWT is much faster than BWA under the requirement of similar alignment accuracy. Guangri Quan, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 3 |
| 2019 | Prediction of Human LncRNAs Based on Integrated Information Entropy Features
Junyi Li 0004, Huinian Li, Qingzhe Xu, Xiaozhu Jing, Wei Jiang 0023, Bo Liu 0023, Yadong Wang 0001 |
ICIC (2) | 8 |
| 2019 | TideHunter: efficient and sensitive tandem repeat detection from noisy long-reads using seed-and-chainabstractMOTIVATION: Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) sequencing technologies can produce long-reads up to tens of kilobases, but with high error rates. In order to reduce sequencing error, Rolling Circle Amplification (RCA) has been used to improve library preparation by amplifying circularized template molecules. Linear products of the RCA contain multiple tandem copies of the template molecule. By integrating additional in silico processing steps, these tandem sequences can be collapsed into a consensus sequence with a higher accuracy than the original raw reads. Existing pipelines using alignment-based methods to discover the tandem repeat patterns from the long-reads are either inefficient or lack sensitivity. RESULTS: We present a novel tandem repeat detection and consensus calling tool, TideHunter, to efficiently discover tandem repeat patterns and generate high-quality consensus sequences from amplified tandemly repeated long-read sequencing data. TideHunter works with noisy long-reads (PacBio and ONT) at error rates of up to 20% and does not have any limitation of the maximal repeat pattern size. We benchmarked TideHunter using simulated and real datasets with varying error rates and repeat pattern sizes. TideHunter is tens of times faster than state-of-the-art methods and has a higher sensitivity and accuracy. AVAILABILITY AND IMPLEMENTATION: TideHunter is written in C, it is open source and is available at https://github.com/yangao07/TideHunter. Yan Gao 0031, Bo Liu 0023, Yadong Wang 0001, Yi Xing |
Bioinform. | 2 |
| 2019 | rMETL: sensitive mobile element insertion detection with long read realignmentabstractSUMMARY: Mobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing technologies provides the opportunity to detect MEIs sensitively. However, the signals of MEI implied by noisy long reads are highly complex due to the repetitiveness of mobile elements as well as the high sequencing error rates. Herein, we propose the Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). Benchmarking results of simulated and real datasets demonstrate that rMETL enables to handle the complex signals to discover MEIs sensitively. It is suited to produce high-quality MEI callsets in many genomics studies. AVAILABILITY AND IMPLEMENTATION: rMETL is available from https://github.com/hitbc/rMETL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Jiang 0021, Bo Liu 0023, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 2 |
| 2019 | Integrated entropy-based approach for analyzing exons and introns in DNA sequencesabstractBACKGROUND: Numerous essential algorithms and methods, including entropy-based quantitative methods, have been developed to analyze complex DNA sequences since the last decade. Exons and introns are the most notable components of DNA and their identification and prediction are always the focus of state-of-the-art research. RESULTS: In this study, we designed an integrated entropy-based analysis approach, which involves modified topological entropy calculation, genomic signal processing (GSP) method and singular value decomposition (SVD), to investigate exons and introns in DNA sequences. We optimized and implemented the topological entropy and the generalized topological entropy to calculate the complexity of DNA sequences, highlighting the characteristics of repetition sequences. By comparing digitalizing entropy values of exons and introns, we observed that they are significantly different. After we converted DNA data to numerical topological entropy value, we applied SVD method to effectively investigate exon and intron regions on a single gene sequence. Additionally, several genes across five species are used for exon predictions. CONCLUSIONS: Our approach not only helps to explore the complexity of DNA sequence and its functional elements, but also provides an entropy-based GSP method to analyze exon and intron regions. Our work is feasible across different species and extendable to analyze other components in both coding and noncoding region of DNA sequences. Junyi Li 0004, Huinian Li, Qingzhe Xu, Renjie Tan, Bo Liu 0023, Yadong Wang 0001 |
BMC Bioinform. | 9 |
| 2018 | rCANID: read Clustering and Assembly-based Novel Insertion Detection tool
Yilei Fu, Tao Jiang 0021, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 3 |
| 2018 | deSPI: efficient classification of metagenomics reads with lightweight de Bruijn graph-based reference indexing
Dengfeng Guan, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 2 |
| 2018 | Fast variation-aware read alignment with deBGA-VARA
Hongzhe Guo, Bo Liu 0023, Dengfeng Guan, Yilei Fu, Yadong Wang 0001 |
BIBM | 2 |
| 2017 | LAMSA: fast split read alignment with long approximate matchesabstractMOTIVATION: Read length is continuously increasing with the development of novel high-throughput sequencing technologies, which has enormous potentials on cutting-edge genomic studies. However, longer reads could more frequently span the breakpoints of structural variants (SVs) than that of shorter reads. This may greatly influence read alignment, since most state-of-the-art aligners are designed for handling relatively small variants in a co-linear alignment framework. Meanwhile, long read alignment is still not as efficient as that of short reads, which could be also a bottleneck for the upcoming wide application. RESULTS: We propose long approximate matches-based split aligner (LAMSA), a novel split read alignment approach. It takes the advantage of the rareness of SVs to implement a specifically designed two-step strategy. That is, LAMSA initially splits the read into relatively long fragments and co-linearly align them to solve the small variations or sequencing errors, and mitigate the effect of repeats. The alignments of the fragments are then used for implementing a sparse dynamic programming-based split alignment approach to handle the large or non-co-linear variants. We benchmarked LAMSA with simulated and real datasets having various read lengths and sequencing error rates, the results demonstrate that it is substantially faster than the state-of-the-art long read aligners; meanwhile, it also has good ability to handle various categories of SVs. AVAILABILITY AND IMPLEMENTATION: LAMSA is available at https://github.com/hitbc/LAMSA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Yan Gao 0031, Yadong Wang 0001 |
Bioinform. | 1 |
| 2017 | rMFilter: acceleration of long read-based structure variation calling by chimeric read filteringabstractMOTIVATION: Long read sequencing technologies provide new opportunities to investigate genome structural variations (SVs) more accurately. However, the state-of-the-art SV calling pipelines are computational intensive and the applications of long reads are restricted. RESULTS: We propose a local region match-based filter (rMFilter) to efficiently nail down chimeric noisy long reads based on short token matches within local genomic regions. rMFilter is able to substantially accelerate long read-based SV calling pipelines without loss of effectiveness. It can be easily integrated into current long read-based pipelines to facilitate SV studies. AVAILABILITY AND IMPLEMENTATION: The C ++ source code of rMFilter is available at https://github.com/hitbc/rMFilter . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Tao Jiang 0021, Siu-Ming Yiu, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 1 |
| 2016 | Accurate annotation of metagenomic data without species-level referencesabstractTaxonomic annotation is a critical first step for analysis of metagenomic data. Despite a lot of tools being developed, the accuracy is still not satisfactory, in particular, when a close species-level reference does not exist in the database. In this paper, we propose a novel annotation tool, MetaAnnotator, to annotate metagenomic reads, which outperforms all existing tools significantly when only genus-level references exist in the database. From our experiments, MetaAnnotator can assign 87.5% reads correctly (67.5% reads are assigned to the exact genus) with only 8.5% reads wrongly assigned. The best existing tool (MetaCluster-TA) can only achieve 73.4% correct read assignment (with only 50.9% reads assigned to the exact genus and 22.6% reads wrongly assigned). The speed of MetaAnnotator is also the second faster (1 hour for 20 million reads). The core concepts behind MetaAnnotator includes: (i) we only consider exact k-mers in coding regions of the references as they should be more significant and accurate; (ii) to assign reads to taxonomy nodes, we construct genome and taxonomy specific probabilistic models from the reference database; and (iii) using the BWT data structure to speed up the k-mer matching process. Haobin Yao, Tak Wah Lam, Hing-Fung Ting, Siu-Ming Yiu, Yadong Wang 0001, Bo Liu 0023 |
BIBM | 6 |
| 2016 | deBGA: read alignment with de Bruijn graph-based seed and extensionabstractMOTIVATION: As high-throughput sequencing (HTS) technology becomes ubiquitous and the volume of data continues to rise, HTS read alignment is becoming increasingly rate-limiting, which keeps pressing the development of novel read alignment approaches. Moreover, promising novel applications of HTS technology require aligning reads to multiple genomes instead of a single reference; however, it is still not viable for the state-of-the-art aligners to align large numbers of reads to multiple genomes. RESULTS: We propose de Bruijn Graph-based Aligner (deBGA), an innovative graph-based seed-and-extension algorithm to align HTS reads to a reference genome that is organized and indexed using a de Bruijn graph. With its well-handling of repeats, deBGA is substantially faster than state-of-the-art approaches while maintaining similar or higher sensitivity and accuracy. This makes it particularly well-suited to handle the rapidly growing volumes of sequencing data. Furthermore, it provides a promising solution for aligning reads to multiple genomes and graph-based references in HTS applications. AVAILABILITY AND IMPLEMENTATION: deBGA is available at: https://github.com/hitbc/deBGA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Hongzhe Guo, Michael Brudno, Yadong Wang 0001 |
Bioinform. | 1 |
| 2016 | rHAT: fast alignment of noisy long reads with regional hashingabstractMOTIVATION: Single Molecule Real-Time (SMRT) sequencing has been widely applied in cutting-edge genomic studies. However, it is still an expensive task to align the noisy long SMRT reads to reference genome by state-of-the-art aligners, which is becoming a bottleneck in applications with SMRT sequencing. Novel approach is on demand for improving the efficiency and effectiveness of SMRT read alignment. RESULTS: We propose Regional Hashing-based Alignment Tool (rHAT), a seed-and-extension-based read alignment approach specifically designed for noisy long reads. rHAT indexes reference genome by regional hash table (RHT), a hash table-based index which describes the short tokens within local windows of reference genome. In the seeding phase, rHAT utilizes RHT for efficiently calculating the occurrences of short token matches between partial read and local genomic windows to find highly possible candidate sites. In the extension phase, a sparse dynamic programming-based heuristic approach is used for reducing the cost of aligning read to the candidate sites. By benchmarking on the real and simulated datasets from various prokaryote and eukaryote genomes, we demonstrated that rHAT can effectively align SMRT reads with outstanding throughput. AVAILABILITY AND IMPLEMENTATION: rHAT is implemented in C++; the source code is available at https://github.com/HIT-Bioinformatics/rHAT CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Dengfeng Guan, Mingxiang Teng, Yadong Wang 0001 |
Bioinform. | 1 |
| 2016 | deBWT: parallel construction of Burrows-Wheeler Transform for large collection of genomes with de Bruijn-branch encodingabstractMOTIVATION: With the development of high-throughput sequencing, the number of assembled genomes continues to rise. It is critical to well organize and index many assembled genomes to promote future genomics studies. Burrows-Wheeler Transform (BWT) is an important data structure of genome indexing, which has many fundamental applications; however, it is still non-trivial to construct BWT for large collection of genomes, especially for highly similar or repetitive genomes. Moreover, the state-of-the-art approaches cannot well support scalable parallel computing owing to their incremental nature, which is a bottleneck to use modern computers to accelerate BWT construction. RESULTS: We propose de Bruijn branch-based BWT constructor (deBWT), a novel parallel BWT construction approach. DeBWT innovatively represents and organizes the suffixes of input sequence with a novel data structure, de Bruijn branch encoding. This data structure takes the advantage of de Bruijn graph to facilitate the comparison between the suffixes with long common prefix, which breaks the bottleneck of the BWT construction of repetitive genomic sequences. Meanwhile, deBWT also uses the structure of de Bruijn graph for reducing unnecessary comparisons between suffixes. The benchmarking suggests that, deBWT is efficient and scalable to construct BWT for large dataset by parallel computing. It is well-suited to index many genomes, such as a collection of individual human genomes, with multiple-core servers or clusters. AVAILABILITY AND IMPLEMENTATION: deBWT is implemented in C language, the source code is available at https://github.com/hitbc/deBWT or https://github.com/DixianZhu/deBWTContact: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Dixian Zhu, Yadong Wang 0001 |
Bioinform. | 1 |
| 2015 | misFinder: identify mis-assemblies in an unbiased manner using reference and paired-end readsabstractBACKGROUND: Because of the short read length of high throughput sequencing data, assembly errors are introduced in genome assembly, which may have adverse impact to the downstream data analysis. Several tools have been developed to eliminate these errors by either 1) comparing the assembled sequences with some similar reference genome, or 2) analyzing paired-end reads aligned to the assembled sequences and determining inconsistent features alone mis-assembled sequences. However, the former approach cannot distinguish real structural variations between the target genome and the reference genome while the latter approach could have many false positive detections (correctly assembled sequence being considered as mis-assembled sequence). RESULTS: We present misFinder, a tool that aims to identify the assembly errors with high accuracy in an unbiased way and correct these errors at their mis-assembled positions to improve the assembly accuracy for downstream analysis. It combines the information of reference (or close related reference) genome and aligned paired-end reads to the assembled sequence. Assembly errors and correct assemblies corresponding to structural variations can be detected by comparing the genome reference and assembled sequence. Different types of assembly errors can then be distinguished from the mis-assembled sequence by analyzing the aligned paired-end reads using multiple features derived from coverage and consistence of insert distance to obtain high confident error calls. CONCLUSIONS: We tested the performance of misFinder on both simulated and real paired-end reads data, and misFinder gave accurate error calls with only very few miscalls. And, we further compared misFinder with QUAST and REAPR. misFinder outperformed QUAST and REAPR by 1) identified more true positive mis-assemblies with very few false positives and false negatives, and 2) distinguished the correct assemblies corresponding to structural variations from mis-assembled sequence. misFinder can be freely downloaded from https://github.com/hitbio/misFinder. Henry C. M. Leung, Francis Y. L. Chin, Siu-Ming Yiu, Guangri Quan, Qinghua Jiang, Bo Liu 0023, Yucui Dong, Yadong Wang 0001 |
BMC Bioinform. | 10 |