Hongzhe Guo

dblp:188/8446 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2024 miniSNV: accurate and fast single nucleotide variant calling from nanopore sequencing data
abstract
Nanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV.
Miao Cui 0005, Yadong Liu 0001, Hongzhe Guo, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023
Briefings Bioinform.4
2023 Comprehensive evaluation of RNA-seq alignment methods based on long-read sequencing data
abstract
Long-read RNA sequencing (RNA-seq) has revolutionized our ability to comprehensively study transcriptomes, enabling the detection of full-length transcripts. The accurate alignment of long-read RNA-seq is a critical step in downstream analysis. However, with the development of alignment tools and each tool declaring the ability to handle long-read RNA-seq alignment on their own, researchers face the challenge of distinguishing and selecting the most suitable tool for their tasks. But to the best of our knowledge, there is still a lack of comprehensive evaluations on the aligners designed for long-read sequencing data. Here, we conducted a benchmark on five state-of-the-art tools on simulated long-read RNA-seq data from three levels to comprehensively assess the performance of each aligner, and provide valuable insights into the strengths and limitations of each aligner, aiding researchers in selecting the most appropriate tool for their long-read RNA-seq studies. We expected our study to advance the field of transcriptomics and promote the accurate analysis of long-read RNA-seq data in diverse biological contexts.
Yadong Liu 0001, Hongzhe Guo, Zhenhao Lu, Yadong Wang 0001, Zhongyu Liu, Tao Jiang 0021
BIBM2
2023 Reply: Correspondence on NanoVar's performance outlined by Jiang T. et al. in 'Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation'
abstract
We published a paper in BMC Bioinformatics comprehensively evaluating the performance of structural variation (SV) calling with long-read SV detection methods based on simulated error-prone long-read data under various sequencing settings. Recently, C.Y.T. et al. wrote a correspondence claiming that the performance of NanoVar was underestimated in our benchmarking and listed some errors in our previous manuscripts. To clarify these matters, we reproduced our previous benchmarking results and carried out a series of parallel experiments on both the newly generated simulated datasets and the ones provided by C.Y.T. et al. The robust benchmark results indicate that NanoVar has unstable performance on simulated data produced from different versions of VISOR, while other tools do not exhibit this phenomenon. Furthermore, the errors proposed by C.Y.T. et al. were due to them using another version of VISOR and Sniffles, which caused many changes in usage and results compared to the versions applied in our previous work. We hope that this commentary proves the validity of our previous publication, clarifies and eliminates the misunderstanding about the commands and results in our benchmarking. Furthermore, we welcome more experts and scholars in the scientific community to pay attention to our research and help us better optimize these valuable works.
Tao Jiang 0021, Hongzhe Guo
BMC Bioinform.3
2022 Evaluation of classification in single cell atac-seq data with machine learning methods
abstract
BACKGROUND: The technologies advances of single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) allowed to generate thousands of single cells in a relatively easy and economic manner and it is rapidly advancing the understanding of the cellular composition of complex organisms and tissues. The data structure and feature in scRNA-seq is similar to that in scATAC-seq, therefore, it's encouraged to identify and classify the cell types in scATAC-seq through traditional supervised machine learning methods, which are proved reliable in scRNA-seq datasets. RESULTS: In this study, we evaluated the classification performance of 6 well-known machine learning methods on scATAC-seq. A total of 4 public scATAC-seq datasets vary in tissues, sizes and technologies were applied to the evaluation of the performance of the methods. We assessed these methods using a 5-folds cross validation experiment, called intra-dataset experiment, based on recall, precision and the percentage of correctly predicted cells. The results show that these methods performed well in some specific types of the cell in a specific scATAC-seq dataset, while the overall performance is not as well as that in scRNA-seq analysis. In addition, we evaluated the classification performance of these methods by training and predicting in different datasets generated from same sample, called inter-datasets experiments, which may help us to assess the performance of these methods in more realistic scenarios. CONCLUSIONS: Both in intra-dataset and in inter-dataset experiment, SVM and NMC are overall outperformed others across all 4 datasets. Thus, we recommend researchers to use SVM and NMC as the underlying classifier when developing an automatic cell-type classification method for scATAC-seq.
Hongzhe Guo, Zhongbo Yang, Tao Jiang 0021, Yadong Wang 0001
BMC Bioinform.1
2021 Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation
abstract
BACKGROUND: With the rapid development of long-read sequencing technologies, it is possible to reveal the full spectrum of genetic structural variation (SV). However, the expensive cost, finite read length and high sequencing error for long-read data greatly limit the widespread adoption of SV calling. Therefore, it is urgent to establish guidance concerning sequencing coverage, read length, and error rate to maintain high SV yields and to achieve the lowest cost simultaneously. RESULTS: In this study, we generated a full range of simulated error-prone long-read datasets containing various sequencing settings and comprehensively evaluated the performance of SV calling with state-of-the-art long-read SV detection methods. The benchmark results demonstrate that almost all SV callers perform better when the long-read data reach 20× coverage, 20 kbp average read length, and approximately 10-7.5% or below 1% error rates. Furthermore, high sequencing coverage is the most influential factor in promoting SV calling, while it also directly determines the expensive costs. CONCLUSIONS: Based on the comprehensive evaluation results, we provide important guidelines for selecting long-read sequencing settings for efficient SV calling. We believe these recommended settings of long-read sequencing will have extraordinary guiding significance in cutting-edge genomic studies and clinical practices.
Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Yadong Wang 0001, Hongzhe Guo
BMC Bioinform.7
2021 deGSM: Memory Scalable Construction Of Large Scale de Bruijn Graph
abstract
The de Bruijn graph, a fundamental data structure to represent and organize genome sequence, plays important roles in various kinds of sequence analysis tasks. With the rapid development of HTS data and ever-increasing number of assembled genomes, there is a high demand to construct the very large de Bruijn graph for sequences up to Tera-base-pair level. Current approaches may have unaffordable memory footprints to handle such a large de Bruijn graph. We propose a lightweight parallel de Bruijn graph construction approach: de Bruijn Graph Constructor in Scalable Memory (deGSM). The main idea of deGSM is to efficiently construct the Burrows-Wheeler Transformation (BWT) of the unipaths of the de Bruijn graph in constant RAM space and transform the BWT into the original unitigs. The experimental results demonstrate that, just with a commonly available machine, deGSM is able to handle very large genome sequence(s), e.g., the contigs (305 Gbp) and scaffolds (1.1 Tbp) recorded in GenBank database and Picea abies HTS dataset (9.7 Tbp). Moreover, deGSM also has faster or comparable construction speed compared with state-of-the-art approaches. With its high scalability and efficiency, deGSM has enormous potential in many large scale genomics studies. The deGSM is publicly available at: https://github.com/hitbc/deGSM.
Hongzhe Guo, Yilei Fu, Yan Gao 0031, Junyi Li 0004, Yadong Wang 0001, Bo Liu 0023
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 Fast variation-aware read alignment with deBGA-VARA
Hongzhe Guo, Bo Liu 0023, Dengfeng Guan, Yilei Fu, Yadong Wang 0001
BIBM1
2016 deBGA: read alignment with de Bruijn graph-based seed and extension
abstract
MOTIVATION: As high-throughput sequencing (HTS) technology becomes ubiquitous and the volume of data continues to rise, HTS read alignment is becoming increasingly rate-limiting, which keeps pressing the development of novel read alignment approaches. Moreover, promising novel applications of HTS technology require aligning reads to multiple genomes instead of a single reference; however, it is still not viable for the state-of-the-art aligners to align large numbers of reads to multiple genomes. RESULTS: We propose de Bruijn Graph-based Aligner (deBGA), an innovative graph-based seed-and-extension algorithm to align HTS reads to a reference genome that is organized and indexed using a de Bruijn graph. With its well-handling of repeats, deBGA is substantially faster than state-of-the-art approaches while maintaining similar or higher sensitivity and accuracy. This makes it particularly well-suited to handle the rapidly growing volumes of sequencing data. Furthermore, it provides a promising solution for aligning reads to multiple genomes and graph-based references in HTS applications. AVAILABILITY AND IMPLEMENTATION: deBGA is available at: https://github.com/hitbc/deBGA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Bo Liu 0023, Hongzhe Guo, Michael Brudno, Yadong Wang 0001
Bioinform.2