VLDB 2026 Research / reviewers in the wild / expert
Tao Jiang 0021
dblp:181/2813-21
· DBLP profile ↗
24ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-0673-8503ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 20 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFSeeker: a splicing-graph-based approach for accurate gene fusion detection from long-read RNA sequencing dataabstractGene fusions are critical oncogenic drivers and therapeutic targets in diverse cancers. Long-read ribonucleic acid sequencing (RNA-seq) offers an unprecedented opportunity to resolve the full-length structure of fusion isoforms, but its high intrinsic error rates pose significant challenges to the precise identification of true fusion events. Here, we developed GFSeeker, an innovative splicing-graph-based computational framework for accurate gene fusion detection from long-read RNA-seq. GFSeeker employs a unique pipeline based on a splicing graph reference and a dual re-alignment validation to effectively overcome data noise from high error rates. Benchmarking across simulated, non-tumor, and cancer cell line datasets demonstrated GFSeeker's state-of-the-art performance, achieving 6%-15% higher F1 score compared to existing methods. Notably, GFSeeker successfully identified the known fusion event, MATN2-POP1, in the MCF-7 cancer cell line, missed by other tools, highlighting its superior sensitivity in resolving complex fusion events. These results validate GFSeeker as a powerful and reliable tool for gene fusion discovery, heralding its significant potential to advance cancer research and precision diagnostics. Heng Hu, Runtian Gao, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 5 |
| 2026 | cuteSV-OL: a real-time structural variation detection framework for nanopore sequencing devicesabstractSUMMARY: Nanopore sequencing technology enables real-time sequencing and is widely used in rapid detection applications. However, in clinical scenarios, existing structural variant (SV) detection tools typically separate sequencing from computation, limiting their timeliness for clinical applications. To address this, we introduce cuteSV-OL, a novel framework designed for real-time SV discovery, which can be embedded within nanopore sequencing instruments to analyze data concurrently with its generation. Additionally, cuteSV-OL features a real-time SV detection rate evaluation module, allowing users to terminate sequencing early when appropriate, thereby reducing time and cost. Experimental results show that on a standard desktop computer, cuteSV-OL can perform real-time analysis during sequencing and complete SV calling within min after sequencing ends, achieving performance comparable to offline methods. This approach has the potential to enhance rapid clinical diagnostics. AVAILABILITY AND IMPLEMENTATION: cuteSV-OL is released under the MIT license and is available at https://github.com/gwmHIT/cuteSV-OL. It can also be installed via Bioconda or accessed through https://doi.org/10.5281/zenodo.17777436. Weimin Guo, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
Bioinform. | 4 |
| 2025 | RVC: A Real-Time Variant Calling Framework for Short-Read Sequencing DataabstractAccurate detection of single-nucleotide variants (SNVs) and small insertions/deletions (indels) from second-generation sequencing (NGS) data is essential for clinical applications such as cancer diagnostics, infectious disease monitoring, and rapid genetic screening. However, conventional variant calling pipelines, such as GATK, decouple analysis from sequencing, deferring detection until sequencing is fully completed. We introduce RVC, a real-time variant calling framework tailored for cycle-based NGS workflows. RVC incrementally processes partially sequenced reads and continuously updates variant evidence using a scanline-based alignment algorithm and a lightweight binomial scoring model. This design enables progressive, low-latency SNVs and indels detection during sequencing, without disrupting the sequencing pipeline. In benchmark experiments using the HG002 dataset, RVC completed variant calling within tens of minutes after sequencing, significantly outperforming GATK in runtime. By tightly integrating analysis with sequencing output, RVC bridges the gap between sequencing speed and clinical responsiveness, offering a scalable and practical solution for real-time genomic diagnostics. Miao Cui 0005, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023, Guohua Wang 0001, Yadong Liu 0001 |
BIBM | 3 |
| 2025 | MethSV: A Long-Read-Based Framework for Profiling DNA Methylation in Structural VariationabstractDNA methylation plays a crucial role in regulating diverse molecular processes in living organisms. However, methylation patterns associated with structural variations (SVs) remain poorly understood due to technical limitations. with the advent of long-read sequencing technologies, it is now feasible to simultaneously detect SVs and methylation at single-molecule resolution. Here, we propose methSV, a novel framework to profile DNA methylation within SV regions, revealing potential interactions between genetic and epigenetic regulation. Using methSV, we achieved a$\sim 3.7$-fold increase in the detection of SV-associated population-level differentially methylated regions (pDMRs) and \~{}4.2-fold increase in high-confidence haplotype-resolved differentially methylated regions (hDMRs). Notably, distinct methylation signatures in high-type deletions (DELs) and low-type insertions (INSs) suggested regulatory potential and were enriched in disease-associated genes such as GP1BA, CHRNE, and HCN2. This framework extends the analytical capabilities of long-read epigenomic studies and offers new insights into the functional consequences of SV. Weize Kong, Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 5 |
| 2025 | SimPG: A Pangenome-Guided Population-Specific Human Genome Simulation ToolabstractThe rapid advancement of high-throughput genome sequencing has enabled large-scale reconstruction of genome sequences at both individual and population levels. However, accurately evaluating genome assemblies and associated variants remains a significant challenge due to sequencing errors, assembly artifacts, and limited experimental validation, which hinder the development of novel algorithms, benchmarking of genomic pipelines, and validation of biological hypotheses. Existing linear genome simulation tools, typically based on a single reference genome, offer limited biological realism and fail to capture the genomic diversity and structural complexity needed for comprehensive evaluation. To address these limitations, we introduce SimPG, a novel simulation framework that generates individual genomes with populationlevel characteristics by leveraging the rich variant and structural information embedded in pangenomes. SimPG produces realistic, high-quality simulated genomes that support diverse applications such as structural variant detection and population genetics research. By providing a reproducible, controllable, and standardized environment, SimPG facilitates the development, testing, and benchmarking of genome analysis tools, enabling robust performance assessment and error analysis, and ultimately advancing computational genomics and genome biology. Yadong Liu 0001, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 5 |
| 2025 | DNAMatch: An Ultra-Fast and Memory-Efficient Deep Learning Framework for Aligning Ultra-Long DNA FragmentsabstractEfficient and accurate alignment of DNA fragments is fundamental to genomics research. With the advent of advanced sequencing technologies, the length of sequencing reads and assembled contigs has increased significantly, posing substantial challenges for existing alignment algorithms. These methods often struggle with megabase-scale DNA fragments due to the computational burden of global searches and exhaustive chromosomal queries. To address this, we propose DNAMatch, a novel alignment framework that integrates the DNABERT2 pre-trained model for feature extraction with a deep residual network for chromosome identification. DNAMatch introduces a chromosome pre-localization strategy, which effectively narrows the search space and significantly reduces the memory footprint required by the downstream aligner, minimap2. This design enables fast and precise alignment of ultra-long DNA fragments. Benchmarking on simulated ultra-long reads from multiple model organisms-including human, Drosophila melanogaster, and Arabidopsis thaliana-demonstrates that DNAMatch achieves 9 8-9 9% accuracy in chromosome identification, accelerates the alignment process by 52.7%, and reduces memory usage by 73%. Importantly, when applied to downstream structural variation detection, DNAMatch maintains high accuracy, with only a marginal 0.05% decrease in F1-score compared to the standard minimap2 pipeline. These results highlight DNAMatch as a powerful and efficient tool for aligning ultra-long reads and analyzing complex genomes. Yuansong Zhu, Chuanmin Wu, Yadong Wang 0001, Tao Jiang 0021, Yadong Liu 0001 |
BIBM | 4 |
| 2025 | SVHunter: long-read-based structural variation detection through the transformer modelabstractStructural variations (SVs) are genomic rearrangements larger than 50 bp, that are widely present in the human genome and are associated with various complex diseases. Existing long-read-based SV detection tools often rely on fixed rules or heuristic algorithms, which can oversimplify the complexity of SV signatures. Therefore, these methods usually lack flexibility and cannot fully capture SV signals, leading to reduced accuracy and robustness. To address these issues, we propose SVHunter, a transformer-based method for long-read SV detection. SVHunter combines convolutional neural networks and transformers to capture both local and global SV signatures, enabling accurate identification of SVs. Additionally, SVHunter employs the mean shift clustering algorithm, which dynamically adjusts bandwidth parameters to accommodate different types of SVs without requiring a preset number of clusters, thus allowing precise breakpoint clustering. Validation across multiple sequencing platforms and datasets demonstrates that SVHunter excels at detecting various types of SVs, with a notable reduction in the false discovery rate. This highlights considerable strong potential for both research and clinical applications. Runtian Gao, Heng Hu, Zhongjun Jiang, Shuqi Cao, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 7 |
| 2025 | CKG-TPI: integrating collaborative knowledge graph with sequence interactions for TCR-peptide binding specificityabstractAccurately identifying interactions between T-cell receptors (TCRs) and peptides is a fundamental challenge in immunology, with significant implications for vaccine design and immunotherapy. While computational methods offer efficient alternatives to labor-intensive experimental screening, achieving robust and accurate TCR-peptide binding prediction remains a challenging task. To address this, we propose collaborative knowledge graph (CKG-TPI), a novel prediction framework based on graph neural networks that integrates both interaction patterns between TCR and peptide sequences and their higher-order biological context through a constructed collaborative knowledge graph. Experimental results on multiple publicly available independent datasets demonstrate that CKG-TPI consistently outperforms state-of-the-art models. Specifically, it achieves a 9.89% improvement in area under the ROC curve compared to the strongest baseline model UnifyImmun, and a 23.93% increase in area under the precision-recall curve over the leading baseline method. Moreover, attention weight visualization and peptide-specific TCR screening validate the model's effectiveness, underscoring its potential as a powerful tool for immunological research and therapeutic discovery. Yue Liu 0034, Haoyan Wang, Guohua Wang 0001, Yadong Liu 0001, Tao Jiang 0021, Yadong Wang 0001 |
Briefings Bioinform. | 5 |
| 2024 | Comprehensive Benchmarking of Genotype Imputation Tools Using a Large-Scale Chinese Reference PanelabstractGenome-wide association studies (GWAS) remain as one of the most essential and potent strategies for identifying genetic markers associated with common human diseases. Gene arrays, which are widely used in GWAS, target predesigned known genomic mutations to enhance clinical diagnostics. The approach of genotype imputation which leverages reference panels has shown great promise in extrapolating additional meaningful variants from existing data. Therefore, accurately broadening the spectrum of mutations beyond those included in gene arrays can unlock significant potential for cost-effective and advanced GWAS analysis. In this study, we focus on the four most popular imputation tools currently to assess their performances based on a large-scale Chinese reference panel. Benchmarking results demonstrate that Minimac4 is the best imputation method, which achieved higher accuracy with more well-imputed loci and maintained lower computational consumption. Moreover, we further revealed the minimum sample size required in the reference panel and the mutations involved in arrays that ensure the acquisition of acceptable imputed call sets. We believe this study can provide practical significance guidelines for the research community to select the suitable imputation tool via an existing reference panel, or construct their own reference panel and genotype array in the most effective way. Yadong Liu 0001, Zhongbo Yang, Yadong Wang 0001, Tao Jiang 0021 |
BIBM | 5 |
| 2024 | Comprehensive evaluation of haplotype phasing tools with different strategies across diverse sequence technologiesabstractHaplotype phasing is a computational technique used to determine the allelic arrangement in a chromosome from genotype data. This process is crucial for understanding genetic diversity and inheritance patterns, which are essential for fields like genetic research, personalized medicine, and population genetic studies. In this study, we conducted a comprehensive evaluation of five state-of-the-art haplotype phasing tools under different strategies including statistical-based Beagle5, Eagle2 and Shapeit5, and read-based WhatsHap, and HapCut2, across various sequencing platforms including Illumina, PacBio, and ONT. Our findings indicate that statistical-based tools generally provide better phasing continuity and are less dependent on the sequencing technology used, whereas read-based tools offer superior performance with long-read datasets but struggle with short reads due to insufficient sequence length. In addition, statistical-based tools were found to be less resource-intensive compared to their read-based counterparts. The insights gained from this benchmarking provide a valuable guide for researchers in selecting the most appropriate phasing tools, potentially enhancing the accuracy and efficiency of genetic analysis and supporting advancements in genetic research methodologies. Zhongbo Yang, Zhenhao Lu, Tao Jiang 0021, Yadong Wang 0001, Yadong Liu 0001 |
BIBM | 4 |
| 2024 | miniSNV: accurate and fast single nucleotide variant calling from nanopore sequencing dataabstractNanopore sequence technology has demonstrated a longer read length and enabled to potentially address the limitations of short-read sequencing including long-range haplotype phasing and accurate variant calling. However, there is still room for improvement in terms of the performance of single nucleotide variant (SNV) identification and computing resource usage for the state-of-the-art approaches. In this work, we introduce miniSNV, a lightweight SNV calling algorithm that simultaneously achieves high performance and yield. miniSNV utilizes known common variants in populations as variation backgrounds and leverages read pileup, read-based phasing, and consensus generation to identify and genotype SNVs for Oxford Nanopore Technologies (ONT) long reads. Benchmarks on real and simulated ONT data under various error profiles demonstrate that miniSNV has superior sensitivity and comparable accuracy on SNV detection and runs faster with outstanding scalability and lower memory than most state-of-the-art variant callers. miniSNV is available from https://github.com/CuiMiao-HIT/miniSNV. Miao Cui 0005, Yadong Liu 0001, Hongzhe Guo, Tao Jiang 0021, Yadong Wang 0001, Bo Liu 0023 |
Briefings Bioinform. | 5 |
| 2024 | SVDF: enhancing structural variation detect from long-read sequencing via automatic filtering strategiesabstractStructural variation (SV) is an important form of genomic variation that influences gene function and expression by altering the structure of the genome. Although long-read data have been proven to better characterize SVs, SVs detected from noisy long-read data still include a considerable portion of false-positive calls. To accurately detect SVs in long-read data, we present SVDF, a method that employs a learning-based noise filtering strategy and an SV signature-adaptive clustering algorithm, for effectively reducing the likelihood of false-positive events. Benchmarking results from multiple orthogonal experiments demonstrate that, across different sequencing platforms and depths, SVDF achieves higher calling accuracy for each sample compared to several existing general SV calling tools. We believe that, with its meticulous and sensitive SV detection capability, SVDF can bring new opportunities and advancements to cutting-edge genomic research. Heng Hu, Runtian Gao, Zhongjun Jiang, Murong Zhou, Guohua Wang 0001, Tao Jiang 0021 |
Briefings Bioinform. | 8 |
| 2024 | Kled: an ultra-fast and sensitive structural variant detection tool for long-read sequencing dataabstractStructural Variants (SVs) are a crucial type of genetic variant that can significantly impact phenotypes. Therefore, the identification of SVs is an essential part of modern genomic analysis. In this article, we present kled, an ultra-fast and sensitive SV caller for long-read sequencing data given the specially designed approach with a novel signature-merging algorithm, custom refinement strategies and a high-performance program structure. The evaluation results demonstrate that kled can achieve optimal SV calling compared to several state-of-the-art methods on simulated and real long-read data for different platforms and sequencing depths. Furthermore, kled excels at rapid SV calling and can efficiently utilize multiple Central Processing Unit (CPU) cores while maintaining low memory usage. The source code for kled can be obtained from https://github.com/CoREse/kled. Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Bo Liu 0023, Yadong Wang 0001 |
Briefings Bioinform. | 2 |
| 2024 | MEHunter: transformer-based mobile element variant detection from long readsabstractSUMMARY: Mobile genetic elements (MEs) are heritable mutagens that significantly contribute to genetic diseases. The advent of long-read sequencing technologies, capable of resolving large DNA fragments, offers promising prospects for the comprehensive detection of ME variants (MEVs). However, achieving high precision while maintaining recall performance remains challenging mainly brought by the variable length and similar content of MEV signatures, which are often obscured by the noise in long reads. Here, we propose MEHunter, a high-performance MEV detection approach utilizing a fine-tuned transformer model adept at identifying potential MEVs with fragmented features. Benchmark experiments on both simulated and real datasets demonstrate that MEHunter consistently achieves higher accuracy and sensitivity than the state-of-the-art tools. Furthermore, it is capable of detecting novel potentially individual-specific MEVs that have been overlooked in published population projects. AVAILABILITY AND IMPLEMENTATION: MEHunter is available from https://github.com/120L021101/MEHunter. Tao Jiang 0021, Zuji Zhou, Shuqi Cao, Yadong Wang 0001, Yadong Liu 0001 |
Bioinform. | 1 |
| 2023 | Comprehensive evaluation of RNA-seq alignment methods based on long-read sequencing dataabstractLong-read RNA sequencing (RNA-seq) has revolutionized our ability to comprehensively study transcriptomes, enabling the detection of full-length transcripts. The accurate alignment of long-read RNA-seq is a critical step in downstream analysis. However, with the development of alignment tools and each tool declaring the ability to handle long-read RNA-seq alignment on their own, researchers face the challenge of distinguishing and selecting the most suitable tool for their tasks. But to the best of our knowledge, there is still a lack of comprehensive evaluations on the aligners designed for long-read sequencing data. Here, we conducted a benchmark on five state-of-the-art tools on simulated long-read RNA-seq data from three levels to comprehensively assess the performance of each aligner, and provide valuable insights into the strengths and limitations of each aligner, aiding researchers in selecting the most appropriate tool for their long-read RNA-seq studies. We expected our study to advance the field of transcriptomics and promote the accurate analysis of long-read RNA-seq data in diverse biological contexts. Yadong Liu 0001, Hongzhe Guo, Zhenhao Lu, Yadong Wang 0001, Zhongyu Liu, Tao Jiang 0021 |
BIBM | 6 |
| 2023 | Reply: Correspondence on NanoVar's performance outlined by Jiang T. et al. in 'Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation'abstractWe published a paper in BMC Bioinformatics comprehensively evaluating the performance of structural variation (SV) calling with long-read SV detection methods based on simulated error-prone long-read data under various sequencing settings. Recently, C.Y.T. et al. wrote a correspondence claiming that the performance of NanoVar was underestimated in our benchmarking and listed some errors in our previous manuscripts. To clarify these matters, we reproduced our previous benchmarking results and carried out a series of parallel experiments on both the newly generated simulated datasets and the ones provided by C.Y.T. et al. The robust benchmark results indicate that NanoVar has unstable performance on simulated data produced from different versions of VISOR, while other tools do not exhibit this phenomenon. Furthermore, the errors proposed by C.Y.T. et al. were due to them using another version of VISOR and Sniffles, which caused many changes in usage and results compared to the versions applied in our previous work. We hope that this commentary proves the validity of our previous publication, clarifies and eliminates the misunderstanding about the commands and results in our benchmarking. Furthermore, we welcome more experts and scholars in the scientific community to pay attention to our research and help us better optimize these valuable works. Tao Jiang 0021, Hongzhe Guo |
BMC Bioinform. | 1 |
| 2022 | Comparison of the Nanopore and PacBio sequencing technologies for DNA 5-methylcytosine detectionabstractDNA methylation provides a pivotal layer of epigenetic regulation in eukaryotes that has significant involvement for numerous biological processes in health and disease. Recent long-read sequencing technology including Oxford Nanopore sequencing and PacBio HiFi sequencing greatly expands the capacity of long-range, single-molecule, and direct DNA modification detection from reads without extra laboratory techniques. A growing number of analytical pipelines including base-calling and 5mC methylation detection have been developed, but there is still a lack of comprehensive evaluations of the two sequencing technologies. Here, we assess the performance of different methylation-calling pipelines based on Nanopore and HiFi sequencing datasets to provide a systematic evaluation to guide researchers on how to select the long-read sequencing technologies in performing human epigenome-wide studies. Yadong Liu 0001, Zhongyu Liu, Tao Jiang 0021, Tianyi Zang, Yadong Wang 0001 |
BIBM | 3 |
| 2022 | Evaluation of classification in single cell atac-seq data with machine learning methodsabstractBACKGROUND: The technologies advances of single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) allowed to generate thousands of single cells in a relatively easy and economic manner and it is rapidly advancing the understanding of the cellular composition of complex organisms and tissues. The data structure and feature in scRNA-seq is similar to that in scATAC-seq, therefore, it's encouraged to identify and classify the cell types in scATAC-seq through traditional supervised machine learning methods, which are proved reliable in scRNA-seq datasets. RESULTS: In this study, we evaluated the classification performance of 6 well-known machine learning methods on scATAC-seq. A total of 4 public scATAC-seq datasets vary in tissues, sizes and technologies were applied to the evaluation of the performance of the methods. We assessed these methods using a 5-folds cross validation experiment, called intra-dataset experiment, based on recall, precision and the percentage of correctly predicted cells. The results show that these methods performed well in some specific types of the cell in a specific scATAC-seq dataset, while the overall performance is not as well as that in scRNA-seq analysis. In addition, we evaluated the classification performance of these methods by training and predicting in different datasets generated from same sample, called inter-datasets experiments, which may help us to assess the performance of these methods in more realistic scenarios. CONCLUSIONS: Both in intra-dataset and in inter-dataset experiment, SVM and NMC are overall outperformed others across all 4 datasets. Thus, we recommend researchers to use SVM and NMC as the underlying classifier when developing an automatic cell-type classification method for scATAC-seq. Hongzhe Guo, Zhongbo Yang, Tao Jiang 0021, Yadong Wang 0001 |
BMC Bioinform. | 3 |
| 2021 | SKSV: ultrafast structural variation detection from circular consensus sequencing readsabstractSUMMARY: Circular consensus sequencing reads are promising for the comprehensive detection of structural variants (SVs). However, alignment-based SV calling pipelines are computationally intensive due to the generation of complete read-alignments and its post-processing. Herein, we propose a SKeleton-based analysis toolkit for Structural Variation detection (SKSV). Benchmarks on real and simulated datasets demonstrate that SKSV has an order of magnitude of faster speed than state-of-the-art SV calling approaches; moreover, it achieves higher F1 scores for various types of SVs. AVAILABILITY AND IMPLEMENTATION: SKSV is available from https://github.com/ydLiu-HIT/SKSV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yadong Liu 0001, Tao Jiang 0021, Junhao Su, Bo Liu 0023, Tianyi Zang, Yadong Wang 0001 |
Bioinform. | 2 |
| 2021 | Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluationabstractBACKGROUND: With the rapid development of long-read sequencing technologies, it is possible to reveal the full spectrum of genetic structural variation (SV). However, the expensive cost, finite read length and high sequencing error for long-read data greatly limit the widespread adoption of SV calling. Therefore, it is urgent to establish guidance concerning sequencing coverage, read length, and error rate to maintain high SV yields and to achieve the lowest cost simultaneously. RESULTS: In this study, we generated a full range of simulated error-prone long-read datasets containing various sequencing settings and comprehensively evaluated the performance of SV calling with state-of-the-art long-read SV detection methods. The benchmark results demonstrate that almost all SV callers perform better when the long-read data reach 20× coverage, 20 kbp average read length, and approximately 10-7.5% or below 1% error rates. Furthermore, high sequencing coverage is the most influential factor in promoting SV calling, while it also directly determines the expensive costs. CONCLUSIONS: Based on the comprehensive evaluation results, we provide important guidelines for selecting long-read sequencing settings for efficient SV calling. We believe these recommended settings of long-read sequencing will have extraordinary guiding significance in cutting-edge genomic studies and clinical practices. Tao Jiang 0021, Shuqi Cao, Yadong Liu 0001, Yadong Wang 0001, Hongzhe Guo |
BMC Bioinform. | 1 |
| 2020 | Assessment of Machine Learning Methods for Classification in Single Cell ATAC-seqabstractSingle-cell assay for transposase accessible chromatin using sequencing(scATAC-seq) is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. The similarity of data structure and feature between scRNA-seq and scATAC-seq makes it feasible to identify the cell types in scATAC-seq through traditional supervised machine learning methods. Here, we evaluated 6 popular machine learning methods for classification in scATAC-seq. The performance of the methods is evaluated using 4 public single cell ATAC-seq datasets of different tissues, sizes and technologies. We evaluated these methods using intradatasets experiments of 5-folds cross validation based on accuracy, recall and percentage of correctly predicted cells. We found that these methods may perform well in some types of cells in a single dataset, but the overall results are not as well as in scRNA-seq analysis. For testing the classification ability of machine learning methods across datasets, we applied inter-dataset experiments to test the performance of machine learning methods in realistic scenarios. SVM and NMC are overall the top 2 best-performing methods across all experiments. We recommend researchers to apply SVM and NMC as the underlying classifier when developing an automatic classification method in scATAC-seq. Bo Liu 0023, Liran Juan, Tianyi Zang, Tao Jiang 0021, Yadong Wang 0001 |
BIBM | 5 |
| 2019 | rMETL: sensitive mobile element insertion detection with long read realignmentabstractSUMMARY: Mobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing technologies provides the opportunity to detect MEIs sensitively. However, the signals of MEI implied by noisy long reads are highly complex due to the repetitiveness of mobile elements as well as the high sequencing error rates. Herein, we propose the Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). Benchmarking results of simulated and real datasets demonstrate that rMETL enables to handle the complex signals to discover MEIs sensitively. It is suited to produce high-quality MEI callsets in many genomics studies. AVAILABILITY AND IMPLEMENTATION: rMETL is available from https://github.com/hitbc/rMETL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Jiang 0021, Bo Liu 0023, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 1 |
| 2018 | rCANID: read Clustering and Assembly-based Novel Insertion Detection tool
Yilei Fu, Tao Jiang 0021, Bo Liu 0023, Yadong Wang 0001 |
BIBM | 2 |
| 2017 | rMFilter: acceleration of long read-based structure variation calling by chimeric read filteringabstractMOTIVATION: Long read sequencing technologies provide new opportunities to investigate genome structural variations (SVs) more accurately. However, the state-of-the-art SV calling pipelines are computational intensive and the applications of long reads are restricted. RESULTS: We propose a local region match-based filter (rMFilter) to efficiently nail down chimeric noisy long reads based on short token matches within local genomic regions. rMFilter is able to substantially accelerate long read-based SV calling pipelines without loss of effectiveness. It can be easily integrated into current long read-based pipelines to facilitate SV studies. AVAILABILITY AND IMPLEMENTATION: The C ++ source code of rMFilter is available at https://github.com/hitbc/rMFilter . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bo Liu 0023, Tao Jiang 0021, Siu-Ming Yiu, Junyi Li 0004, Yadong Wang 0001 |
Bioinform. | 2 |