Ruibang Luo

dblp:55/11133 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-9711-6533ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Accelerated long-read variant calling with Clair3 for whole-genome sequencing
abstract
SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30× whole-genome sequence in 12-20 minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15 minutes on an Apple Mac Studio (32 threads), which is ∼10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30× ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.
Zhenxian Zheng, Minggao He, Lei Chen 0072, Angel On Ki Wong, Yekai Zhou, Ruibang Luo
Bioinform.9
2025 Repun: an accurate small variant representation unification method for multiple sequencing platforms
abstract
Ensuring a unified variant representation aligning the sequencing data is critical for downstream analysis as variant representation may differ across platforms and sequencing conditions. Current approaches typically treat variant unification as a post-step following variant calling and are incapable of measuring the correct variant representation from the outset. Aligning variant representations with the alignment before variant calling has benefits like providing reliable training labels for deep learning-based variant caller model training and enabling direct assessment of alignment quality. However, it also poses challenges due to the large number of candidates to handle. Here, we present Repun, a haplotype-aware variant-alignment unification algorithm that harmonizes the variant representation between provided variants and alignments in different sequencing platforms. Repun leverages phasing to facilitate equivalent haplotype matches between variants and alignments. Our approach reduced the comparisons between variant haplotypes and candidate haplotypes by utilizing haplotypes with read evidence to speed up the unification process. Repun achieved >99.99% precision and > 99.5% recall through extensive evaluations of various Genome in a Bottle Consortium samples encompassing three sequencing platforms: Oxford Nanopore Technology, Pacific Biosciences, and Illumina. Repun is open-source and available at (https://github.com/zhengzhenxian/Repun).
Zhenxian Zheng, Yingxuan Ren, Lei Chen 0072, Angel On Ki Wong, Tak Wah Lam, Ruibang Luo
Briefings Bioinform.8
2025 AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature
abstract
MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.
Chi-Man Liu, Yuanhua Huang, Tak Wah Lam, Ruibang Luo
Bioinform.6
2024 ShiftCAM: A Time-Domain Content Addressable Memory Utilizing Shifted Hamming Distance for Robust Genome Analysis
abstract
Fast and efficient genome analysis can have a significant impact in areas such as scientific discovery and personalized medicine. Given the extensive data produced by sequencing machines, in-memory computing is considered a strong candidate to tackle the frequent data movement issue. Previous research has introduced many designs based on Content Addressable Memories (CAM), mainly optimized for tolerating edit distance; however, these systems struggle when there are a few insertion or deletion errors. This limitation presents a significant challenge for genome analysis, as current Third-Generation Sequencing still has high error rates. In this work, we introduce ShiftCAM, a time-domain Content Addressable Memory, designed to accommodate the high error rates in practical scenarios. Utilizing time-domain comparison, ShiftCAM effectively calculates the Shifted Hamming Distance to better approximate the computationally expensive edit distance. Additionally, the Modification to Accidental Match strategy specially designed for hardware implementation is introduced to eliminate accidental matches of single base pairs, further reducing false positives and improving edit distance approximation. Monte Carlo simulations based on physical ReRAM device statistical measurements and commercial PDK are also conducted to validate the robustness of the ShiftCAM design. Our experiments demonstrate that ShiftCAM can achieve an average of 2.1× (from 40.1% to 83.8%) higher F1 score in contamination analysis, 21.3% estimation error in relative abundance analysis, 51.2% reduction in cell area, 29.5× speed up, and 9.4× higher energy efficiency, compared to state-of-the-art in-memory DNA classification accelerators.
Peiyi He, Ruibin Mao, Keyi Shan, Yunwei Tong, Muyuan Peng, Ruibang Luo, Can Li 0024
ICCAD7
2024 Unveiling promising drug targets for autism spectrum disorder: insights from genetics, transcriptomics, and proteomics
abstract
Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder for which current treatments are limited and drug development costs are prohibitive. Identifying drug targets for ASD is crucial for the development of targeted therapies. Summary-level data of expression quantitative trait loci obtained from GTEx, protein quantitative trait loci data from the ROSMAP project, and two ASD genome-wide association studies datasets were utilized for discovery and replication. We conducted a combined analysis using Mendelian randomization (MR), transcriptome-wide association studies, Bayesian colocalization, and summary-data-based MR to identify potential therapeutic targets associated with ASD and examine whether there are shared causal variants among them. Furthermore, pathway and drug enrichment analyses were performed to further explore the underlying mechanisms and summarize the current status of pharmacological targets for developing drugs to treat ASD. The protein-protein interaction (PPI) network and mouse knockout models were performed to estimate the effect of therapeutic targets. A total of 17 genes revealed causal associations with ASD and were identified as potential targets for ASD patients. Cathepsin B (CTSB) [odd ratio (OR) = 2.66 95, confidence interval (CI): 1.28-5.52, P = 8.84 × 10-3], gamma-aminobutyric acid type B receptor subunit 1 (GABBR1) (OR = 1.99, 95CI: 1.06-3.75, P = 3.24 × 10-2), and formin like 1 (FMNL1) (OR = 0.15, 95CI: 0.04-0.58, P = 5.59 × 10-3) were replicated in the proteome-wide MR analyses. In Drugbank, two potential therapeutic drugs, Acamprosate (GABBR1 inhibitor) and Bryostatin 1 (CASP8 inhibitor), were inferred as potential influencers of autism. Knockout mouse models suggested the involvement of the CASP8, GABBR1, and PLEKHM1 genes in neurological processes. Our findings suggest 17 candidate therapeutic targets for ASD and provide novel drug targets for therapy development and critical drug repurposing opportunities.
Xinqi Qiu, Jianyi Chen, Ruibang Luo, Ruijie Zeng, Shuangshuang Tong, Yanlin Lyu, Panpan Sun, Qizhou Lian, Felix W. Leung, Weihong Sha
Briefings Bioinform.5
2023 Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP
abstract
BACKGROUND: With the continuous advances in third-generation sequencing technology and the increasing affordability of next-generation sequencing technology, sequencing data from different sequencing technology platforms is becoming more common. While numerous benchmarking studies have been conducted to compare variant-calling performance across different platforms and approaches, little attention has been paid to the potential of leveraging the strengths of different platforms to optimize overall performance, especially integrating Oxford Nanopore and Illumina sequencing data. RESULTS: We investigated the impact of multi-platform data on the performance of variant calling through carefully designed experiments with a deep learning-based variant caller named Clair3-MP (Multi-Platform). Through our research, we not only demonstrated the capability of ONT-Illumina data for improved variant calling, but also identified the optimal scenarios for utilizing ONT-Illumina data. In addition, we revealed that the improvement in variant calling using ONT-Illumina data comes from an improvement in difficult genomic regions, such as the large low-complexity regions and segmental and collapse duplication regions. Moreover, Clair3-MP can incorporate reference genome stratification information to achieve a small but measurable improvement in variant calling. Clair3-MP is accessible as an open-source project at: https://github.com/HKU-BAL/Clair3-MP . CONCLUSIONS: These insights have important implications for researchers and practitioners alike, providing valuable guidance for improving the reliability and efficiency of genomic analysis in diverse applications.
Huijing Yu, Zhenxian Zheng, Junhao Su, Tak Wah Lam, Ruibang Luo
BMC Bioinform.5
2022 Clair3-trio: high-performance Nanopore long-read variant calling in family trios with trio-to-trio deep neural networks
abstract
Accurate identification of genetic variants from family child-mother-father trio sequencing data is important in genomics. However, state-of-the-art approaches treat variant calling from trios as three independent tasks, which limits their calling accuracy for Nanopore long-read sequencing data. For better trio variant calling, we introduce Clair3-Trio, the first variant caller tailored for family trio data from Nanopore long-reads. Clair3-Trio employs a Trio-to-Trio deep neural network model, which allows it to input the trio sequencing information and output all of the trio's predicted variants within a single model to improve variant calling. We also present MCVLoss, a novel loss function tailor-made for variant calling in trios, leveraging the explicit encoding of the Mendelian inheritance. Clair3-Trio showed comprehensive improvement in experiments. It predicted far fewer Mendelian inheritance violation variations than current state-of-the-art methods. We also demonstrated that our Trio-to-Trio model is more accurate than competing architectures. Clair3-Trio is accessible as a free, open-source project at https://github.com/HKU-BAL/Clair3-Trio.
Junhao Su, Zhenxian Zheng, Syed Shakeel Ahmed, Tak Wah Lam, Ruibang Luo
Briefings Bioinform.5
2022 Duet: SNP-assisted structural variant calling and phasing using Oxford nanopore sequencing
abstract
BACKGROUND: Whole genome sequencing using the long-read Oxford Nanopore Technologies (ONT) MinION sequencer provides a cost-effective option for structural variant (SV) detection in clinical applications. Despite the advantage of using long reads, however, accurate SV calling and phasing are still challenging. RESULTS: We introduce Duet, an SV detection tool optimized for SV calling and phasing using ONT data. The tool uses novel features integrated from both SV signatures and single-nucleotide polymorphism signatures, which can accurately distinguish SV haplotype from a false signal. Duet was benchmarked against state-of-the-art tools on multiple ONT sequencing datasets of sequencing coverage ranging from 8× to 40×. At low sequencing coverage of 8×, Duet performs better than all other tools in SV calling, SV genotyping and SV phasing. When the sequencing coverage is higher (20× to 40×), the F1-score for SV phasing is further improved in comparison to the performance of other tools, while its performance of SV genotyping and SV calling remains higher than other tools. CONCLUSION: Duet can perform accurate SV calling, SV genotyping and SV phasing using low-coverage ONT data, making it very useful for low-coverage genomes. It has great performance when scaled to high-coverage genomes, which is adaptable to various clinical applications. Duet is open source and is available at https://github.com/yekaizhou/duet .
Yekai Zhou, Amy Wing-Sze Leung, Syed Shakeel Ahmed, Tak Wah Lam, Ruibang Luo
BMC Bioinform.5
2020 ChromSeg: Two-Stage Framework for Overlapping Chromosome Segmentation and Reconstruction
abstract
Karyotyping is the most commonly used genetic tool for diagnosing diseases associated with chromosomal abnormalities. It generates images of the chromosomes of a patient in which quantity or shape discrepancies against normal chromosomes might suggest chromosomal abnormalities. However, the current methods are cumbersome and require manual or half-automatic separation of overlapping chromosomes, significantly limiting the productivity of clinical geneticists and cytologists. In this project, we implemented a fully automatic method, called ChromSeg, which efficiently separates crossing-overlap chromosomes. It uses a new neural network architecture called “region-guided UNet++” to accurately detect crossing-overlap chromosomes from metaphase cell images. A new heuristic algorithm, called “crossing-partition”, is then applied to splice and reconstruct the crossing-overlap chromosomes into single chromosomes. While there are a very limited number of publicly accessible annotations on overlapping chromosomes, we manually annotated 345 images for our model training and performance testing. Benchmarking results showed that our method achieved 99.1% overlap detection on crossing-overlap chromosomes and outperformed the second best method by 3.1%. Notably, this is the first tool to provide an image of the reconstructed chromosomes; other tools provide only segmentation suggestions, which are of less value to end-users. The source code of ChromSeg is available at https://github.com/HKU-BAL/ChromSeg, and the 345 annotated images are available at http://www.bio8.cs.hku.hk/bibm/.
Fangzhou Lan, Chi-Man Liu, Tak Wah Lam, Ruibang Luo
BIBM5
2020 MegaPath-Nano: Accurate Compositional Analysis and Drug-level Antimicrobial Resistance Detection Software for Oxford Nanopore Long-read Metagenomics
abstract
Accurate and sensitive taxonomic profiling is essential for any metagenomic analysis to reveal microbial community structure and for potential functional prediction. Antimicrobial resistance (AMR) detection is also a critical task in the clinical diagnosis of infection and antimicrobial therapy. By incorporating Oxford Nanopore Technologies (ONT) sequencing, users benefit from the high-confidence alignment of long reads for taxonomic classification, even among bacteria with similar genomes. Portable ONT devices, such as VolTRAX with MinION, allow short turnaround time for detection and can be used in a lightweight laboratory setting. However, error-prone ONT sequencing reads are still challenging for existing software for accurate taxonomic classification of microbes and detection of AMR down to the drug level. In this paper, we present MegaPath-Nano, the successor to NGS-based MegaPath. It is a high-precision compositional analysis software with drug-level AMR detection for ONT metagenomic sequencing data. MegaPath-Nano performs 1) thorough multi-level filtering against decoy and human reads while removing noisy alignments, 2) alignment-based taxonomic classification with RefSeq down to strain-level, with an alignment-reassignment algorithm to tackle the challenge of non-unique alignments, based on global alignment distribution, and 3) comprehensive downstream drug-level AMR detection, integrating five AMR databases. In our benchmarks using the Zymo metagenomic dataset, MegaPath-Nano performed better than other existing software for taxonomic classification. We also sequenced five real patient isolates using MinION to benchmark its performance of AMR detection. MegaPath-Nano was the most accurate and provided the most comprehensive output at both the drug and class level of AMR prediction against other state-of-the-art software. MegaPath-Nano is open-source and available at https://github.com/HKU-BAL/MegaPath-Nano.
Wui Wang Lui, Amy Wing-Sze Leung, Henry C. M. Leung, Jade L. L. Teng, Patrick C. Y. Woo, Tak Wah Lam, Ruibang Luo
BIBM8
2020 MC-Explorer: Analyzing and Visualizing Motif-Cliques on Large Networks
abstract
Large networks with labeled nodes are prevalent in various applications, such as biological graphs, social networks, and e-commerce graphs. To extract insight from this rich information source, we propose MC-Explorer, which is an advanced analysis and visualization system. A highlight of MC-Explorer is its ability to discover motif-cliques from a graph with labeled nodes. A motif, such as a 3-node triangle, is a fundamental building block of a graph. A motif-clique is a "complete" subgraph in a network with respect to a desired higher-order connection pattern. For example, on a large biological graph, we found out some motif-cliques, which disclose new side effects of a drug, and potential drugs for healing diseases. MC-Explorer includes online and interactive facilities for exploring a large labeled network through the use of motif-cliques. We will demonstrate how MC-Explorer can facilitate the analysis and visualization of a labeled biological network.An online demo video of MC-Explorer can be accessed from https://www.dropbox.com/s/vkalumc28wqp8yl/demo.mov.
Boxuan Li, Reynold Cheng, Jiafeng Hu, Yixiang Fang, Min Ou, Ruibang Luo, Kevin Chen-Chuan Chang, Xuemin Lin 0001
ICDE6
2019 RENET: A Deep Learning Approach for Extracting Gene-Disease Associations from Literature
Ye Wu 0007, Ruibang Luo, Henry C. M. Leung, Hing-Fung Ting, Tak Wah Lam
RECOMB2
2018 AC-DIAMOND v1: accelerating large-scale DNA-protein alignment
abstract
Summary: AC-DIAMOND (v1) is a DNA-protein alignment tool designed to tackle the efficiency challenge of aligning large amount of reads or contigs to protein databases. When compared with the previously most efficient method DIAMOND, AC-DIAMOND gains a 6- to 7-fold speed-up, while retaining a similar degree of sensitivity. The improvement is rooted at two aspects: first, using a compressed index of seeds with adaptive-length to speed-up the matching between query and reference sequences; second, adopting a compact form of dynamic programing to fully utilize the parallelism of the SIMD capability. Availability and implementation: Software source codes and binaries available at https://github.com/Maihj/AC-DIAMOND/. Supplementary information: Supplementary data are available at Bioinformatics online.
Huijun Mai, Dinghua Li, Henry C. M. Leung, Ruibang Luo, Chi-Kwong Wong, Hing-Fung Ting, Tak Wah Lam
Bioinform.5
2017 MegaGTA: a sensitive and accurate metagenomic gene-targeted assembler using iterative de Bruijn graphs
abstract
BACKGROUND: The recent release of the gene-targeted metagenomics assembler Xander has demonstrated that using the trained Hidden Markov Model (HMM) to guide the traversal of de Bruijn graph gives obvious advantage over other assembly methods. Xander, as a pilot study, indeed has a lot of room for improvement. Apart from its slow speed, Xander uses only 1 k-mer size for graph construction and whatever choice of k will compromise either sensitivity or accuracy. Xander uses a Bloom-filter representation of de Bruijn graph to achieve a lower memory footprint. Bloom filters bring in false positives, and it is not clear how this would impact the quality of assembly. Xander does not keep track of the multiplicity of k-mers, which would have been an effective way to differentiate between erroneous k-mers and correct k-mers. RESULTS: In this paper, we present a new gene-targeted assembler MegaGTA, which attempts to improve Xander in different aspects. Quality-wise, it utilizes iterative de Bruijn graphs to take full advantage of multiple k-mer sizes to make the best of both sensitivity and accuracy. Computation-wise, it employs succinct de Bruijn graphs (SdBG) to achieve low memory footprint and high speed (the latter is benefited from a highly efficient parallel algorithm for constructing SdBG). Unlike Bloom filters, an SdBG is an exact representation of a de Bruijn graph. It enables MegaGTA to avoid false-positive contigs and to easily incorporate the multiplicity of k-mers for building better HMM model. We have compared MegaGTA and Xander on an HMP-defined mock metagenomic dataset, and showed that MegaGTA excelled in both sensitivity and accuracy. On a large rhizosphere soil metagenomic sample (327Gbp), MegaGTA produced 9.7-19.3% more contigs than Xander, and these contigs were assigned to 10-25% more gene references. In our experiments, MegaGTA, depending on the number of k-mers used, is two to ten times faster than Xander. CONCLUSION: MegaGTA improves on the algorithm of Xander and achieves higher sensitivity, accuracy and speed. Moreover, it is capable of assembling gene sequences from ultra-large metagenomic datasets. Its source code is freely available at https://github.com/HKU-BAL/megagta .
Dinghua Li, Henry C. M. Leung, Ruibang Luo, Hing-Fung Ting, Tak Wah Lam
BMC Bioinform.4
2015 MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph
abstract
Abstract Summary: MEGAHIT is a NGS de novo assembler for assembling large and complex metagenomics data in a time- and cost-efficient manner. It finished assembling a soil metagenomics dataset with 252 Gbps in 44.1 and 99.6 h on a single computing node with and without a graphics processing unit, respectively. MEGAHIT assembles the data as a whole, i.e. no pre-processing like partitioning and normalization was needed. When compared with previous methods on assembling the soil data, MEGAHIT generated a three-time larger assembly, with longer contig N50 and average contig length; furthermore, 55.8% of the reads were aligned to the assembly, giving a fourfold improvement. Availability and implementation: The source code of MEGAHIT is freely available at https://github.com/voutcn/megahit under GPLv3 license. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Dinghua Li, Chi-Man Liu, Ruibang Luo, Kunihiko Sadakane, Tak Wah Lam
Bioinform.3
2015 database.bio: a web application for interpreting human variations
abstract
UNLABELLED: Rapid advances of next-generation sequencing technology have led to the integration of genetic information with clinical care. Genetic basis of diseases and response to drugs provide new ways of disease diagnosis and safer drug usage. This integration reveals the urgent need for effective and accurate tools to analyze genetic variants. Due to the number and diversity of sources for annotation, automating variant analysis is a challenging task. Here, we present database.bio, a web application that combines variant annotation, prioritization and visualization so as to support insight into the individual genetic characteristics. It enhances annotation speed by preprocessing data on a supercomputer, and reduces database space via a unified database representation with compressed fields. AVAILABILITY AND IMPLEMENTATION: Freely available at https://database.bio.
Min Ou, Ricky Ma, Jeanno Cheung, Katie Lo, Patrick Yee, Tewei Luo, T. L. Chan, Chun Hang Au, Ava Kwong, Ruibang Luo, Tak Wah Lam
Bioinform.10
2015 MICA: A fast short-read aligner that takes full advantage of Many Integrated Core Architecture (MIC)
abstract
BACKGROUND: Short-read aligners have recently gained a lot of speed by exploiting the massive parallelism of GPU. An uprising alterative to GPU is Intel MIC; supercomputers like Tianhe-2, currently top of TOP500, is built with 48,000 MIC boards to offer ~55 PFLOPS. The CPU-like architecture of MIC allows CPU-based software to be parallelized easily; however, the performance is often inferior to GPU counterparts as an MIC card contains only ~60 cores (while a GPU card typically has over a thousand cores). RESULTS: To better utilize MIC-enabled computers for NGS data analysis, we developed a new short-read aligner MICA that is optimized in view of MIC's limitation and the extra parallelism inside each MIC core. By utilizing the 512-bit vector units in the MIC and implementing a new seeding strategy, experiments on aligning 150 bp paired-end reads show that MICA using one MIC card is 4.9 times faster than BWA-MEM (using 6 cores of a top-end CPU), and slightly faster than SOAP3-dp (using a GPU). Furthermore, MICA's simplicity allows very efficient scale-up when multiple MIC cards are used in a node (3 cards give a 14.1-fold speedup over BWA-MEM). SUMMARY: MICA can be readily used by MIC-enabled supercomputers for production purpose. We have tested MICA on Tianhe-2 with 90 WGS samples (17.47 Tera-bases), which can be aligned in an hour using 400 nodes. MICA has impressive performance even though MIC is only in its initial stage of development. AVAILABILITY AND IMPLEMENTATION: MICA's source code is freely available at http://sourceforge.net/projects/mica-aligner under GPL v3. SUPPLEMENTARY INFORMATION: Supplementary information is available as "Additional File 1". Datasets are available at www.bio8.cs.hku.hk/dataset/mica.
Ruibang Luo, Jeanno Cheung, Edward Wu, Sze-Hang Chan, Wai-Chun Law, Guangzhu He, Chi-Man Liu, Dazong Zhou, Yingrui Li, Ruiqiang Li, Jun Wang 0004, Xiaoqian Zhu, Shaoliang Peng, Tak Wah Lam
BMC Bioinform.1
2014 FaSD-somatic: a fast and accurate somatic SNV detection algorithm for cancer genome sequencing data
abstract
UNLABELLED: Recent advances in high-throughput sequencing technologies have enabled us to sequence large number of cancer samples to reveal novel insights into oncogenetic mechanisms. However, the presence of intratumoral heterogeneity, normal cell contamination and insufficient sequencing depth, together pose a challenge for detecting somatic mutations. Here we propose a fast and an accurate somatic single-nucleotide variations (SNVs) detection program, FaSD-somatic. The performance of FaSD-somatic is extensively assessed on various types of cancer against several state-of-the-art somatic SNV detection programs. Benchmarked by somatic SNVs from either existing databases or de novo higher-depth sequencing data, FaSD-somatic has the best overall performance. Furthermore, FaSD-somatic is efficient, it finishes somatic SNV calling within 14 h on 50X whole genome sequencing data in paired samples. AVAILABILITY AND IMPLEMENTATION: The program, datasets and supplementary files are available at http://jjwanglab.org/FaSD-somatic/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Weixin Wang 0004, Panwen Wang, Ruibang Luo, Maria P. Wong, Tak Wah Lam, Junwen Wang
Bioinform.4
2014 SOAPdenovo-Trans: de novo transcriptome assembly with short RNA-Seq reads
abstract
MOTIVATION: Transcriptome sequencing has long been the favored method for quickly and inexpensively obtaining a large number of gene sequences from an organism with no reference genome. Owing to the rapid increase in throughputs and decrease in costs of next- generation sequencing, RNA-Seq in particular has become the method of choice. However, the very short reads (e.g. 2 × 90 bp paired ends) from next generation sequencing makes de novo assembly to recover complete or full-length transcript sequences an algorithmic challenge. RESULTS: Here, we present SOAPdenovo-Trans, a de novo transcriptome assembler designed specifically for RNA-Seq. We evaluated its performance on transcriptome datasets from rice and mouse. Using as our benchmarks the known transcripts from these well-annotated genomes (sequenced a decade ago), we assessed how SOAPdenovo-Trans and two other popular transcriptome assemblers handled such practical issues as alternative splicing and variable expression levels. Our conclusion is that SOAPdenovo-Trans provides higher contiguity, lower redundancy and faster execution. AVAILABILITY AND IMPLEMENTATION: Source code and user manual are available at http://sourceforge.net/projects/soapdenovotrans/.
Yinlong Xie, Gengxiong Wu, Jingbo Tang, Ruibang Luo, Jordan Patterson, Shanlin Liu, Weihua Huang, Guangzhu He, Shengchang Gu, Shengkang Li, Tak Wah Lam, Yingrui Li, Gane Ka-Shu Wong, Jun Wang 0004
Bioinform.4
2012 SOAP3: ultra-fast GPU-based parallel alignment tool for short reads
abstract
Abstract Summary: SOAP3 is the first short read alignment tool that leverages the multi-processors in a graphic processing unit (GPU) to achieve a drastic improvement in speed. We adapted the compressed full-text index (BWT) used by SOAP2 in view of the advantages and disadvantages of GPU. When tested with millions of Illumina Hiseq 2000 length-100 bp reads, SOAP3 takes < 30 s to align a million read pairs onto the human reference genome and is at least 7.5 and 20 times faster than BWA and Bowtie, respectively. For aligning reads with up to four mismatches, SOAP3 aligns slightly more reads than BWA and Bowtie; this is because SOAP3, unlike BWA and Bowtie, is not heuristic-based and always reports all answers. Availability: SOAP3 is available at: http://www.cs.hku.hk/2bwt-tools/soap3; http://soap.genomics.org.cn/soap3.html. Contact: [email protected], [email protected]
Chi-Man Liu, Thomas K. F. Wong, Edward Wu, Ruibang Luo, Siu-Ming Yiu, Yingrui Li, Bingqiang Wang, Xiaowen Chu 0001, Kaiyong Zhao, Ruiqiang Li, Tak Wah Lam
Bioinform.4
2012 COPE: an accurate k-mer-based pair-end reads connection tool to facilitate genome assembly
abstract
MOTIVATION: The boost of next-generation sequencing technologies provides us with an unprecedented opportunity for elucidating genetic mysteries, yet the short-read length hinders us from better assembling the genome from scratch. New protocols now exist that can generate overlapping pair-end reads. By joining the 3' ends of each read pair, one is able to construct longer reads for assembling. However, effectively joining two overlapped pair-end reads remains a challenging task. RESULT: In this article, we present an efficient tool called Connecting Overlapped Pair-End (COPE) reads, to connect overlapping pair-end reads using k-mer frequencies. We evaluated our tool on 30× simulated pair-end reads from Arabidopsis thaliana with 1% base error. COPE connected over 99% of reads with 98.8% accuracy, which is, respectively, 10 and 2% higher than the recently published tool FLASH. When COPE is applied to real reads for genome assembly, the resulting contigs are found to have fewer errors and give a 14-fold improvement in the N50 measurement when compared with the contigs produced using unconnected reads. AVAILABILITY AND IMPLEMENTATION: COPE is implemented in C++ and is freely available as open-source code at ftp://ftp.genomics.org.cn/pub/cope. CONTACT: [email protected] or [email protected]
Binghang Liu, Jianying Yuan, Siu-Ming Yiu, Yinlong Xie, Yujian Shi, Yingrui Li, Tak Wah Lam, Ruibang Luo
Bioinform.11