Inanç Birol

dblp:69/1147 · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-0950-7839ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 ntStat: k-mer characterization using occurrence statistics in raw sequencing data
abstract
K-mer counts are fundamental in many genomic data analysis tasks, providing valuable information for genome assembly, error correction, and variant detection. State-of-the-art k-mer counting tools employ various techniques, such as parallelism, probabilistic data structures, and disk utilization, to efficiently extract k-mer frequencies from large datasets. The distribution of k-mer counts in raw sequencing reads reveals key genomic characteristics such as genome size, heterozygosity, and basecalling quality. The number of reads containing a k-mer has also shown application in genome assembly and sequence analysis. We present ntStat, a toolkit that employs succinct Bloom filter data structures to track both k-mer count and depth information and use in downstream applications. ntStat models the k-mer count histogram using evolutionary computation, and infers valuable insights about the genome, sequencing data, and individual k-mers, de novo. ntStat consistently ran faster than DSK, BFCounter, hackgap, and Squeakr in all of our tests. Jellyfish performed faster than ntStat for human data with k = 25 but fell behind with k = 64. KMC3 was faster overall but at a high disk usage and memory cost. ntStat also used less memory than other non-disk-based k-mer counters and typically, 99.5-99.9% of the k-mers processed by ntStat are counted correctly. ntStat's histogram analysis module detected heterozygosity percentages and k-mer coverage for long-read datasets simulated from a diploid human genome with less than 1% and 0.5-fold difference to the ground truth. The analysis of simulated long read datasets showed an average error of just 2% in k-mer robustness estimates.
Parham Kazemi, Lauren Coombe, René L. Warren, Inanç Birol
PLoS Comput. Biol.4
2025 GoldPolish-target: targeted long-read genome assembly polishing
abstract
BACKGROUND: Advanced long-read sequencing technologies, such as those from Oxford Nanopore Technologies and Pacific Biosciences, are finding a wide use in de novo genome sequencing projects. However, long reads typically have higher error rates relative to short reads. If left unaddressed, subsequent genome assemblies may exhibit high base error rates that compromise the reliability of downstream analysis. Several specialized error correction tools for genome assemblies have since emerged, employing a range of algorithms and strategies to improve base quality. However, despite these efforts, many genome assembly workflows still produce regions with elevated error rates, such as gaps filled with unpolished or ambiguous bases. To address this, we introduce GoldPolish-Target, a modular targeted sequence polishing pipeline. Coupled with GoldPolish, a linear-time genome assembly algorithm, GoldPolish-Target isolates and polishes user-specified assembly loci, offering a resource-efficient means for polishing targeted regions of draft genomes. RESULTS: Experiments using Drosophila melanogaster and Homo sapiens datasets demonstrate that GoldPolish-Target can reduce insertion/deletion (indel) and mismatch errors by up to 49.2% and 55.4% respectively, achieving base accuracy values upwards of 99.9% (Phred score Q > 30). This polishing accuracy is comparable to the current state-of-the-art, Medaka, while exhibiting up to 27-fold shorter run times and consuming 95% less memory, on average. CONCLUSION: GoldPolish-Target, in contrast to most other polishing tools, offers the ability to target specific regions of a genome assembly for polishing, providing a computationally light-weight and highly scalable solution for base error correction.
Emily Zhang, Lauren Coombe, Johnathan Wong, René L. Warren, Inanç Birol
BMC Bioinform.5
2022 ntHash2: recursive spaced seed hashing for nucleotide sequences
abstract
MOTIVATION: Spaced seeds are robust alternatives to k-mers in analyzing nucleotide sequences with high base mismatch rates. Hashing is also crucial for efficiently storing abundant sequence data. Here, we introduce ntHash2, a fast algorithm for spaced seed hashing that can be integrated into various bioinformatics tools for efficient sequence analysis with applications in genome research. RESULTS: ntHash2 is up to 2.1× faster at hashing various spaced seeds than the previous version and 3.8× faster than conventional hashing algorithms with naïve adaptation. Additionally, we reduced the collision rate of ntHash for longer k-mer lengths and improved the uniformity of the hash distribution by modifying the canonical hashing mechanism. AVAILABILITY AND IMPLEMENTATION: ntHash2 is freely available online at github.com/bcgsc/ntHash under an MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Parham Kazemi, Johnathan Wong, Vladimir Nikolic, Hamid Mohamadi, René L. Warren, Inanç Birol
Bioinform.6
2022 RResolver: efficient short-read repeat resolution within ABySS
abstract
BACKGROUND: De novo genome assembly is essential to modern genomics studies. As it is not biased by a reference, it is also a useful method for studying genomes with high variation, such as cancer genomes. De novo short-read assemblers commonly use de Bruijn graphs, where nodes are sequences of equal length k, also known as k-mers. Edges in this graph are established between nodes that overlap by [Formula: see text] bases, and nodes along unambiguous walks in the graph are subsequently merged. The selection of k is influenced by multiple factors, and optimizing this value results in a trade-off between graph connectivity and sequence contiguity. Ideally, multiple k sizes should be used, so lower values can provide good connectivity in lesser covered regions and higher values can increase contiguity in well-covered regions. However, current approaches that use multiple k values do not address the scalability issues inherent to the assembly of large genomes. RESULTS: Here we present RResolver, a scalable algorithm that takes a short-read de Bruijn graph assembly with a starting k as input and uses a k value closer to that of the read length to resolve repeats. RResolver builds a Bloom filter of sequencing reads which is used to evaluate the assembly graph path support at branching points and removes paths with insufficient support. RResolver runs efficiently, taking only 26 min on average for an ABySS human assembly with 48 threads and 60 GiB memory. Across all experiments, compared to a baseline assembly, RResolver improves scaffold contiguity (NGA50) by up to 15% and reduces misassemblies by up to 12%. CONCLUSIONS: RResolver adds a missing component to scalable de Bruijn graph genome assembly. By improving the initial and fundamental graph traversal outcome, all downstream ABySS algorithms greatly benefit by working with a more accurate and less complex representation of the genome. The RResolver code is integrated into ABySS and is available at https://github.com/bcgsc/abyss/tree/master/RResolver .
Vladimir Nikolic, Amirhossein Afshinfard, Justin Chu, Johnathan Wong, Lauren Coombe, Ka Ming Nip, René L. Warren, Inanç Birol
BMC Bioinform.8
2021 HLA predictions from the bronchoalveolar lavage fluid and blood samples of eight COVID-19 patients at the pandemic onset
abstract
Severe acute respiratory syndrome (SARS)-CoV-2 infections have reached global pandemic proportions in early 2020, affecting over 21 M people worldwide (as of this writing in August 2020; Source: Johns Hopkins University) and showing no signs of easing, except in a few jurisdictions where strict quarantine measures were implemented early on. The resulting coronavirus disease (COVID-19) has a relatively high (∼3.4%) mortality rate (Rajgor et al., 2020)—a figure that varies widely between jurisdictions due to factors yet to be determined. Currently, no vaccines or effective treatments are available. Most current data analysis efforts are, understandably, focussed on the virus itself for the purpose of vaccine development and tracking its evolution for diagnostics and infection monitoring purposes. Curiously, it is estimated that as high as 18–30% or more of the population may be asymptomatic to SARS-CoV-2 infections (Mizumoto et al., 2020; Nishiura et al., 2020), while other affected individuals exhibit mild to severe to critical symptoms of infection. Thus, gaining insights on host susceptibility to the coronavirus is clearly another important aspect that needs to be worked on and understood (Shi et al., 2020). One would expect a link between host immunity genes and susceptibility or resistance to infection. The Human Leukocyte Antigen (HLA) gene complex includes two classes of such genes, which encode the Major Histocompatibility Complex (MHC). Proteins of the MHC present (class I) internally- or (class II) externally-derived antigenic determinants (epitopes) to T cells, which upon recognition of the epitope-complex, will mount an immune response to defend against viral and bacterial infections. HLA genes are, therefore, cornerstone to acquired immunity in humans. HLA alleles have also been shown to be factors in susceptibility or resistance to certain diseases, and their frequency and composition in human populations vary widely (http://allelefrequencies.net/). A previous study found HLA class I (HLA-I) genes HLA-B*46:01 and HLA-B*54:01 to be associated with the 2003 SARS coronavirus infections in Taiwan (Lin et al., 2003)—a related disease to the current pandemic. For over a decade, high-throughput transcriptome sequencing has proven a worthy instrument for measuring changes of gene expression in human diseases and beyond (Wang et al., 2009). Transcriptome analysis has the potential to reveal key genes that are modulated in response to infections, but also has the potential to reveal the HLA composition of affected individuals. A few years ago, we developed an approach for mining high-throughput next-generation shotgun sequencing data for the purpose of HLA determination (Warren et al., 2012), which has since been applied in a broader clinical context (Brown et al., 2014). Here, we report our initial observations based on transcriptome sequencing (RNA-Seq) libraries prepared from the bronchoalveolar lavage (BAL) fluid and peripheral blood mononuclear cell (PMBC) samples of five and three COVID-19 patients at the early stage of the pneumonia coronavirus outbreak in Wuhan, China, respectively (Xiong et al., 2020; Zhou et al., 2020) (see Section 2). Of note, we identified the HLA-I group A allele A*24:02 in four out of five individuals from the first cohort and class II haplotype DPA1*02:02-DPB1*05:01 in seven out of eight individuals from both cohorts. We downloaded MGISEQ-2000RS paired-end (150 bp) RNA-Seq reads from libraries prepared from the BAL fluid samples of five patients [https://www.ebi.ac.uk/ena/browser/view/PRJNA605983 Accessions: SRX7730880-SRX7730884 denoted in the tables as Patients 1–5, respectively (Zhou et al., 2020)] and BGISEQ-500 paired-end (100 bp) RNA-Seq reads derived from the PMBC samples of three COVID-19 patients from a different study [Run accessions: CRR119891-3 from BIG Data Center (https://bigd.big.ac.cn/) project CRA002390 denoted in the tables as Patients 6–8, respectively (Xiong et al., 2020)]. We note that these are metatranscriptomic RNA samples prepared for the primary purpose of identifying/characterizing the novel coronavirus and identifying host response genes. On each dataset, we ran HLAminer (Warren et al., 2012) in targeted assembly mode with default values (v1.4; contig length ≥200 bp, seq. identity ≥99%, score ≥1000, e-value ≥25), predicting HLA-I and class II (HLA-II) alleles and report 4-digit (HLA allele/protein) resolution when top-scoring predictions are unambiguous. Otherwise the 2-digit (allele group) resolution is reported. We also ran HLA prediction software seq2HLA (Boegel et al., 2012; v2.3), OptiType (Szolek et al., 2014; v1.3.4) and arcasHLA (Orenbuch et al., 2020; v0.2.0 with latest code commit 301085e) on the RNA-Seq data derived from the BAL samples of COVID-19 patients (Supplementary Methods). The tool used to perform the reported analysis results in Tables 1 and 2, HLAminer, is available from https://github.com/bcgsc/hlaminer. Predictions are available for download from https://www.bcgsc.ca/downloads/btl/SARS-CoV-2/BAL. We predicted and compiled the likely HLA-I (Table 1) and HLA-II (Table 2) alleles of eight patients at the early stage of the COVID-19 outbreak in Wuhan, China. In the first cohort comprised of five patients, although the BAL fluid samples were initially utilized to identify and characterize the novel coronavirus [Zhou et al. (2020) with similar justification in Wu et al. (2020)], BAL metagenomics samples are expected to contain host cells/nucleic acids (DNA/RNA). Because HLA genes are expressed at the surface of all human nucleated cells, RNA-Seq data can be employed to determine HLA profiles from BAL samples. HLA-I predictions from the BAL fluid samples of five patients at the early stage of the Wuhan seafood market pneumonia coronavirus outbreak and from the PBMC samples of three COVID-19 patients from a different cohort/study Note: Highest-scoring HLAminer predictions are shown for each HLA-I genes A, B and C. Missing class I genes or (—) denote the absence of a second prediction. Common HLA alleles between two or more patients of a given cohort are highlighted in bold. Ambiguous predictions are shown at the group (2-digit) resolution. HLA-I predictions from the BAL fluid samples of five patients at the early stage of the Wuhan seafood market pneumonia coronavirus outbreak and from the PBMC samples of three COVID-19 patients from a different cohort/study Note: Highest-scoring HLAminer predictions are shown for each HLA-I genes A, B and C. Missing class I genes or (—) denote the absence of a second prediction. Common HLA alleles between two or more patients of a given cohort are highlighted in bold. Ambiguous predictions are shown at the group (2-digit) resolution. HLA-II predictions from the BAL fluid samples of five patients at the early stage of the Wuhan seafood market pneumonia coronavirus outbreak and from the PBMC samples of three COVID-19 patients from a different cohort/study Note: Highest-scoring HLAminer predictions are shown for HLA-II genes DPA1, DPB1, DQA1, DQB1 and DRB4. Missing class II genes or (—) denote the absence of predictions. Common HLA alleles between two or more patients of a given cohort are highlighted in bold. Ambiguous predictions are shown at the group (2-digit) resolution. HLA-II predictions from the BAL fluid samples of five patients at the early stage of the Wuhan seafood market pneumonia coronavirus outbreak and from the PBMC samples of three COVID-19 patients from a different cohort/study Note: Highest-scoring HLAminer predictions are shown for HLA-II genes DPA1, DPB1, DQA1, DQB1 and DRB4. Missing class II genes or (—) denote the absence of predictions. Common HLA alleles between two or more patients of a given cohort are highlighted in bold. Ambiguous predictions are shown at the group (2-digit) resolution. We observe the HLA-A*24:02 allele in four out of five (80%) patients of the first cohort, but this allele was not predicted in the second cohort whose SARS-CoV-2 positive patients were also admitted in a Wuhan hospital (Table 1) (Xiong et al., 2020). In the absence of an equivalent BAL metatranscriptomic dataset with known HLA genotypes, we opted to run the established prediction software seq2HLA (Boegel et al., 2012) and OptiType (Szolek et al., 2014) and the recently published utility arcasHLA (Orenbuch et al., 2020), in an effort to help validate the predictions we obtained from HLAminer on the BAL cohort RNA-Seq (Supplementary Table S1). We observe good concordance among the tools, with the highest concordance observed between HLAminer and seq2HLA when predicting HLA-I gene A (100% concordance) and between HLAminer and OptiType when predicting HLA-I B and C genes (100% and 88.9%, respectively; Supplementary Table S1). We note that arcasHLA failed to output predictions altogether on these metatranscriptomic data, likely due to low HLA signal. Of interest, the HLA-I allele A*24:02 prediction that we report in Table 1 was recapitulated with seq2HLA. HLA-A*24 is a common group of alleles in South-Eastern Asian populations, and the frequency of the HLA-A*24:02 allele can be high, especially in indigenous Taiwanese populations, reaching as high as 86.3% (http://allelefrequencies.net/). It has been reported that the five patients from the first cohort were sellers and delivery workers at the Wuhan seafood market, but since no information on patient ethnicity/ancestry is available (Zhou et al., 2020), no inferences can be made with respect to population frequency, especially given the small sample size. Also, of note, we observe the HLA-II DPA1*02:02 and DPB1*05:01 haplotype predicted in seven out of the eight (87.5%) patients (Table 2). We point out that HLA-A*24 has not been previously reported as a risk factor for SARS infection (Sun and Xi, 2014). There are reports of other disease association with HLA-A*24:02, notably with diabetes (Adamashvili et al., 1997; Kronenberg et al., 2012; Nakanishi and Inoko 2006; Noble et al., 2002), which is a recorded potential risk factor in COVID-19 patients (Guan et al., 2020). Both DPA1*02:02 and DPB1*05:01 occur at relative high frequency (44.8% and 31.3%, n=1490) in Han Chinese (Chu et al., 2018), and associations of those particular type II alleles with narcolepsy (Ollila et al., 2015) and Graves’ disease (Chu et al., 2018), both autoimmune disorders, have been reported in that population. Further, a genome-wide association study found a link between HLA-DPB1*05:01 and chronic hepatitis B in Asians, and it has been suggested that this risk allele may impact one’s ability to clear viral infections (Kamatani et al., 2009; Ollila et al., 2015). HLA also informs vaccine development. This knowledge would help prioritize SARS-CoV-2 derived epitopes predicted to be stable HLA binders (Kiyotani et al., 2020; Nguyen et al., 2020; Prachar et al., 2020; Yarmarkovich et al., 2020). HLA-I A*24:02 was reported to be among just a few allotypes that showed stable binding with more than 10 epitopes derived from the SARS-CoV-2 proteome (Kiyotani et al., 2020; Prachar et al., 2020). In contrast, previously reported SARS risk allele HLA-B*46:01 (Lin et al., 2003) had amongst the fewest number of predicted binding SARS-CoV-2 peptides (Nguyen et al., 2020). Further research into host susceptibility and resistance to SARS-CoV-2 infections on larger population cohorts and from different jurisdictions is sorely needed as it may help us better manage and mitigate risks of infections. We stress that our observations were derived from small sample sets, and caution that host susceptibility gene inferences require larger cohorts and properly designed data collection experiments with controls, to help quantify the false positive rate and confidence in predictions. Our letter highlights the technical feasibility and challenges associated with deriving HLA types directly from metatranscriptomic RNA-Seq libraries prepared from COVID-19 patient samples and not collected specifically for that purpose. We chose to communicate our early findings in this domain to facilitate rapid development of response strategies. This work was supported by Genome BC and Genome Canada [281ANV]; and the National Institutes of Health [2R01HG007182-04A1]. The content of this paper is solely the responsibility of the authors, and does not necessarily represent the official views of the National Institutes of Health or other funding organizations. Conflict of Interest: none declared.
René L. Warren, Inanç Birol
Bioinform.2
2021 LongStitch: high-quality genome assembly correction and scaffolding using long reads
abstract
BACKGROUND: Generating high-quality de novo genome assemblies is foundational to the genomics study of model and non-model organisms. In recent years, long-read sequencing has greatly benefited genome assembly and scaffolding, a process by which assembled sequences are ordered and oriented through the use of long-range information. Long reads are better able to span repetitive genomic regions compared to short reads, and thus have tremendous utility for resolving problematic regions and helping generate more complete draft assemblies. Here, we present LongStitch, a scalable pipeline that corrects and scaffolds draft genome assemblies exclusively using long reads. RESULTS: LongStitch incorporates multiple tools developed by our group and runs in up to three stages, which includes initial assembly correction (Tigmint-long), followed by two incremental scaffolding stages (ntLink and ARKS-long). Tigmint-long and ARKS-long are misassembly correction and scaffolding utilities, respectively, previously developed for linked reads, that we adapted for long reads. Here, we describe the LongStitch pipeline and introduce our new long-read scaffolder, ntLink, which utilizes lightweight minimizer mappings to join contigs. LongStitch was tested on short and long-read assemblies of Caenorhabditis elegans, Oryza sativa, and three different human individuals using corresponding nanopore long-read data, and improves the contiguity of each assembly from 1.2-fold up to 304.6-fold (as measured by NGA50 length). Furthermore, LongStitch generates more contiguous and correct assemblies compared to state-of-the-art long-read scaffolder LRScaf in most tests, and consistently improves upon human assemblies in under five hours using less than 23 GB of RAM. CONCLUSIONS: Due to its effectiveness and efficiency in improving draft assemblies using long reads, we expect LongStitch to benefit a wide variety of de novo genome assembly projects. The LongStitch pipeline is freely available at https://github.com/bcgsc/longstitch .
Lauren Coombe, Janet X. Li, Theodora Lo, Johnathan Wong, Vladimir Nikolic, René L. Warren, Inanç Birol
BMC Bioinform.7
2021 GapPredict - A Language Model for Resolving Gaps in Draft Genome Assemblies
abstract
bases per run, typically composed of reads 150 bases long. Despite this high throughput, de novo assembly algorithms have difficulty reconstructing contiguous genome sequences using short reads due to both repetitive and difficult-to-sequence regions in these genomes. Some of the short read assembly challenges are mitigated by scaffolding assembled sequences using paired-end reads. However, unresolved sequences in these scaffolds appear as "gaps". Here, we introduce GapPredict - An implementation of a proof of concept that uses a character-level language model to predict unresolved nucleotides in scaffold gaps. We benchmarked GapPredict against the state-of-the-art gap-filling tool Sealer, and observed that the former can fill 65.6% of the sampled gaps that were left unfilled by the latter with high similarity to the reference genome, demonstrating the practical utility of deep learning approaches to the gap-filling problem in genome assembly.
Justin Chu, René L. Warren, Inanç Birol
IEEE ACM Trans. Comput. Biol. Bioinform.5
2020 ENANO: Encoder for NANOpore FASTQ files
abstract
MOTIVATION: The amount of genomic data generated globally is seeing explosive growth, leading to increasing needs for processing, storage and transmission resources, which motivates the development of efficient compression tools for these data. Work so far has focused mainly on the compression of data generated by short-read technologies. However, nanopore sequencing technologies are rapidly gaining popularity due to the advantages offered by the large increase in the average size of the produced reads, the reduction in their cost and the portability of the sequencing technology. We present ENANO (Encoder for NANOpore), a novel lossless compression algorithm especially designed for nanopore sequencing FASTQ files. RESULTS: The main focus of ENANO is on the compression of the quality scores, as they dominate the size of the compressed file. ENANO offers two modes, Maximum Compression and Fast (default), which trade-off compression efficiency and speed. We tested ENANO, the current state-of-the-art compressor SPRING and the general compressor pigz on several publicly available nanopore datasets. The results show that the proposed algorithm consistently achieves the best compression performance (in both modes) on every considered nanopore dataset, with an average improvement over pigz and SPRING of >24.7% and 6.3%, respectively. In addition, in terms of encoding and decoding speeds, ENANO is 2.9× and 1.7× times faster than SPRING, respectively, with memory consumption up to 0.2 GB. AVAILABILITY AND IMPLEMENTATION: ENANO is freely available for download at: https://github.com/guilledufort/EnanoFASTQ. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guillermo Dufort, Gadiel Seroussi, Pablo Smircich, José Sotelo, Idoia Ochoa, Alvaro Martín, Inanç Birol
Bioinform.7
2020 Fusion-Bloom: fusion detection in assembled transcriptomes
abstract
SUMMARY: Presence or absence of gene fusions is one of the most important diagnostic markers in many cancer types. Consequently, fusion detection methods using various genomics data types, such as RNA sequencing (RNA-seq) are valuable tools for research and clinical applications. While information-rich RNA-seq data have proven to be instrumental in discovery of a number of hallmark fusion events, bioinformatics tools to detect fusions still have room for improvement. Here, we present Fusion-Bloom, a fusion detection method that leverages recent developments in de novo transcriptome assembly and assembly-based structural variant calling technologies (RNA-Bloom and PAVFinder, respectively). We benchmarked Fusion-Bloom against the performance of five other state-of-the-art fusion detection tools using multiple datasets. Overall, we observed Fusion-Bloom to display a good balance between detection sensitivity and specificity. We expect the tool to find applications in translational research and clinical genomics pipelines. AVAILABILITY AND IMPLEMENTATION: Fusion-Bloom is implemented as a UNIX Make utility, available at https://github.com/bcgsc/pavfinder and released under a Creative Commons License (Attribution 4.0 International), as described at http://creativecommons.org/licenses/by/4.0/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Readman Chiu, Ka Ming Nip, Inanç Birol
Bioinform.3
2020 ntJoin: Fast and lightweight assembly-guided scaffolding using minimizer graphs
abstract
SUMMARY: The ability to generate high-quality genome sequences is cornerstone to modern biological research. Even with recent advancements in sequencing technologies, many genome assemblies are still not achieving reference-grade. Here, we introduce ntJoin, a tool that leverages structural synteny between a draft assembly and reference sequence(s) to contiguate and correct the former with respect to the latter. Instead of alignments, ntJoin uses a lightweight mapping approach based on a graph data structure generated from ordered minimizer sketches. The tool can be used in a variety of different applications, including improving a draft assembly with a reference-grade genome, a short-read assembly with a draft long-read assembly and a draft assembly with an assembly from a closely related species. When scaffolding a human short-read assembly using the reference human genome or a long-read assembly, ntJoin improves the NGA50 length 23- and 13-fold, respectively, in under 13 m, using <11 GB of RAM. Compared to existing reference-guided scaffolders, ntJoin generates highly contiguous assemblies faster and using less memory. AVAILABILITY AND IMPLEMENTATION: ntJoin is written in C++ and Python and is freely available at https://github.com/bcgsc/ntjoin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lauren Coombe, Vladimir Nikolic, Justin Chu, Inanç Birol, René L. Warren
Bioinform.4
2020 Portable nanopore analytics: are we there yet?
abstract
MOTIVATION: Oxford Nanopore technologies (ONT) add miniaturization and real time to high-throughput sequencing. All available software for ONT data analytics run on cloud/clusters or personal computers. Instead, a linchpin to true portability is software that works on mobile devices of internet connections. Smartphones' and tablets' chipset/memory/operating systems differ from desktop computers, but software can be recompiled. We sought to understand how portable current ONT analysis methods are. RESULTS: Several tools, from base-calling to genome assembly, were ported and benchmarked on an Android smartphone. Out of 23 programs, 11 succeeded. Recompilation failures included lack of standard headers and unsupported instruction sets. Only DSK, BCALM2 and Kraken were able to process files up to 16 GB, with linearly scaling CPU-times. However, peak CPU temperatures were high. In conclusion, the portability scenario is not favorable. Given the fast market growth, attention of developers to ARM chipsets and Android/iOS is warranted, as well as initiatives to implement mobile-specific libraries. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at: https://github.com/marco-oliva/portable-nanopore-analytics.
Marco Oliva, Franco Milicchio, Kaden King, Grace Benson, Christina Boucher 0001, Mattia Prosperi, Inanç Birol
Bioinform.7
2019 ORCA: a comprehensive bioinformatics container environment for education and research
abstract
SUMMARY: The ORCA bioinformatics environment is a Docker image that contains hundreds of bioinformatics tools and their dependencies. The ORCA image and accompanying server infrastructure provide a comprehensive bioinformatics environment for education and research. The ORCA environment on a server is implemented using Docker containers, but without requiring users to interact directly with Docker, suitable for novices who may not yet have familiarity with managing containers. ORCA has been used successfully to provide a private bioinformatics environment to external collaborators at a large genome institute, for teaching an undergraduate class on bioinformatics targeted at biologists, and to provide a ready-to-go bioinformatics suite for a hackathon. Using ORCA eliminates time that would be spent debugging software installation issues, so that time may be better spent on education and research. AVAILABILITY AND IMPLEMENTATION: The ORCA Docker image is available at https://hub.docker.com/r/bcgsc/orca/. The source code of ORCA is available at https://github.com/bcgsc/orca under the MIT license.
Shaun D. Jackman, Tatyana Mozgacheva, Susie Chen, Brendan O'Huiginn, Lance Bailey, Inanç Birol, Steven J. M. Jones
Bioinform.6
2019 ntEdit: scalable genome sequence polishing
abstract
MOTIVATION: In the modern genomics era, genome sequence assemblies are routine practice. However, depending on the methodology, resulting drafts may contain considerable base errors. Although utilities exist for genome base polishing, they work best with high read coverage and do not scale well. We developed ntEdit, a Bloom filter-based genome sequence editing utility that scales to large mammalian and conifer genomes. RESULTS: We first tested ntEdit and the state-of-the-art assembly improvement tools GATK, Pilon and Racon on controlled Escherichia coli and Caenorhabditis elegans sequence data. Generally, ntEdit performs well at low sequence depths (<20×), fixing the majority (>97%) of base substitutions and indels, and its performance is largely constant with increased coverage. In all experiments conducted using a single CPU, the ntEdit pipeline executed in <14 s and <3 m, on average, on E.coli and C.elegans, respectively. We performed similar benchmarks on a sub-20× coverage human genome sequence dataset, inspecting accuracy and resource usage in editing chromosomes 1 and 21, and whole genome. ntEdit scaled linearly, executing in 30-40 m on those sequences. We show how ntEdit ran in <2 h 20 m to improve upon long and linked read human genome assemblies of NA12878, using high-coverage (54×) Illumina sequence data from the same individual, fixing frame shifts in coding sequences. We also generated 17-fold coverage spruce sequence data from haploid sequence sources (seed megagametophyte), and used it to edit our pseudo haploid assemblies of the 20 Gb interior and white spruce genomes in <4 and <5 h, respectively, making roughly 50M edits at a (substitution+indel) rate of 0.0024. AVAILABILITY AND IMPLEMENTATION: https://github.com/bcgsc/ntedit. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
René L. Warren, Lauren Coombe, Hamid Mohamadi, Barry Jaquish, Nathalie Isabel, Steven J. M. Jones, Jean Bousquet, Jörg Bohlmann, Inanç Birol
Bioinform.10
2018 ntPack: A Software Package for Big Data in Genomics
abstract
Establishing a fundamental understanding of the specific molecular biology of a species benefits substantially from reconstructing its genome (DNA) and transcriptome (RNA). These efforts are enabled by modern high throughput sequencing technologies. For over a decade, assembling the generated data into coherent information has been a primary focus of the bioinformatics field. However, the expanding data volume in the field and growing read lengths from evolving sequencing platforms require adapting bioinformatics tools to properly leverage the potential of new genomics technologies. This study is about efficient and scalable algorithms to perform a set of unit operations in genomics studies to guide sequence assembly. Here, we report on a software package, ntPack, with two components: ntHash, for nucleotide hashing, and ntCard for cardinality estimation. We characterize the statistical properties of these algorithms, and demonstrate their application on whole genome shotgun sequencing datasets describing the roundworm, human, and Canadian white spruce genomes. The software that implements these algorithms can be downloaded from our github repository at https://github.com/bcgsc.
Inanç Birol, Hamid Mohamadi, Justin Chu
BDCAT1
2018 ChopStitch: exon annotation and splice graph construction using transcriptome assembly and whole genome sequencing data
abstract
Motivation: Sequencing studies on non-model organisms often interrogate both genomes and transcriptomes with massive amounts of short sequences. Such studies require de novo analysis tools and techniques, when the species and closely related species lack high quality reference resources. For certain applications such as de novo annotation, information on putative exons and alternative splicing may be desirable. Results: Here we present ChopStitch, a new method for finding putative exons de novo and constructing splice graphs using an assembled transcriptome and whole genome shotgun sequencing (WGSS) data. ChopStitch identifies exon-exon boundaries in de novo assembled RNA-Seq data with the help of a Bloom filter that represents the k-mer spectrum of WGSS reads. The algorithm also accounts for base substitutions in transcript sequences that may be derived from sequencing or assembly errors, haplotype variations, or putative RNA editing events. The primary output of our tool is a FASTA file containing putative exons. Further, exon edges are interrogated for alternative exon-exon boundaries to detect transcript isoforms, which are represented as splice graphs in DOT output format. Availability and implementation: ChopStitch is written in Python and C++ and is released under the GPL license. It is freely available at https://github.com/bcgsc/ChopStitch. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Hamid Mohamadi, Benjamin P. Vandervalk, René L. Warren, Justin Chu, Inanç Birol
Bioinform.6
2018 ARCS: scaffolding genome drafts with linked reads
abstract
Motivation: Sequencing of human genomes is now routine, and assembly of shotgun reads is increasingly feasible. However, assemblies often fail to inform about chromosome-scale structure due to a lack of linkage information over long stretches of DNA-a shortcoming that is being addressed by new sequencing protocols, such as the GemCode and Chromium linked reads from 10 × Genomics. Results: Here, we present ARCS, an application that utilizes the barcoding information contained in linked reads to further organize draft genomes into highly contiguous assemblies. We show how the contiguity of an ABySS H.sapiens genome assembly can be increased over six-fold, using moderate coverage (25-fold) Chromium data. We expect ARCS to have broad utility in harnessing the barcoding information contained in linked read data for connecting high-quality sequences in genome assembly drafts. Availability and implementation: https://github.com/bcgsc/ARCS/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Sarah Yeo, Lauren Coombe, René L. Warren, Justin Chu, Inanç Birol
Bioinform.5
2018 ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers
abstract
BACKGROUND: The long-range sequencing information captured by linked reads, such as those available from 10× Genomics (10xG), helps resolve genome sequence repeats, and yields accurate and contiguous draft genome assemblies. We introduce ARKS, an alignment-free linked read genome scaffolding methodology that uses linked reads to organize genome assemblies further into contiguous drafts. Our approach departs from other read alignment-dependent linked read scaffolders, including our own (ARCS), and uses a kmer-based mapping approach. The kmer mapping strategy has several advantages over read alignment methods, including better usability and faster processing, as it precludes the need for input sequence formatting and draft sequence assembly indexing. The reliance on kmers instead of read alignments for pairing sequences relaxes the workflow requirements, and drastically reduces the run time. RESULTS: Here, we show how linked reads, when used in conjunction with Hi-C data for scaffolding, improve a draft human genome assembly of PacBio long-read data five-fold (baseline vs. ARKS NG50 = 4.6 vs. 23.1 Mbp, respectively). We also demonstrate how the method provides further improvements of a megabase-scale Supernova human genome assembly (NG50 = 14.74 Mbp vs. 25.94 Mbp before and after ARKS), which itself exclusively uses linked read data for assembly, with an execution speed six to nine times faster than competitive linked read scaffolders (~ 10.5 h compared to 75.7 h, on average). Following ARKS scaffolding of a human genome 10xG Supernova assembly (of cell line NA12878), fewer than 9 scaffolds cover each chromosome, except the largest (chromosome 1, n = 13). CONCLUSIONS: ARKS uses a kmer mapping strategy instead of linked read alignments to record and associate the barcode information needed to order and orient draft assembly sequences. The simplified workflow, when compared to that of our initial implementation, ARCS, markedly improves run time performances on experimental human genome datasets. Furthermore, the novel distance estimator in ARKS utilizes barcoding information from linked reads to estimate gap sizes. It accomplishes this by modeling the relationship between known distances of a region within contigs and calculating associated Jaccard indices. ARKS has the potential to provide correct, chromosome-scale genome assemblies, promptly. We expect ARKS to have broad utility in helping refine draft genomes.
Lauren Coombe, Benjamin P. Vandervalk, Justin Chu, Shaun D. Jackman, Inanç Birol, René L. Warren
BMC Bioinform.6
2018 Tigmint: correcting assembly errors using linked reads from large molecules
abstract
BACKGROUND: Genome sequencing yields the sequence of many short snippets of DNA (reads) from a genome. Genome assembly attempts to reconstruct the original genome from which these reads were derived. This task is difficult due to gaps and errors in the sequencing data, repetitive sequence in the underlying genome, and heterozygosity. As a result, assembly errors are common. In the absence of a reference genome, these misassemblies may be identified by comparing the sequencing data to the assembly and looking for discrepancies between the two. Once identified, these misassemblies may be corrected, improving the quality of the assembled sequence. Although tools exist to identify and correct misassemblies using Illumina paired-end and mate-pair sequencing, no such tool yet exists that makes use of the long distance information of the large molecules provided by linked reads, such as those offered by the 10x Genomics Chromium platform. We have developed the tool Tigmint to address this gap. RESULTS: To demonstrate the effectiveness of Tigmint, we applied it to assemblies of a human genome using short reads assembled with ABySS 2.0 and other assemblers. Tigmint reduced the number of misassemblies identified by QUAST in the ABySS assembly by 216 (27%). While scaffolding with ARCS alone more than doubled the scaffold NGA50 of the assembly from 3 to 8 Mbp, the combination of Tigmint and ARCS improved the scaffold NGA50 of the assembly over five-fold to 16.4 Mbp. This notable improvement in contiguity highlights the utility of assembly correction in refining assemblies. We demonstrate the utility of Tigmint in correcting the assemblies of multiple tools, as well as in using Chromium reads to correct and scaffold assemblies of long single-molecule sequencing. CONCLUSIONS: Scaffolding an assembly that has been corrected with Tigmint yields a final assembly that is both more correct and substantially more contiguous than an assembly that has not been corrected. Using single-molecule sequencing in combination with linked reads enables a genome sequence assembly that achieves both a high sequence contiguity as well as high scaffold contiguity, a feat not currently achievable with either technology alone.
Shaun D. Jackman, Lauren Coombe, Justin Chu, René L. Warren, Benjamin P. Vandervalk, Sarah Yeo, Zhuyi Xue, Hamid Mohamadi, Jörg Bohlmann, Steven J. M. Jones, Inanç Birol
BMC Bioinform.11
2017 Innovations and challenges in detecting long read overlaps: an evaluation of the state-of-the-art
abstract
Identifying overlaps between error-prone long reads, specifically those from Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PB), is essential for certain downstream applications, including error correction and de novo assembly. Though akin to the read-to-reference alignment problem, read-to-read overlap detection is a distinct problem that can benefit from specialized algorithms that perform efficiently and robustly on high error rate long reads. Here, we review the current state-of-the-art read-to-read overlap tools for error-prone long reads, including BLASR, DALIGNER, MHAP, GraphMap and Minimap. These specialized bioinformatics tools differ not just in their algorithmic designs and methodology, but also in their robustness of performance on a variety of datasets, time and memory efficiency and scalability. We highlight the algorithmic features of these tools, as well as their potential issues and biases when utilizing any particular method. To supplement our review of the algorithms, we benchmarked these tools, tracking their resource needs and computational performance, and assessed the specificity and precision of each. In the versions of the tools tested, we observed that Minimap is the most computationally efficient, specific and sensitive method on the ONT datasets tested; whereas GraphMap and DALIGNER are the most specific and sensitive methods on the tested PB datasets. The concepts surveyed may apply to future sequencing technologies, as scalability is becoming more relevant with increased sequencing throughput. CONTACT: [email protected] , [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Justin Chu, Hamid Mohamadi, René L. Warren, Inanç Birol
Bioinform.5
2017 Kollector: transcript-informed, targeted de novo assembly of gene loci
abstract
MOTIVATION: Despite considerable advancements in sequencing and computing technologies, de novo assembly of whole eukaryotic genomes is still a time-consuming task that requires a significant amount of computational resources and expertise. A targeted assembly approach to perform local assembly of sequences of interest remains a valuable option for some applications. This is especially true for gene-centric assemblies, whose resulting sequence can be readily utilized for more focused biological research. Here we describe Kollector, an alignment-free targeted assembly pipeline that uses thousands of transcript sequences concurrently to inform the localized assembly of corresponding gene loci. Kollector robustly reconstructs introns and novel sequences within these loci, and scales well to large genomes-properties that makes it especially useful for researchers working on non-model eukaryotic organisms. RESULTS: We demonstrate the performance of Kollector for assembling complete or near-complete Caenorhabditis elegans and Homo sapiens gene loci from their respective, input transcripts. In a time- and memory-efficient manner, the Kollector pipeline successfully reconstructs respectively 99% and 80% (compared to 86% and 73% with standard de novo assembly techniques) of C.elegans and H.sapiens transcript targets in their corresponding genomic space using whole genome shotgun sequencing reads. We also show that Kollector outperforms both established and recently released targeted assembly tools. Finally, we demonstrate three use cases for Kollector, including comparative and cancer genomics applications. AVAILABILITY AND IMPLEMENTATION: Kollector is implemented as a bash script, and is available at https://github.com/bcgsc/kollector. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erdi Kucuk, Justin Chu, Benjamin P. Vandervalk, S. Austin Hammond, René L. Warren, Inanç Birol
Bioinform.6
2017 Kollector: transcript-informed, targeted de novo assembly of gene loci
abstract
Bioinformatics (2017) 33(12), 1782–1788 The authors of the above article wish to inform readers that the following was omitted after Table 1: “The white spruce whole-genome shotgun sequence read data reported in Table 1 dataset no. 3 is available at the SRA under accession SRP041401 and Genome Annotation from the high-confidence genes are available at ftp://ftp.bcgsc.ca/supplementary/PG29_20140822/high_confidence_genes.fasta”. The paper has now been corrected online.
Erdi Kucuk, Justin Chu, Benjamin P. Vandervalk, S. Austin Hammond, René L. Warren, Inanç Birol
Bioinform.6
2017 ntCard: a streaming algorithm for cardinality estimation in genomics data
abstract
Motivation: Many bioinformatics algorithms are designed for the analysis of sequences of some uniform length, conventionally referred to as k -mers. These include de Bruijn graph assembly methods and sequence alignment tools. An efficient algorithm to enumerate the number of unique k -mers, or even better, to build a histogram of k -mer frequencies would be desirable for these tools and their downstream analysis pipelines. Among other applications, estimated frequencies can be used to predict genome sizes, measure sequencing error rates, and tune runtime parameters for analysis tools. However, calculating a k -mer histogram from large volumes of sequencing data is a challenging task. Results: Here, we present ntCard, a streaming algorithm for estimating the frequencies of k -mers in genomics datasets. At its core, ntCard uses the ntHash algorithm to efficiently compute hash values for streamed sequences. It then samples the calculated hash values to build a reduced representation multiplicity table describing the sample distribution. Finally, it uses a statistical model to reconstruct the population distribution from the sample distribution. We have compared the performance of ntCard and other cardinality estimation algorithms. We used three datasets of 480 GB, 500 GB and 2.4 TB in size, where the first two representing whole genome shotgun sequencing experiments on the human genome and the last one on the white spruce genome. Results show ntCard estimates k -mer coverage frequencies >15× faster than the state-of-the-art algorithms, using similar amount of memory, and with higher accuracy rates. Thus, our benchmarks demonstrate ntCard as a potentially enabling technology for large-scale genomics applications. Availability and Implementation: ntCard is written in C ++ and is released under the GPL license. It is freely available at https://github.com/bcgsc/ntCard. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Hamid Mohamadi, Inanç Birol
Bioinform.3
2016 Ensemble Minimum Sum of Squared Similarities sampling for Nyström-based spectral clustering
abstract
Spectral clustering is a powerful approach for clustering, with applications across multiple disciplines, including bioinformatics. However, the way its computational complexity scales limits its application in analyzing large datasets. This complexity can be reduced using the Nyström method, which subsamples the input data in a way that preserves its representational diversity. There are different established strategies for subsampling, yet they may have performance limitations for certain complex datasets. This paper we propose an alternative to those methods, introducing a new sampling procedure called Ensemble Minimum Sum of Squared Similarities (EMS3). We further improve on this method by using weight mixtures in subsample selection, yielding more accurate low-rank approximations than existing ensemble Nyström methods. We also provide a theoretical analysis of the upper error bound of the EMS3 algorithm, and demonstrate its performance in comparison to the leading spectral clustering methods that use Nyström sampling.
Djallel Bouneffouf 0001, Inanç Birol
IJCNN2
2016 Theoretical analysis of the Minimum Sum of Squared Similarities sampling for Nyström-based spectral clustering
abstract
Spectral clustering has shown a superior performance in analyzing the cluster structure. However, the exponentially computational complexity limits its application in analyzing large-scale data. To tackle this problem, many low-rank matrix approximating algorithms are proposed, of which the Nyström method is an approach with proved lower approximate errors. The algorithms commonly combine two powerful techniques in machine learning: spectral clustering algorithms and Nyström methods commonly used to obtain good quality low rank approximations of large matrices. This paper proposes to analyze a scalable Nyström-based clustering algorithm with a Minimum Sum of Squared Similarities (MSSS) sampling procedure. We provide theoretical analysis of the performance of the algorithm MSSS and demonstrate its theoretical performance in comparison to the leading spectral clustering methods that use Nyström sampling.
Djallel Bouneffouf 0001, Inanç Birol
IJCNN2
2016 ntHash: recursive nucleotide hashing
abstract
MOTIVATION: Hashing has been widely used for indexing, querying and rapid similarity search in many bioinformatics applications, including sequence alignment, genome and transcriptome assembly, k-mer counting and error correction. Hence, expediting hashing operations would have a substantial impact in the field, making bioinformatics applications faster and more efficient. RESULTS: We present ntHash, a hashing algorithm tuned for processing DNA/RNA sequences. It performs the best when calculating hash values for adjacent k-mers in an input sequence, operating an order of magnitude faster than the best performing alternatives in typical use cases. AVAILABILITY AND IMPLEMENTATION: ntHash is available online at http://www.bcgsc.ca/platform/bioinfo/software/nthash and is free for academic use. CONTACTS: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online.
Hamid Mohamadi, Justin Chu, Benjamin P. Vandervalk, Inanç Birol
Bioinform.4
2015 Sampling with Minimum Sum of Squared Similarities for Nystrom-Based Large Scale Spectral Clustering
Djallel Bouneffouf 0001, Inanç Birol
IJCAI2
2015 Sealer: a scalable gap-closing application for finishing draft genomes
abstract
BACKGROUND: While next-generation sequencing technologies have made sequencing genomes faster and more affordable, deciphering the complete genome sequence of an organism remains a significant bioinformatics challenge, especially for large genomes. Low sequence coverage, repetitive elements and short read length make de novo genome assembly difficult, often resulting in sequence and/or fragment "gaps" - uncharacterized nucleotide (N) stretches of unknown or estimated lengths. Some of these gaps can be closed by re-processing latent information in the raw reads. Even though there are several tools for closing gaps, they do not easily scale up to processing billion base pair genomes. RESULTS: Here we describe Sealer, a tool designed to close gaps within assembly scaffolds by navigating de Bruijn graphs represented by space-efficient Bloom filter data structures. We demonstrate how it scales to successfully close 50.8% and 13.8% of gaps in human (3 Gbp) and white spruce (20 Gbp) draft assemblies in under 30 and 27 h, respectively - a feat that is not possible with other leading tools with the breadth of data used in our study. CONCLUSION: Sealer is an automated finishing application that uses the succinct Bloom filter representation of a de Bruijn graph to close gaps in draft assemblies, including that of very large genomes. We expect Sealer to have broad utility for finishing genomes across the tree of life, from bacterial genomes to large plant genomes and beyond. Sealer is available for download at https://github.com/bcgsc/abyss/tree/sealer-release.
Daniel Paulino, René L. Warren, Benjamin P. Vandervalk, Anthony Raymond, Shaun D. Jackman, Inanç Birol
BMC Bioinform.6
2014 Spaced seed data structures
abstract
This past decade, genome sciences have benefitted from rapid advances in DNA sequencing technologies, and development of efficient algorithms for processing short nucleotide sequences played a key role in enabling their uptake in the field. In particular, reassembly of human genomes (de novo or reference-guided) from short DNA sequence reads had a substantial impact on health research. De novo assembly of a genome is essential in the absence of a reference genome sequence of a species. It is also gaining traction even when one is available, due to the utility of the method to resolve ambiguous or rearranged genomic regions with high specificity. With commercial high-throughput sequencing technologies increasing their throughput and their read lengths, the de Bruijn graph (DBG) paradigm used by many assembly algorithms needs to be revisited. DBG uses a table of k-mers, sequences of length k base pairs derived from the reads, and their k-1 base pair overlaps to assemble sequences. Despite longer k-mers unlocking longer genomic features for assembly, associated increases in memory usage and other compute resources are tradeoffs that limit the practicability of DBG over other assembly archetypes already designed for longer reads. Here, we introduce three data structure designs for paired k-mers, or spaced seeds, each addressing memory and run time constraints imposed by longer reads. In spaced seeds, a fixed distance separates k-mer pairs, providing increased sequence specificity with increased distance, while keeping memory usage low. Further, we describe a data structure based on Bloom filters that would be suitable to implicitly store spaced seeds, and would be tolerant to sequencing errors. Building on the spaced seeds Bloom filter, we describe a data structure for tracking the frequencies of observed spaced seeds. We expect the data structure designs we introduce in this study to have broad applications in genomics research, with niche applications in genome, transcriptome and metagenome assemblies, and in read error correction.
Inanç Birol, Hamid Mohamadi, Anthony Raymond, Karthika Raghavan, Justin Chu, Benjamin P. Vandervalk, Shaun D. Jackman, René L. Warren
BIBM1
2014 Konnector: Connecting paired-end reads using a bloom filter de Bruijn graph
abstract
Paired-end sequencing yields a read from each end of a DNA fragment, typically leaving a gap of unsequenced nucleotides in the middle. Closing this gap using information from other reads in the same sequencing experiment offers the potential to generate longer “pseudo-reads” using short read sequencing platforms. Such long reads may benefit downstream applications such as de novo sequence assembly, gap filling, and variant detection. With these possible applications in mind, we have developed Konnector, a software tool to fill in the nucleotides of the sequence gap between read pairs by navigating a de Bruijn graph. Konnector represents the de Bruijn graph using a Bloom filter, a probabilistic and memory-efficient data structure. Our implementation is able to store the de Bruijn graph using a mean 1.5 bytes of memory per k-mer, which represents a marked improvement over the typical hash table data structure. The memory usage per k-mer is independent of the k-mer length, enabling application of the tool to large genomes. We report the performance of the tool on simulated and experimental datasets, and discuss its utility for downstream analysis. Availability-Konnector is open-source software, free for academic use, released under the British Columbia Cancer Agency's academic license. The tool is included with ABySS version 1.5.2 and later, and is available for download from http://www.bcgsc.ca/platform/bioinfo/software/abyss.
Benjamin P. Vandervalk, Shaun D. Jackman, Anthony Raymond, Hamid Mohamadi, Dean A. Attali, Justin Chu, René L. Warren, Inanç Birol
BIBM9
2014 BioBloom tools: fast, accurate and memory-efficient host species sequence screening using bloom filters
abstract
Abstract Large datasets can be screened for sequences from a specific organism, quickly and with low memory requirements, by a data structure that supports time- and memory-efficient set membership queries. Bloom filters offer such queries but require that false positives be controlled. We present BioBloom Tools, a Bloom filter-based sequence-screening tool that is faster than BWA, Bowtie 2 (popular alignment algorithms) and FACS (a membership query algorithm). It delivers accuracies comparable with these tools, controls false positives and has low memory requirements. Availability and implementaion: www.bcgsc.ca/platform/bioinfo/software/biobloomtools Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Justin Chu, Sara Sadeghi, Anthony Raymond, Shaun D. Jackman, Ka Ming Nip, Richard Mar, Hamid Mohamadi, Yaron S. N. Butterfield, Gordon Robertson, Inanç Birol
Bioinform.10
2013 Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data
abstract
UNLABELLED: White spruce (Picea glauca) is a dominant conifer of the boreal forests of North America, and providing genomics resources for this commercially valuable tree will help improve forest management and conservation efforts. Sequencing and assembling the large and highly repetitive spruce genome though pushes the boundaries of the current technology. Here, we describe a whole-genome shotgun sequencing strategy using two Illumina sequencing platforms and an assembly approach using the ABySS software. We report a 20.8 giga base pairs draft genome in 4.9 million scaffolds, with a scaffold N50 of 20,356 bp. We demonstrate how recent improvements in the sequencing technology, especially increasing read lengths and paired end reads from longer fragments have a major impact on the assembly contiguity. We also note that scalable bioinformatics tools are instrumental in providing rapid draft assemblies. AVAILABILITY: The Picea glauca genome sequencing and assembly data are available through NCBI (Accession#: ALWZ0100000000 PID: PRJNA83435). http://www.ncbi.nlm.nih.gov/bioproject/83435.
Inanç Birol, Anthony Raymond, Shaun D. Jackman, Stephen Pleasance, Robin Coope, Greg A. Taylor, Macaire Man Saint Yuen, Christopher I. Keeling, Dana Brand, Benjamin P. Vandervalk, Heather Kirk, Pawan Pandoh, Richard A. Moore, Yongjun Zhao 0002, Andrew J. Mungall, Barry Jaquish, Alvin Yanchuk, Carol Ritland, Brian Boyle, Jean Bousquet, Kermit Ritland, John MacKay, Jörg Bohlmann, Steven J. M. Jones
Bioinform.1
2013 Identifying cancer mutation targets across thousands of samples: MuteProc, a high throughput mutation analysis pipeline
abstract
BACKGROUND: In the past decade, bioinformatics tools have matured enough to reliably perform sophisticated primary data analysis on Next Generation Sequencing (NGS) data, such as mapping, assemblies and variant calling, however, there is still a dire need for improvements in the higher level analysis such as NGS data organization, analysis of mutation patterns and Genome Wide Association Studies (GWAS). RESULTS: We present a high throughput pipeline for identifying cancer mutation targets, capable of processing billions of variations across thousands of samples. This pipeline is coupled with our Human Variation Database to provide more complex down stream analysis on the variations hosted in the database. Most notably, these analysis include finding significantly mutated regions across multiple genomes and regions with mutational preferences within certain types of cancers. The results of the analysis is presented in HTML summary reports that incorporate gene annotations from various resources for the reported regions. CONCLUSION: MuteProc is available for download through the Vancouver Short Read Analysis Package on Sourceforge: http://vancouvershortr.sourceforge.net. Instructions for use and a tutorial are provided on the accompanying wiki pages at https://sourceforge.net/apps/mediawiki/vancouvershortr/index.php?title=Pipeline_introduction.
Alireza Hadj Khodabakhshi, Anthony P. Fejes, Inanç Birol, Steven J. M. Jones
BMC Bioinform.3
2012 Hive plots - rational approach to visualizing networks
abstract
Networks are typically visualized with force-based or spectral layouts. These algorithms lack reproducibility and perceptual uniformity because they do not use a node coordinate system. The layouts can be difficult to interpret and are unsuitable for assessing differences in networks. To address these issues, we introduce hive plots (http://www.hiveplot.com) for generating informative, quantitative and comparable network layouts. Hive plots depict network structure transparently, are simple to understand and can be easily tuned to identify patterns of interest. The method is computationally straightforward, scales well and is amenable to a plugin for existing tools.
Martin Krzywinski, Inanç Birol, Steven J. M. Jones, Marco A. Marra
Briefings Bioinform.2
2012 Dissect: detection and characterization of novel structural alterations in transcribed sequences
abstract
MOTIVATION: Computational identification of genomic structural variants via high-throughput sequencing is an important problem for which a number of highly sophisticated solutions have been recently developed. With the advent of high-throughput transcriptome sequencing (RNA-Seq), the problem of identifying structural alterations in the transcriptome is now attracting significant attention. In this article, we introduce two novel algorithmic formulations for identifying transcriptomic structural variants through aligning transcripts to the reference genome under the consideration of such variation. The first formulation is based on a nucleotide-level alignment model; a second, potentially faster formulation is based on chaining fragments shared between each transcript and the reference genome. Based on these formulations, we introduce a novel transcriptome-to-genome alignment tool, Dissect (DIScovery of Structural Alteration Event Containing Transcripts), which can identify and characterize transcriptomic events such as duplications, inversions, rearrangements and fusions. Dissect is suitable for whole transcriptome structural variation discovery problems involving sufficiently long reads or accurately assembled contigs. RESULTS: We tested Dissect on simulated transcripts altered via structural events, as well as assembled RNA-Seq contigs from human prostate cancer cell line C4-2. Our results indicate that Dissect has high sensitivity and specificity in identifying structural alteration events in simulated transcripts as well as uncovering novel structural alterations in cancer transcriptomes. AVAILABILITY: Dissect is available for public use at: http://dissect-trans.sourceforge.net.
Deniz Yörükoglu, Faraz Hach, Lucas Swanson, Colin C. Collins, Inanç Birol, Süleyman Cenk Sahinalp
Bioinform.5
2011 Human variation database: an open-source database template for genomic discovery
abstract
MOTIVATION: Current public variation databases are based upon collaboratively pooling data into a single database with a single interface available to the public. This gives little control to the collaborator to mine the database and requires that they freely share their data with the owners of the repository. We aim to provide an alternative mechanism: providing the source code and application programming interface (API) of a database, enabling researchers to set up local versions without investing heavily in the development of the resource and allowing for confidential information to remain secure. RESULTS: We describe an open-source database that can be installed easily at any research facility for the storage and analysis of thousands of next-generation sequencing variations. This database is built using PostgreSQL 8.4 (The PostgreSQL Global Development Group. postgres 8.4: http://www.postgresql.org) and provides a novel method for collating and searching across the reported results from thousands of next-generation sequence samples, as well as rapidly accessing vital information on the origin of the samples. The schema of the database makes rapid and insightful queries simple and enables easy annotation of novel or known genetic variations. A modular and cross-platform Java API is provided to perform common functions, such as generation of standard experimental reports and graphical summaries of modifications to genes. Included libraries allow adopters of the database to quickly develop their own queries. AVAILABILITY: The software is available for download through the Vancouver Short Read Analysis Package on Sourceforge, http://vancouvershortr.sourceforge.net. Instructions for use and deployment are provided on the accompanying wiki pages. CONTACT: [email protected].
Anthony P. Fejes, Alireza Hadj Khodabakhshi, Inanç Birol, Steven J. M. Jones
Bioinform.3
2010 Detection and characterization of novel sequence insertions using paired-end next-generation sequencing
abstract
MOTIVATION: In the past few years, human genome structural variation discovery has enjoyed increased attention from the genomics research community. Many studies were published to characterize short insertions, deletions, duplications and inversions, and associate copy number variants (CNVs) with disease. Detection of new sequence insertions requires sequence data, however, the 'detectable' sequence length with read-pair analysis is limited by the insert size. Thus, longer sequence insertions that contribute to our genetic makeup are not extensively researched. RESULTS: We present NovelSeq: a computational framework to discover the content and location of long novel sequence insertions using paired-end sequencing data generated by the next-generation sequencing platforms. Our framework can be built as part of a general sequence analysis pipeline to discover multiple types of genetic variation (SNPs, structural variation, etc.), thus it requires significantly less-computational resources than de novo sequence assembly. We apply our methods to detect novel sequence insertions in the genome of an anonymous donor and validate our results by comparing with the insertions discovered in the same genome using various sources of sequence data. AVAILABILITY: The implementation of the NovelSeq pipeline is available at http://compbio.cs.sfu.ca/strvar.htm CONTACT: [email protected]; [email protected]
Iman Hajirasouliha, Fereydoun Hormozdiari, Can Alkan, Jeffrey M. Kidd, Inanç Birol, Evan E. Eichler, Süleyman Cenk Sahinalp
Bioinform.5
2010 Genomic analysis of a rare human tumor
abstract
The introduction of next-generation DNA sequencing devices into the field of oncology provides an unprecedented mechanism to determine the underlying genetic changes that have occurred within a tumor and also the changes that accrue during treatment. An enhanced understanding of the oncogenic mechanisms could have an immediate clinical role in the treatment of rare tumors - where treatment protocols do not exist and their rarity would indicate that clinical trials would be unlikely to be undertaken for their establishment. We have investigated the utility of massively parallel sequencing to characterize a rare adenocarcinoma of the tongue, before and after treatment. In the pre-treatment tumor we identified 7,629 genes within regions of copy number gain, 1,078 genes exhibited increased expression relative to the blood and unrelated tumors and four genes contained somatic protein-coding mutations. Our analysis suggested the tumor cells were driven by the RET oncogene and its other pathway constituents. Genes whose protein products are targeted by the RET inhibitors sunitinib and sorafenib correlated with being amplified and or highly expressed. Consistent with our observations subsequent administration of sunitinib was associated with stable disease lasting 4 months, after which the lung lesions began to grow. Administration of sorafenib and sulindac provided disease stabilization for an additional 3 months after which the cancer progressed and new lesions appeared. A metastasis recurring in the skin was determined to possess 7,288 genes within copy number amplicons, 385 genes exhibiting increased expression relative to other tumours and 9 new somatic protein coding mutations. The observed mutations and amplifications were found to be consistent with resistance to therapy arising through further activation of RET pathway and nascent activation of the AKT pathway. Our results provide evidence for the clinical utility of complete genomic characterization and direct in-vivo genome-wide characterization of the mutations accruing within a tumor under drug selection.
Steven J. M. Jones, Janessa Laskin, Yvonne Y. Li, Obi L. Griffith, Jianghong An, Mikhail Bilenky, Yaron S. N. Butterfield, Timothee Cezard, Eric Chuah, Richard Corbett, Anthony P. Fejes, Malachi Griffith, John Yee, Montgomery Martin, Michael Mayo, Nataliya Melnyk, Ryan D. Morin, Trevor J. Pugh, Tesa Severson, Sohrab P. Shah, Margaret Sutcliffe, Angela Tam, Jefferson Terry, Nina Thiessen, Thomas Thomson, Richard Varhol, Thomas Zeng 0002, Yongjun Zhao 0002, Richard A. Moore, David G. Huntsman, Inanç Birol, Martin Hirst, Robert A. Holt, Marco A. Marra
BMC Bioinform.31
2010 LaneRuler: Automated Lane Tracking for DNA Electrophoresis Gel Images
abstract
We present a novel method for correctly identifying and straightening one dimensional agarose electrophoretic lanes. Our method has been shown to yield comparable accuracy with manual lane tracking results, and to successfully process 98% of DNA fingerprinting gels with no human intervention.
R. T. F. Wong, Stephane Flibotte, Richard Corbett, Parvaneh Saeedi, Steven J. M. Jones, Marco A. Marra, Jacqueline E. Schein, Inanç Birol
IEEE Trans Autom. Sci. Eng.8
2009 De novo transcriptome assembly with ABySS
abstract
MOTIVATION: Whole transcriptome shotgun sequencing data from non-normalized samples offer unique opportunities to study the metabolic states of organisms. One can deduce gene expression levels using sequence coverage as a surrogate, identify coding changes or discover novel isoforms or transcripts. Especially for discovery of novel events, de novo assembly of transcriptomes is desirable. RESULTS: Transcriptome from tumor tissue of a patient with follicular lymphoma was sequenced with 36 base pair (bp) single- and paired-end reads on the Illumina Genome Analyzer II platform. We assembled approximately 194 million reads using ABySS into 66 921 contigs 100 bp or longer, with a maximum contig length of 10 951 bp, representing over 30 million base pairs of unique transcriptome sequence, or roughly 1% of the genome. AVAILABILITY AND IMPLEMENTATION: Source code and binaries of ABySS are freely available for download at http://www.bcgsc.ca/platform/bioinfo/software/abyss. Assembler tool is implemented in C++. The parallel version uses Open MPI. ABySS-Explorer tool is implemented in Java using the Java universal network/graph framework. CONTACT: [email protected].
Inanç Birol, Shaun D. Jackman, Cydney B. Nielsen, Jenny Q. Qian, Richard Varhol, Greg Stazyk, Ryan D. Morin, Yongjun Zhao 0002, Martin Hirst, Jacqueline E. Schein, Douglas E. Horsman, Joseph M. Connors, Randy D. Gascoyne, Marco A. Marra, Steven J. M. Jones
Bioinform.1
2009 ABySS-Explorer: Visualizing Genome Sequence Assemblies
abstract
One bottleneck in large-scale genome sequencing projects is reconstructing the full genome sequence from the short subsequences produced by current technologies. The final stages of the genome assembly process inevitably require manual inspection of data inconsistencies and could be greatly aided by visualization. This paper presents our design decisions in translating key data features identified through discussions with analysts into a concise visual encoding. Current visualization tools in this domain focus on local sequence errors making high-level inspection of the assembly difficult if not impossible. We present a novel interactive graph display, ABySS-Explorer, that emphasizes the global assembly structure while also integrating salient data features such as sequence length. Our tool replaces manual and in some cases pen-and-paper based analysis tasks, and we discuss how user feedback was incorporated into iterative design refinements. Finally, we touch on applications of this representation not initially considered in our design phase, suggesting the generality of this encoding for DNA sequence data.
Cydney B. Nielsen, Shaun D. Jackman, Inanç Birol, Steven J. M. Jones
IEEE Trans. Vis. Comput. Graph.3
2008 Optimal pooling for genome re-sequencing with ultra-high-throughput short-read technologies
abstract
New generation sequencing technologies offer unique opportunities and challenges for re-sequencing studies. In this article, we focus on re-sequencing experiments using the Solexa technology, based on bacterial artificial chromosome (BAC) clones, and address an experimental design problem. In these specific experiments, approximate coordinates of the BACs on a reference genome are known, and fine-scale differences between the BAC sequences and the reference are of interest. The high-throughput characteristics of the sequencing technology makes it possible to multiplex BAC sequencing experiments by pooling BACs for a cost-effective operation. However, the way BACs are pooled in such re-sequencing experiments has an effect on the downstream analysis of the generated data, mostly due to subsequences common to multiple BACs. The experimental design strategy we develop in this article offers combinatorial solutions based on approximation algorithms for the well-known max n-cut problem and the related max n-section problem on hypergraphs. Our algorithms, when applied to a number of sample cases give more than a 2-fold performance improvement over random partitioning.
Iman Hajirasouliha, Fereydoun Hormozdiari, Süleyman Cenk Sahinalp, Inanç Birol
ISMB4