EDBT 2026 Demo / reviewers in the wild / expert
Steven J. M. Jones
dblp:02/2455
· DBLP profile ↗
43ranked-venue papers
1as first author
2since 2021 · last 2024
0000-0003-3394-2208ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
26 papers |
Bioinformatics and computational biology · 99% Computing education · 1% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 100% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.8 | 6 | 2019 | ntEdit: scalable genome sequence polishing · Bioinform. 2019 Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data · Bioinform. 2013 ABySS-Explorer: Visualizing Genome Sequence Assemblies · IEEE Trans. Vis. Comput. Graph. 2009 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.5 | 3 | 2016 | ChAsE: chromatin analysis and exploration tool · Bioinform. 2016 ALEA: a toolbox for allele-specific epigenomics analysis · Bioinform. 2014 FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology · Bioinform. 2008 |
Bioinformatics and computational biology
epigenomics |
0.4 | 2 | 2016 | ChAsE: chromatin analysis and exploration tool · Bioinform. 2016 ALEA: a toolbox for allele-specific epigenomics analysis · Bioinform. 2014 |
Bioinformatics and computational biology
gene regulation |
0.4 | 2 | 2019 | RTNduals: an R/Bioconductor package for analysis of co-regulation and inference of dual regulons · Bioinform. 2019 ORegAnno: an open access database and curation system for literature-derived promoters, transcription factor binding sites and regulatory variation · Bioinform. 2006 |
Bioinformatics and computational biology
bioinformatics infrastructure |
0.4 | 1 | 2019 | ORCA: a comprehensive bioinformatics container environment for education and research · Bioinform. 2019 |
Bioinformatics and computational biology
cancer genomics |
0.4 | 1 | 2019 | RTNsurvival: an R/Bioconductor package for regulatory network survival analysis · Bioinform. 2019 |
Bioinformatics and computational biology › gene regulation
gene regulatory network |
0.4 | 1 | 2019 | RTNduals: an R/Bioconductor package for analysis of co-regulation and inference of dual regulons · Bioinform. 2019 |
Bioinformatics and computational biology › gene regulation › gene regulatory network
gene regulatory network analysis |
0.4 | 1 | 2019 | RTNsurvival: an R/Bioconductor package for regulatory network survival analysis · Bioinform. 2019 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome polishing |
0.4 | 1 | 2019 | ntEdit: scalable genome sequence polishing · Bioinform. 2019 |
Bioinformatics and computational biology › genomics
genomic variant analysis |
0.4 | 1 | 2019 | MAVIS: merging, annotation, validation, and illustration of structural variants · Bioinform. 2019 |
Bioinformatics and computational biology › genomics › structural variation
structural variant detection |
0.4 | 1 | 2019 | MAVIS: merging, annotation, validation, and illustration of structural variants · Bioinform. 2019 |
Bioinformatics and computational biology › gene regulation › transcription factor analysis
transcription factor co-regulation |
0.4 | 1 | 2019 | RTNduals: an R/Bioconductor package for analysis of co-regulation and inference of dual regulons · Bioinform. 2019 |
Bioinformatics and computational biology
gene expression analysis |
0.3 | 4 | 2019 | On the Deep Order-Preserving Submatrix Problem: A Best Effort Approach · IEEE Trans. Knowl. Data Eng. 2012 RTNsurvival: an R/Bioconductor package for regulatory network survival analysis · Bioinform. 2019 A methodology for analyzing SAGE libraries for cancer profiling · ACM Trans. Inf. Syst. 2005 |
Bioinformatics and computational biology › biomedical text mining
literature-based discovery |
0.3 | 1 | 2018 | A collaborative filtering-based approach to biomedical knowledge discovery · Bioinform. 2018 |
Bioinformatics and computational biology
sequence alignment |
0.3 | 3 | 2010 | High quality SNP calling using Illumina data at shallow coverage · Bioinform. 2010 Slider - maximum use of probability information for alignment of short sequence reads and SNP detection · Bioinform. 2009 Assembling millions of short DNA sequences using SSAKE · Bioinform. 2007 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
de novo assembly |
0.2 | 2 | 2013 | Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data · Bioinform. 2013 Assembling millions of short DNA sequences using SSAKE · Bioinform. 2007 |
Data mining
pattern mining |
0.2 | 2 | 2012 | On the Deep Order-Preserving Submatrix Problem: A Best Effort Approach · IEEE Trans. Knowl. Data Eng. 2012 Discovering significant OPSM subspace clusters in massive gene expression data · KDD 2006 |
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment |
0.2 | 2 | 2010 | High quality SNP calling using Illumina data at shallow coverage · Bioinform. 2010 Slider - maximum use of probability information for alignment of short sequence reads and SNP detection · Bioinform. 2009 |
Bioinformatics and computational biology › genomics › computational genomics
SNP detection |
0.2 | 2 | 2010 | High quality SNP calling using Illumina data at shallow coverage · Bioinform. 2010 Slider - maximum use of probability information for alignment of short sequence reads and SNP detection · Bioinform. 2009 |
Bioinformatics and computational biology › statistical genetics
allele-specific analysis |
0.2 | 1 | 2014 | ALEA: a toolbox for allele-specific epigenomics analysis · Bioinform. 2014 |
Visualization and visual analytics
biological data visualization |
0.2 | 2 | 2016 | ABySS-Explorer: Visualizing Genome Sequence Assemblies · IEEE Trans. Vis. Comput. Graph. 2009 ChAsE: chromatin analysis and exploration tool · Bioinform. 2016 |
Bioinformatics and computational biology
sequence analysis |
0.2 | 2 | 2009 | De novo transcriptome assembly with ABySS · Bioinform. 2009 Assembling millions of short DNA sequences using SSAKE · Bioinform. 2007 |
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
short read assembly |
0.2 | 2 | 2009 | De novo transcriptome assembly with ABySS · Bioinform. 2009 Assembling millions of short DNA sequences using SSAKE · Bioinform. 2007 |
Bioinformatics and computational biology › genomics › genome sequencing
whole genome shotgun sequencing |
0.2 | 1 | 2013 | Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing data · Bioinform. 2013 |
Bioinformatics and computational biology › gene expression analysis
biclustering |
0.1 | 1 | 2012 | On the Deep Order-Preserving Submatrix Problem: A Best Effort Approach · IEEE Trans. Knowl. Data Eng. 2012 |
Bioinformatics and computational biology › cancer genomics
gene fusion detection |
0.1 | 1 | 2012 | BreakFusion: targeted assembly-based identification of gene fusions in whole transcriptome paired-end sequencing data · Bioinform. 2012 |
Bioinformatics and computational biology › transcriptomics
transcriptome sequencing |
0.1 | 1 | 2012 | BreakFusion: targeted assembly-based identification of gene fusions in whole transcriptome paired-end sequencing data · Bioinform. 2012 |
Bioinformatics and computational biology › biological database
genetic variation database |
0.1 | 1 | 2011 | Human variation database: an open-source database template for genomic discovery · Bioinform. 2011 |
Bioinformatics and computational biology
genomics |
0.1 | 2 | 2008 | FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology · Bioinform. 2008 IslandPath: aiding detection of genomic islands in prokaryotes · Bioinform. 2003 |
Computing education › STEM education
bioinformatics education |
0.1 | 1 | 2019 | ORCA: a comprehensive bioinformatics container environment for education and research · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
k-means clustering · 0.5interactive visualization · 0.5statistical significance testing · 0.4kaplan-meier analysis · 0.4gene set enrichment analysis · 0.4docker containerization · 0.4cox proportional hazards regression · 0.4bloom filter · 0.4singular value decomposition · 0.3collaborative filtering · 0.3resource-bounded mining · 0.1best-effort search space pruning · 0.1interactive graph display · 0.1sequential pattern mining · 0.1wilcoxon rank sum test · 0.1subspace selection · 0.1normalization · 0.1imputation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing clinical genomic accuracy with panelGC: a novel metric and tool for quantifying and monitoring GC biases in hybridization capture panel sequencingabstractAccurate assessment of fragment abundance within a genome is crucial in clinical genomics applications such as the analysis of copy number variation (CNV). However, this task is often hindered by biased coverage in regions with varying guanine-cytosine (GC) content. These biases are particularly exacerbated in hybridization capture sequencing due to GC effects on probe hybridization and polymerase chain reaction (PCR) amplification efficiency. Such GC content-associated variations can exert a negative impact on the fidelity of CNV calling within hybridization capture panels. In this report, we present panelGC, a novel metric, to quantify and monitor GC biases in hybridization capture sequencing data. We establish the efficacy of panelGC, demonstrating its proficiency in identifying and flagging potential procedural anomalies, even in situations where instrument and experimental monitoring data may not be readily accessible. Validation using real-world datasets demonstrates that panelGC enhances the quality control and reliability of hybridization capture panel sequencing. Xuanjin Cheng, Murathan T. Goktas, Laura M. Williamson, Martin Krzywinski, David T. Mulder, Lucas Swanson, Jill Slind, Jelena Sihvonen, Cynthia R. Chow, Amy Carr, Ian Bosdet, Tracy Tucker, Sean Young, Richard A. Moore, Karen L. Mungall, Stephen Yip, Steven J. M. Jones |
Briefings Bioinform. | 17 |
| 2022 | cSurvival: a web resource for biomarker interactions in cancer outcomes and in cell linesabstractSurvival analysis is a technique for identifying prognostic biomarkers and genetic vulnerabilities in cancer studies. Large-scale consortium-based projects have profiled >11 000 adult and >4000 pediatric tumor cases with clinical outcomes and multiomics approaches. This provides a resource for investigating molecular-level cancer etiologies using clinical correlations. Although cancers often arise from multiple genetic vulnerabilities and have deregulated gene sets (GSs), existing survival analysis protocols can report only on individual genes. Additionally, there is no systematic method to connect clinical outcomes with experimental (cell line) data. To address these gaps, we developed cSurvival (https://tau.cmmt.ubc.ca/cSurvival). cSurvival provides a user-adjustable analytical pipeline with a curated, integrated database and offers three main advances: (i) joint analysis with two genomic predictors to identify interacting biomarkers, including new algorithms to identify optimal cutoffs for two continuous predictors; (ii) survival analysis not only at the gene, but also the GS level; and (iii) integration of clinical and experimental cell line studies to generate synergistic biological insights. To demonstrate these advances, we report three case studies. We confirmed findings of autophagy-dependent survival in colorectal cancers and of synergistic negative effects between high expression of SLC7A11 and SLC2A1 on outcomes in several cancers. We further used cSurvival to identify high expression of the Nrf2-antioxidant response element pathway as a main indicator for lung cancer prognosis and for cellular resistance to oxidative stress-inducing drugs. Altogether, these analyses demonstrate cSurvival's ability to support biomarker prognosis and interaction analysis via gene- and GS-level approaches and to integrate clinical and experimental biomedical studies. Xuanjin Cheng, Yongxing Liu, Andrew Gordon Robertson, Xuekui Zhang, Steven J. M. Jones, Stefan Taubert |
Briefings Bioinform. | 7 |
| 2019 | RTNduals: an R/Bioconductor package for analysis of co-regulation and inference of dual regulonsabstractMOTIVATION: Transcription factors (TFs) are key regulators of gene expression, and can activate or repress multiple target genes, forming regulatory units, or regulons. Understanding downstream effects of these regulators includes evaluating how TFs cooperate or compete within regulatory networks. Here we present RTNduals, an R/Bioconductor package that implements a general method for analyzing pairs of regulons. RESULTS: RTNduals identifies a dual regulon when the number of targets shared between a pair of regulators is statistically significant. The package extends the RTN (Reconstruction of Transcriptional Networks) package, and uses RTN transcriptional networks to identify significant co-regulatory associations between regulons. The Supplementary Information reports two case studies for TFs using the METABRIC and TCGA breast cancer cohorts. AVAILABILITY AND IMPLEMENTATION: RTNduals is written in the R language, and is available from the Bioconductor project at http://bioconductor.org/packages/RTNduals/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vinicius S. Chagas, Clarice S. Groeneveld, Kelin G. Oliveira, Sheyla Trefflich, Rodrigo C. de Almeida, Bruce A. J. Ponder, Kerstin B. Meyer, Steven J. M. Jones, Gordon Robertson, Mauro A. A. Castro |
Bioinform. | 8 |
| 2019 | RTNsurvival: an R/Bioconductor package for regulatory network survival analysisabstractMOTIVATION: Transcriptional networks are models that allow the biological state of cells or tumours to be described. Such networks consist of connected regulatory units known as regulons, each comprised of a regulator and its targets. Inferring a transcriptional network can be a helpful initial step in characterizing the different phenotypes within a cohort. While the network itself provides no information on molecular differences between samples, the per-sample state of each regulon, i.e. the regulon activity, can be used for describing subtypes in a cohort. Integrating regulon activities with clinical data and outcomes would extend this characterization of differences between subtypes. RESULTS: We describe RTNsurvival, an R/Bioconductor package that calculates regulon activity profiles using transcriptional networks reconstructed by the RTN package, gene expression data, and a two-tailed Gene Set Enrichment Analysis. Given regulon activity profiles across a cohort, RTNsurvival can perform Kaplan-Meier analyses and Cox Proportional Hazards regressions, while also considering confounding variables. The Supplementary Information provides two case studies that use data from breast and liver cancer cohorts and features uni- and multivariate regulon survival analysis. AVAILABILITY AND IMPLEMENTATION: RTNsurvival is written in the R language, and is available from the Bioconductor project at http://bioconductor.org/packages/RTNsurvival/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Clarice S. Groeneveld, Vinicius S. Chagas, Steven J. M. Jones, Gordon Robertson, Bruce A. J. Ponder, Kerstin B. Meyer, Mauro A. A. Castro |
Bioinform. | 3 |
| 2019 | ORCA: a comprehensive bioinformatics container environment for education and researchabstractSUMMARY: The ORCA bioinformatics environment is a Docker image that contains hundreds of bioinformatics tools and their dependencies. The ORCA image and accompanying server infrastructure provide a comprehensive bioinformatics environment for education and research. The ORCA environment on a server is implemented using Docker containers, but without requiring users to interact directly with Docker, suitable for novices who may not yet have familiarity with managing containers. ORCA has been used successfully to provide a private bioinformatics environment to external collaborators at a large genome institute, for teaching an undergraduate class on bioinformatics targeted at biologists, and to provide a ready-to-go bioinformatics suite for a hackathon. Using ORCA eliminates time that would be spent debugging software installation issues, so that time may be better spent on education and research. AVAILABILITY AND IMPLEMENTATION: The ORCA Docker image is available at https://hub.docker.com/r/bcgsc/orca/. The source code of ORCA is available at https://github.com/bcgsc/orca under the MIT license. Shaun D. Jackman, Tatyana Mozgacheva, Susie Chen, Brendan O'Huiginn, Lance Bailey, Inanç Birol, Steven J. M. Jones |
Bioinform. | 7 |
| 2019 | MAVIS: merging, annotation, validation, and illustration of structural variantsabstractSummary: Reliably identifying genomic rearrangements and interpreting their impact is a key step in understanding their role in human cancers and inherited genetic diseases. Many short read algorithmic approaches exist but all have appreciable false negative rates. A common approach is to evaluate the union of multiple tools increasing sensitivity, followed by filtering to retain specificity. Here we describe an application framework for the rapid generation of structural variant consensus, unique in its ability to visualize the genetic impact and context as well as process both genome and transcriptome data. Availability and implementation: http://mavis.bcgsc.ca. Supplementary information: Supplementary data are available at Bioinformatics online. Caralyn Reisle, Karen L. Mungall, Caleb Choo, Daniel Paulino, Dustin W. Bleile, Amir Muhammadzadeh, Andrew J. Mungall, Richard A. Moore, Inna Shlafman, Robin Coope, Stephen Pleasance, Yussanne Ma, Steven J. M. Jones |
Bioinform. | 13 |
| 2019 | ntEdit: scalable genome sequence polishingabstractMOTIVATION: In the modern genomics era, genome sequence assemblies are routine practice. However, depending on the methodology, resulting drafts may contain considerable base errors. Although utilities exist for genome base polishing, they work best with high read coverage and do not scale well. We developed ntEdit, a Bloom filter-based genome sequence editing utility that scales to large mammalian and conifer genomes. RESULTS: We first tested ntEdit and the state-of-the-art assembly improvement tools GATK, Pilon and Racon on controlled Escherichia coli and Caenorhabditis elegans sequence data. Generally, ntEdit performs well at low sequence depths (<20×), fixing the majority (>97%) of base substitutions and indels, and its performance is largely constant with increased coverage. In all experiments conducted using a single CPU, the ntEdit pipeline executed in <14 s and <3 m, on average, on E.coli and C.elegans, respectively. We performed similar benchmarks on a sub-20× coverage human genome sequence dataset, inspecting accuracy and resource usage in editing chromosomes 1 and 21, and whole genome. ntEdit scaled linearly, executing in 30-40 m on those sequences. We show how ntEdit ran in <2 h 20 m to improve upon long and linked read human genome assemblies of NA12878, using high-coverage (54×) Illumina sequence data from the same individual, fixing frame shifts in coding sequences. We also generated 17-fold coverage spruce sequence data from haploid sequence sources (seed megagametophyte), and used it to edit our pseudo haploid assemblies of the 20 Gb interior and white spruce genomes in <4 and <5 h, respectively, making roughly 50M edits at a (substitution+indel) rate of 0.0024. AVAILABILITY AND IMPLEMENTATION: https://github.com/bcgsc/ntedit. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. René L. Warren, Lauren Coombe, Hamid Mohamadi, Barry Jaquish, Nathalie Isabel, Steven J. M. Jones, Jean Bousquet, Jörg Bohlmann, Inanç Birol |
Bioinform. | 7 |
| 2018 | A collaborative filtering-based approach to biomedical knowledge discoveryabstractMotivation: The increase in publication rates makes it challenging for an individual researcher to stay abreast of all relevant research in order to find novel research hypotheses. Literature-based discovery methods make use of knowledge graphs built using text mining and can infer future associations between biomedical concepts that will likely occur in new publications. These predictions are a valuable resource for researchers to explore a research topic. Current methods for prediction are based on the local structure of the knowledge graph. A method that uses global knowledge from across the knowledge graph needs to be developed in order to make knowledge discovery a frequently used tool by researchers. Results: We propose an approach based on the singular value decomposition (SVD) that is able to combine data from across the knowledge graph through a reduced representation. Using cooccurrence data extracted from published literature, we show that SVD performs better than the leading methods for scoring discoveries. We also show the diminishing predictive power of knowledge discovery as we compare our predictions with real associations that appear further into the future. Finally, we examine the strengths and weaknesses of the SVD approach against another well-performing system using several predicted associations. Availability and implementation: All code and results files for this analysis can be accessed at https://github.com/jakelever/knowledgediscovery. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Jake Lever, Sitanshu Gakkhar, Michael Gottlieb, Tahereh Rashnavadi, Santina Lin, Celia Siu, Maia Smith, Martin R. Jones, Martin Krzywinski, Steven J. M. Jones |
Bioinform. | 10 |
| 2018 | Tigmint: correcting assembly errors using linked reads from large moleculesabstractBACKGROUND: Genome sequencing yields the sequence of many short snippets of DNA (reads) from a genome. Genome assembly attempts to reconstruct the original genome from which these reads were derived. This task is difficult due to gaps and errors in the sequencing data, repetitive sequence in the underlying genome, and heterozygosity. As a result, assembly errors are common. In the absence of a reference genome, these misassemblies may be identified by comparing the sequencing data to the assembly and looking for discrepancies between the two. Once identified, these misassemblies may be corrected, improving the quality of the assembled sequence. Although tools exist to identify and correct misassemblies using Illumina paired-end and mate-pair sequencing, no such tool yet exists that makes use of the long distance information of the large molecules provided by linked reads, such as those offered by the 10x Genomics Chromium platform. We have developed the tool Tigmint to address this gap. RESULTS: To demonstrate the effectiveness of Tigmint, we applied it to assemblies of a human genome using short reads assembled with ABySS 2.0 and other assemblers. Tigmint reduced the number of misassemblies identified by QUAST in the ABySS assembly by 216 (27%). While scaffolding with ARCS alone more than doubled the scaffold NGA50 of the assembly from 3 to 8 Mbp, the combination of Tigmint and ARCS improved the scaffold NGA50 of the assembly over five-fold to 16.4 Mbp. This notable improvement in contiguity highlights the utility of assembly correction in refining assemblies. We demonstrate the utility of Tigmint in correcting the assemblies of multiple tools, as well as in using Chromium reads to correct and scaffold assemblies of long single-molecule sequencing. CONCLUSIONS: Scaffolding an assembly that has been corrected with Tigmint yields a final assembly that is both more correct and substantially more contiguous than an assembly that has not been corrected. Using single-molecule sequencing in combination with linked reads enables a genome sequence assembly that achieves both a high sequence contiguity as well as high scaffold contiguity, a feat not currently achievable with either technology alone. Shaun D. Jackman, Lauren Coombe, Justin Chu, René L. Warren, Benjamin P. Vandervalk, Sarah Yeo, Zhuyi Xue, Hamid Mohamadi, Jörg Bohlmann, Steven J. M. Jones, Inanç Birol |
BMC Bioinform. | 10 |
| 2016 | ChAsE: chromatin analysis and exploration toolabstract: We present ChAsE, a cross-platform desktop application developed for interactive visualization, exploration and clustering of epigenomic data such as ChIP-seq experiments. ChAsE is designed and developed in close collaboration with several groups of biologists and bioinformaticians with a focus on usability and interactivity. Data can be analyzed through k-means clustering, specifying presence or absence of signal in epigenetic data and performing set operations between clusters. Results can be explored in an interactive heat map and profile plot interface and exported for downstream analysis or as high quality figures suitable for publications. AVAILABILITY AND IMPLEMENTATION: Software, source code (MIT License), data and video tutorials available at http://chase.cs.univie.ac.at CONTACT: : [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Hamid Younesy, Cydney B. Nielsen, Matthew C. Lorincz, Steven J. M. Jones, Mohammad M. Karimi, Torsten Möller |
Bioinform. | 4 |
| 2015 | VisRseq: R-based visual framework for analysis of sequencing dataabstractBACKGROUND: Several tools have been developed to enable biologists to perform initial browsing and exploration of sequencing data. However the computational tool set for further analyses often requires significant computational expertise to use and many of the biologists with the knowledge needed to interpret these data must rely on programming experts. RESULTS: We present VisRseq, a framework for analysis of sequencing datasets that provides a computationally rich and accessible framework for integrative and interactive analyses without requiring programming expertise. We achieve this aim by providing R apps, which offer a semi-auto generated and unified graphical user interface for computational packages in R and repositories such as Bioconductor. To address the interactivity limitation inherent in R libraries, our framework includes several native apps that provide exploration and brushing operations as well as an integrated genome browser. The apps can be chained together to create more powerful analysis workflows. CONCLUSIONS: To validate the usability of VisRseq for analysis of sequencing data, we present two case studies performed by our collaborators and report their workflow and insights. Hamid Younesy, Torsten Möller, Matthew C. Lorincz, Mohammad M. Karimi, Steven J. M. Jones |
BMC Bioinform. | 5 |
| 2014 | ALEA: a toolbox for allele-specific epigenomics analysisabstractThe assessment of expression and epigenomic status using sequencing based methods provides an unprecedented opportunity to identify and correlate allelic differences with epigenomic status. We present ALEA, a computational toolbox for allele-specific epigenomics analysis, which incorporates allelic variation data within existing resources, allowing for the identification of significant associations between epigenetic modifications and specific allelic variants in human and mouse cells. ALEA provides a customizable pipeline of command line tools for allele-specific analysis of next-generation sequencing data (ChIP-seq, RNA-seq, etc.) that takes the raw sequencing data and produces separate allelic tracks ready to be viewed on genome browsers. The pipeline has been validated using human and hybrid mouse ChIP-seq and RNA-seq data. AVAILABILITY: The package, test data and usage instructions are available online at http://www.bcgsc.ca/platform/bioinfo/software/alea CONTACT: : [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Hamid Younesy, Torsten Möller, Alireza Heravi Moussavi, Jeffrey B. Cheng, Joseph F. Costello, Matthew C. Lorincz, Mohammad M. Karimi, Steven J. M. Jones |
Bioinform. | 8 |
| 2013 | Assembling the 20 Gb white spruce (Picea glauca) genome from whole-genome shotgun sequencing dataabstractUNLABELLED: White spruce (Picea glauca) is a dominant conifer of the boreal forests of North America, and providing genomics resources for this commercially valuable tree will help improve forest management and conservation efforts. Sequencing and assembling the large and highly repetitive spruce genome though pushes the boundaries of the current technology. Here, we describe a whole-genome shotgun sequencing strategy using two Illumina sequencing platforms and an assembly approach using the ABySS software. We report a 20.8 giga base pairs draft genome in 4.9 million scaffolds, with a scaffold N50 of 20,356 bp. We demonstrate how recent improvements in the sequencing technology, especially increasing read lengths and paired end reads from longer fragments have a major impact on the assembly contiguity. We also note that scalable bioinformatics tools are instrumental in providing rapid draft assemblies. AVAILABILITY: The Picea glauca genome sequencing and assembly data are available through NCBI (Accession#: ALWZ0100000000 PID: PRJNA83435). http://www.ncbi.nlm.nih.gov/bioproject/83435. Inanç Birol, Anthony Raymond, Shaun D. Jackman, Stephen Pleasance, Robin Coope, Greg A. Taylor, Macaire Man Saint Yuen, Christopher I. Keeling, Dana Brand, Benjamin P. Vandervalk, Heather Kirk, Pawan Pandoh, Richard A. Moore, Yongjun Zhao 0002, Andrew J. Mungall, Barry Jaquish, Alvin Yanchuk, Carol Ritland, Brian Boyle, Jean Bousquet, Kermit Ritland, John MacKay, Jörg Bohlmann, Steven J. M. Jones |
Bioinform. | 24 |
| 2013 | Identifying cancer mutation targets across thousands of samples: MuteProc, a high throughput mutation analysis pipelineabstractBACKGROUND: In the past decade, bioinformatics tools have matured enough to reliably perform sophisticated primary data analysis on Next Generation Sequencing (NGS) data, such as mapping, assemblies and variant calling, however, there is still a dire need for improvements in the higher level analysis such as NGS data organization, analysis of mutation patterns and Genome Wide Association Studies (GWAS). RESULTS: We present a high throughput pipeline for identifying cancer mutation targets, capable of processing billions of variations across thousands of samples. This pipeline is coupled with our Human Variation Database to provide more complex down stream analysis on the variations hosted in the database. Most notably, these analysis include finding significantly mutated regions across multiple genomes and regions with mutational preferences within certain types of cancers. The results of the analysis is presented in HTML summary reports that incorporate gene annotations from various resources for the reported regions. CONCLUSION: MuteProc is available for download through the Vancouver Short Read Analysis Package on Sourceforge: http://vancouvershortr.sourceforge.net. Instructions for use and a tutorial are provided on the accompanying wiki pages at https://sourceforge.net/apps/mediawiki/vancouvershortr/index.php?title=Pipeline_introduction. Alireza Hadj Khodabakhshi, Anthony P. Fejes, Inanç Birol, Steven J. M. Jones |
BMC Bioinform. | 4 |
| 2013 | An Interactive Analysis and Exploration Tool for Epigenomic DataabstractAbstract In this design study, we present an analysis and abstraction of the data and tasks related to the domain of epigenomics, and the design and implementation of an interactive tool to facilitate data analysis and visualization in this domain. Epigenomic data can be grouped into subsets either by k‐means clustering or by querying for combinations of presence or absence of signal (on/off) in different epigenomic experiments. These steps can easily be interleaved and the comparison of different workflows is explicitly supported. We took special care to contain the exponential expansion of possible on/off combinations by creating a novel querying interface. An interactive heat map facilitates the exploration and comparison of different clusters. We validated our iterative design by working closely with two groups of biologists on different biological problems. Both groups quickly found new insight into their data as well as claimed that our tool would save them several hours or days of work over using existing tools. Hamid Younesy, Cydney B. Nielsen, Torsten Möller, Olivia Alder, Rebecca Cullum, Matthew C. Lorincz, Mohammad M. Karimi, Steven J. M. Jones |
Comput. Graph. Forum | 8 |
| 2012 | Hive plots - rational approach to visualizing networksabstractNetworks are typically visualized with force-based or spectral layouts. These algorithms lack reproducibility and perceptual uniformity because they do not use a node coordinate system. The layouts can be difficult to interpret and are unsuitable for assessing differences in networks. To address these issues, we introduce hive plots (http://www.hiveplot.com) for generating informative, quantitative and comparable network layouts. Hive plots depict network structure transparently, are simple to understand and can be easily tuned to identify patterns of interest. The method is computationally straightforward, scales well and is amenable to a plugin for existing tools. Martin Krzywinski, Inanç Birol, Steven J. M. Jones, Marco A. Marra |
Briefings Bioinform. | 3 |
| 2012 | BreakFusion: targeted assembly-based identification of gene fusions in whole transcriptome paired-end sequencing dataabstractAbstract Summary: Despite recent progress, computational tools that identify gene fusions from next-generation whole transcriptome sequencing data are often limited in accuracy and scalability. Here, we present a software package, BreakFusion that combines the strength of reference alignment followed by read-pair analysis and de novo assembly to achieve a good balance in sensitivity, specificity and computational efficiency. Availability: http://bioinformatics.mdanderson.org/main/BreakFusion Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online Ken Chen 0001, John W. Wallis, Cyriac Kandoth, Joelle M. Kalicki-Veizer, Karen L. Mungall, Andrew J. Mungall, Steven J. M. Jones, Marco A. Marra, Timothy J. Ley, Elaine R. Mardis, Richard K. Wilson, John N. Weinstein |
Bioinform. | 7 |
| 2012 | On the Deep Order-Preserving Submatrix Problem: A Best Effort ApproachabstractOrder-preserving submatrix (OPSM) has been widely accepted as a biologically meaningful cluster model, capturing the general tendency of gene expression across a subset of experiments. In an OPSM, the expression levels of all genes induce the same linear ordering of the experiments. The OPSM problem is to discover those statistically significant OPSMs from a given data matrix. The problem is reducible to a special case of the sequential pattern mining problem, where a pattern and its supporting sequences uniquely specify an OPSM. Unfortunately, existing methods do not scale well to massive data sets containing thousands of experiments and hundreds of thousands of genes, which are common in today's gene expression analysis. In particular, deep OPSMs, corresponding to long patterns with few supporting sequences, incur explosive computational costs in their discovery and are completely pruned off by existing methods. However, it is of particular interest of biologists to determine small groups of genes that are tightly coregulated across many experiments, and some pathways or processes may require as few as two genes to act in concert. In this paper, we study the discovery of deep OPSMs from massive data sets. We propose a novel best effort mining framework Kiwi that exploits two parameters k and w to bound the available computational resources and search a selected search space, and does what it can to find as many as possible deep OPSMs. Extensive biological and computational evaluations on real data sets demonstrate the validity and importance of the deep OPSM problem, and the efficiency and effectiveness of the Kiwi mining framework. Byron J. Gao, Obi L. Griffith, Martin Ester, Hui Xiong 0001, Steven J. M. Jones |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2011 | Human variation database: an open-source database template for genomic discoveryabstractMOTIVATION: Current public variation databases are based upon collaboratively pooling data into a single database with a single interface available to the public. This gives little control to the collaborator to mine the database and requires that they freely share their data with the owners of the repository. We aim to provide an alternative mechanism: providing the source code and application programming interface (API) of a database, enabling researchers to set up local versions without investing heavily in the development of the resource and allowing for confidential information to remain secure. RESULTS: We describe an open-source database that can be installed easily at any research facility for the storage and analysis of thousands of next-generation sequencing variations. This database is built using PostgreSQL 8.4 (The PostgreSQL Global Development Group. postgres 8.4: http://www.postgresql.org) and provides a novel method for collating and searching across the reported results from thousands of next-generation sequence samples, as well as rapidly accessing vital information on the origin of the samples. The schema of the database makes rapid and insightful queries simple and enables easy annotation of novel or known genetic variations. A modular and cross-platform Java API is provided to perform common functions, such as generation of standard experimental reports and graphical summaries of modifications to genes. Included libraries allow adopters of the database to quickly develop their own queries. AVAILABILITY: The software is available for download through the Vancouver Short Read Analysis Package on Sourceforge, http://vancouvershortr.sourceforge.net. Instructions for use and deployment are provided on the accompanying wiki pages. CONTACT: [email protected]. Anthony P. Fejes, Alireza Hadj Khodabakhshi, Inanç Birol, Steven J. M. Jones |
Bioinform. | 4 |
| 2011 | A Computational Approach to Finding Novel Targets for Existing DrugsabstractRepositioning existing drugs for new therapeutic uses is an efficient approach to drug discovery. We have developed a computational drug repositioning pipeline to perform large-scale molecular docking of small molecule drugs against protein drug targets, in order to map the drug-target interaction space and find novel interactions. Our method emphasizes removing false positive interaction predictions using criteria from known interaction docking, consensus scoring, and specificity. In all, our database contains 252 human protein drug targets that we classify as reliable-for-docking as well as 4621 approved and experimental small molecule drugs from DrugBank. These were cross-docked, then filtered through stringent scoring criteria to select top drug-target interactions. In particular, we used MAPK14 and the kinase inhibitor BIM-8 as examples where our stringent thresholds enriched the predicted drug-target interactions with known interactions up to 20 times compared to standard score thresholds. We validated nilotinib as a potent MAPK14 inhibitor in vitro (IC50 40 nM), suggesting a potential use for this drug in treating inflammatory diseases. The published literature indicated experimental evidence for 31 of the top predicted interactions, highlighting the promising nature of our approach. Novel interactions discovered may lead to the drug being repositioned as a therapeutic treatment for its off-target's associated disease, added insight into the drug's mechanism of action, and added insight into the drug's side effects. Yvonne Y. Li, Jianghong An, Steven J. M. Jones |
PLoS Comput. Biol. | 3 |
| 2010 | High quality SNP calling using Illumina data at shallow coverageabstractMOTIVATION: Detection of single nucleotide polymorphisms (SNPs) has been a major application in processing second generation sequencing (SGS) data. In principle, SNPs are called on single base differences between a reference genome and a sequence generated from SGS short reads of a sample genome. However, this exercise is far from trivial; several parameters related to sequencing quality, and/or reference genome properties, play essential effect on the accuracy of called SNPs especially at shallow coverage data. In this work, we present Slider II, an alignment and SNP calling approach that demonstrates improved algorithmic approaches enabling larger number of called SNPs with lower false positive rate. In addition to the regular alignment and SNP calling, as an optional feature, Slider II is capable of utilizing information about known SNPs of a target genome, as priors, in the alignment and SNPs calling to enhance it's capability of detecting these known SNPs and novel SNPs and mutations in their vicinity. Nawar Malhis, Steven J. M. Jones |
Bioinform. | 2 |
| 2010 | Genomic analysis of a rare human tumorabstractThe introduction of next-generation DNA sequencing devices into the field of oncology provides an unprecedented mechanism to determine the underlying genetic changes that have occurred within a tumor and also the changes that accrue during treatment. An enhanced understanding of the oncogenic mechanisms could have an immediate clinical role in the treatment of rare tumors - where treatment protocols do not exist and their rarity would indicate that clinical trials would be unlikely to be undertaken for their establishment. We have investigated the utility of massively parallel sequencing to characterize a rare adenocarcinoma of the tongue, before and after treatment. In the pre-treatment tumor we identified 7,629 genes within regions of copy number gain, 1,078 genes exhibited increased expression relative to the blood and unrelated tumors and four genes contained somatic protein-coding mutations. Our analysis suggested the tumor cells were driven by the RET oncogene and its other pathway constituents. Genes whose protein products are targeted by the RET inhibitors sunitinib and sorafenib correlated with being amplified and or highly expressed. Consistent with our observations subsequent administration of sunitinib was associated with stable disease lasting 4 months, after which the lung lesions began to grow. Administration of sorafenib and sulindac provided disease stabilization for an additional 3 months after which the cancer progressed and new lesions appeared. A metastasis recurring in the skin was determined to possess 7,288 genes within copy number amplicons, 385 genes exhibiting increased expression relative to other tumours and 9 new somatic protein coding mutations. The observed mutations and amplifications were found to be consistent with resistance to therapy arising through further activation of RET pathway and nascent activation of the AKT pathway. Our results provide evidence for the clinical utility of complete genomic characterization and direct in-vivo genome-wide characterization of the mutations accruing within a tumor under drug selection. Steven J. M. Jones, Janessa Laskin, Yvonne Y. Li, Obi L. Griffith, Jianghong An, Mikhail Bilenky, Yaron S. N. Butterfield, Timothee Cezard, Eric Chuah, Richard Corbett, Anthony P. Fejes, Malachi Griffith, John Yee, Montgomery Martin, Michael Mayo, Nataliya Melnyk, Ryan D. Morin, Trevor J. Pugh, Tesa Severson, Sohrab P. Shah, Margaret Sutcliffe, Angela Tam, Jefferson Terry, Nina Thiessen, Thomas Thomson, Richard Varhol, Thomas Zeng 0002, Yongjun Zhao 0002, Richard A. Moore, David G. Huntsman, Inanç Birol, Martin Hirst, Robert A. Holt, Marco A. Marra |
BMC Bioinform. | 1 |
| 2010 | LaneRuler: Automated Lane Tracking for DNA Electrophoresis Gel ImagesabstractWe present a novel method for correctly identifying and straightening one dimensional agarose electrophoretic lanes. Our method has been shown to yield comparable accuracy with manual lane tracking results, and to successfully process 98% of DNA fingerprinting gels with no human intervention. R. T. F. Wong, Stephane Flibotte, Richard Corbett, Parvaneh Saeedi, Steven J. M. Jones, Marco A. Marra, Jacqueline E. Schein, Inanç Birol |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2009 | De novo transcriptome assembly with ABySSabstractMOTIVATION: Whole transcriptome shotgun sequencing data from non-normalized samples offer unique opportunities to study the metabolic states of organisms. One can deduce gene expression levels using sequence coverage as a surrogate, identify coding changes or discover novel isoforms or transcripts. Especially for discovery of novel events, de novo assembly of transcriptomes is desirable. RESULTS: Transcriptome from tumor tissue of a patient with follicular lymphoma was sequenced with 36 base pair (bp) single- and paired-end reads on the Illumina Genome Analyzer II platform. We assembled approximately 194 million reads using ABySS into 66 921 contigs 100 bp or longer, with a maximum contig length of 10 951 bp, representing over 30 million base pairs of unique transcriptome sequence, or roughly 1% of the genome. AVAILABILITY AND IMPLEMENTATION: Source code and binaries of ABySS are freely available for download at http://www.bcgsc.ca/platform/bioinfo/software/abyss. Assembler tool is implemented in C++. The parallel version uses Open MPI. ABySS-Explorer tool is implemented in Java using the Java universal network/graph framework. CONTACT: [email protected]. Inanç Birol, Shaun D. Jackman, Cydney B. Nielsen, Jenny Q. Qian, Richard Varhol, Greg Stazyk, Ryan D. Morin, Yongjun Zhao 0002, Martin Hirst, Jacqueline E. Schein, Douglas E. Horsman, Joseph M. Connors, Randy D. Gascoyne, Marco A. Marra, Steven J. M. Jones |
Bioinform. | 15 |
| 2009 | Slider - maximum use of probability information for alignment of short sequence reads and SNP detectionabstractMOTIVATION: A plethora of alignment tools have been created that are designed to best fit different types of alignment conditions. While some of these are made for aligning Illumina Sequence Analyzer reads, none of these are fully utilizing its probability (prb) output. In this article, we will introduce a new alignment approach (Slider) that reduces the alignment problem space by utilizing each read base's probabilities given in the prb files. RESULTS: Compared with other aligners, Slider has higher alignment accuracy and efficiency. In addition, given that Slider matches bases with probabilities other than the most probable, it significantly reduces the percentage of base mismatches. The result is that its SNP predictions are more accurate than other SNP prediction approaches used today that start from the most probable sequence, including those using base quality. Nawar Malhis, Yaron S. N. Butterfield, Martin Ester, Steven J. M. Jones |
Bioinform. | 4 |
| 2009 | ABySS-Explorer: Visualizing Genome Sequence AssembliesabstractOne bottleneck in large-scale genome sequencing projects is reconstructing the full genome sequence from the short subsequences produced by current technologies. The final stages of the genome assembly process inevitably require manual inspection of data inconsistencies and could be greatly aided by visualization. This paper presents our design decisions in translating key data features identified through discussions with analysts into a concise visual encoding. Current visualization tools in this domain focus on local sequence errors making high-level inspection of the assembly difficult if not impossible. We present a novel interactive graph display, ABySS-Explorer, that emphasizes the global assembly structure while also integrating salient data features such as sequence length. Our tool replaces manual and in some cases pen-and-paper based analysis tasks, and we discuss how user feedback was incorporated into iterative design refinements. Finally, we touch on applications of this representation not initially considered in our design phase, suggesting the generality of this encoding for DNA sequence data. Cydney B. Nielsen, Shaun D. Jackman, Inanç Birol, Steven J. M. Jones |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technologyabstractSUMMARY: Next-generation sequencing can provide insight into protein-DNA association events on a genome-wide scale, and is being applied in an increasing number of applications in genomics and meta-genomics research. However, few software applications are available for interpreting these experiments. We present here an efficient application for use with chromatin-immunoprecipitation (ChIP-Seq) experimental data that includes novel functionality for identifying areas of gene enrichment and transcription factor binding site locations, as well as for estimating DNA fragment size distributions in enriched areas. The FindPeaks application can generate UCSC compatible custom 'WIG' track files from aligned-read files for short-read sequencing technology. The software application can be executed on any platform capable of running a Java Runtime Environment. Memory requirements are proportional to the number of sequencing reads analyzed; typically 4 GB permits processing of up to 40 million reads. AVAILABILITY: The FindPeaks 3.1 package and manual, containing algorithm descriptions, usage instructions and examples, are available at http://www.bcgsc.ca/platform/bioinfo/software/findpeaks Source files for FindPeaks 3.1 are available for academic use. Anthony P. Fejes, Gordon Robertson, Mikhail Bilenky, Richard Varhol, Matthew N. Bainbridge, Steven J. M. Jones |
Bioinform. | 6 |
| 2007 | THOR: targeted high-throughput ortholog reconstructorabstractAbstract Summary: Low-coverage genomes (LCGs) are becoming an increasingly important source of data for phylogenetic studies. However, assembly of these genomes is time consuming, difficult and lags behind sequence generation. THOR is a fast, stringent application for targeted reconstruction of sequence orthologs in unassembled LCGs. Using a 4× coverage set of mouse whole-genome sequence reads, THOR could partially or completely reconstruct 416/1000 human promoter ortholog regions in ∼7.3 min/promoter. THOR's reconstruction rate improves markedly with both higher-coverage, and less divergent target species. Availability: THOR is implemented in java and is currently available as source code and as a web service (www.bcgsc.ca/services/thor) for reconstructing human sequences. Contact: [email protected] Matthew N. Bainbridge, René L. Warren, An He, Mikhail Bilenky, Gordon Robertson, Steven J. M. Jones |
Bioinform. | 6 |
| 2007 | Assembling millions of short DNA sequences using SSAKEabstractUNLABELLED: Novel DNA sequencing technologies with the potential for up to three orders magnitude more sequence throughput than conventional Sanger sequencing are emerging. The instrument now available from Solexa Ltd, produces millions of short DNA sequences of 25 nt each. Due to ubiquitous repeats in large genomes and the inability of short sequences to uniquely and unambiguously characterize them, the short read length limits applicability for de novo sequencing. However, given the sequencing depth and the throughput of this instrument, stringent assembly of highly identical sequences can be achieved. We describe SSAKE, a tool for aggressively assembling millions of short nucleotide sequences by progressively searching through a prefix tree for the longest possible overlap between any two sequences. SSAKE is designed to help leverage the information from short sequence reads by stringently assembling them into contiguous sequences that can be used to characterize novel sequencing targets. AVAILABILITY: http://www.bcgsc.ca/bioinfo/software/ssake. René L. Warren, Granger G. Sutton, Steven J. M. Jones, Robert A. Holt |
Bioinform. | 3 |
| 2007 | A Survey of Genomic Properties for the Detection of Regulatory PolymorphismsabstractAdvances in the computational identification of functional noncoding polymorphisms will aid in cataloging novel determinants of health and identifying genetic variants that explain human evolution. To date, however, the development and evaluation of such techniques has been limited by the availability of known regulatory polymorphisms. We have attempted to address this by assembling, from the literature, a computationally tractable set of regulatory polymorphisms within the ORegAnno database (http://www.oreganno.org). We have further used 104 regulatory single-nucleotide polymorphisms from this set and 951 polymorphisms of unknown function, from 2-kb and 152-bp noncoding upstream regions of genes, to investigate the discriminatory potential of 23 properties related to gene regulation and population genetics. Among the most important properties detected in this region are distance to transcription start site, local repetitive content, sequence conservation, minor and derived allele frequencies, and presence of a CpG island. We further used the entire set of properties to evaluate their collective performance in detecting regulatory polymorphisms. Using a 10-fold cross-validation approach, we were able to achieve a sensitivity and specificity of 0.82 and 0.71, respectively, and we show that this performance is strongly influenced by the distance to the transcription start site. Stephen B. Montgomery, Obi L. Griffith, Johanna M. Schuetz, Angela Brooks-Wilson, Steven J. M. Jones |
PLoS Comput. Biol. | 5 |
| 2006 | Discovering significant OPSM subspace clusters in massive gene expression dataabstractOrder-preserving submatrixes (OPSMs) have been accepted as a biologically meaningful subspace cluster model, capturing the general tendency of gene expressions across a subset of conditions. In an OPSM, the expression levels of all genes induce the same linear ordering of the conditions. OPSM mining is reducible to a special case of the sequential pattern mining problem, in which a pattern and its supporting sequences uniquely specify an OPSM cluster. Those small twig clusters, specified by long patterns with naturally low support, incur explosive computational costs and would be completely pruned off by most existing methods for massive datasets containing thousands of conditions and hundreds of thousands of genes, which are common in today's gene expression analysis. However, it is in particular interest of biologists to reveal such small groups of genes that are tightly coregulated under many conditions, and some pathways or processes might require only two genes to act in concert. In this paper, we introduce the KiWi mining framework for massive datasets, that exploits two parameters k and w to provide a biased testing on a bounded number of candidates, substantially reducing the search space and problem scale, targeting on highly promising seeds that lead to significant clusters and twig clusters. Extensive biological and computational evaluations on real datasets demonstrate that KiWi can effectively mine biologically meaningful OPSM subspace clusters with good efficiency and scalability. Byron J. Gao, Obi L. Griffith, Martin Ester, Steven J. M. Jones |
KDD | 4 |
| 2006 | ORegAnno: an open access database and curation system for literature-derived promoters, transcription factor binding sites and regulatory variationabstractMOTIVATION: Our understanding of gene regulation is currently limited by our ability to collectively synthesize and catalogue transcriptional regulatory elements stored in scientific literature. Over the past decade, this task has become increasingly challenging as the accrual of biologically validated regulatory sequences has accelerated. To meet this challenge, novel community-based approaches to regulatory element annotation are required. SUMMARY: Here, we present the Open Regulatory Annotation (ORegAnno) database as a dynamic collection of literature-curated regulatory regions, transcription factor binding sites and regulatory mutations (polymorphisms and haplotypes). ORegAnno has been designed to manage the submission, indexing and validation of new annotations from users worldwide. Submissions to ORegAnno are immediately cross-referenced to EnsEMBL, dbSNP, Entrez Gene, the NCBI Taxonomy database and PubMed, where appropriate. AVAILABILITY: ORegAnno is available directly through MySQL, Web services, and online at http://www.oreganno.org. All software is licensed under the Lesser GNU Public License (LGPL). Stephen B. Montgomery, Obi L. Griffith, Monica C. Sleumer, Casey M. Bergman, Misha Bilenky, Erin Pleasance, Y. Prychyna, Steven J. M. Jones |
Bioinform. | 9 |
| 2006 | An interactive tool for visualization of relationships between gene expression profilesabstractBACKGROUND: Application of phenetic methods to gene expression analysis proved to be a successful approach. Visualizing the results in a 3-dimentional space may further enhance these techniques. RESULTS: We designed and built TreeBuilder3D, an interactive viewer for visualizing the hierarchical relationships between expression profiles such as SAGE libraries or microarrays. The program allows loading expression data as plain text files and visualizing the relative differences of the analyzed datasets in 3-dimensional space using various distance metrics. CONCLUSION: TreeBuilder3D provides a simple interface and has a small size. Written in Java, TreeBuilder3D is a platform-independent, open source application, which may be useful in analysis of large-scale gene expression data. Peter Ruzanov, Steven J. M. Jones |
BMC Bioinform. | 2 |
| 2005 | CGMIM: Automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genesabstractBACKGROUND: Online Mendelian Inheritance in Man (OMIM) is a computerized database of information about genes and heritable traits in human populations, based on information reported in the scientific literature. Our objective was to establish an automated text-mining system for OMIM that will identify genetically-related cancers and cancer-related genes. We developed the computer program CGMIM to search for entries in OMIM that are related to one or more cancer types. We performed manual searches of OMIM to verify the program results. RESULTS: In the OMIM database on September 30, 2004, CGMIM identified 1943 genes related to cancer. BRCA2 (OMIM *164757), BRAF (OMIM *164757) and CDKN2A (OMIM *600160) were each related to 14 types of cancer. There were 45 genes related to cancer of the esophagus, 121 genes related to cancer of the stomach, and 21 genes related to both. Analysis of CGMIM results indicate that fewer than three gene entries in OMIM should mention both, and the more than seven-fold discrepancy suggests cancers of the esophagus and stomach are more genetically related than current literature suggests. CONCLUSION: CGMIM identifies genetically-related cancers and cancer-related genes. In several ways, cancers with shared genetic etiology are anticipated to lead to further etiologic hypotheses and advances regarding environmental agents. CGMIM results are posted monthly and the source code can be obtained free of charge from the BC Cancer Research Centre website http://www.bccrc.ca/ccr/CGMIM Chris D. Bajdik, Byron Kuo, Shawn Rusaw, Steven J. M. Jones, Angela Brooks-Wilson |
BMC Bioinform. | 4 |
| 2005 | A methodology for analyzing SAGE libraries for cancer profilingabstractSerial Analysis of Gene Expression (SAGE) has proven to be an important alternative to microarray techniques for global profiling of mRNA populations. We have developed preprocessing methodologies to address problems in analyzing SAGE data due to noise caused by sequencing error, normalization methodologies to account for libraries sampled at different depths, and missing tag imputation methodologies to aid in the analysis of poorly sampled SAGE libraries. We have also used subspace selection using the Wilcoxon rank sum test to exclude tags that have similar expression levels regardless of source. Using these methodologies we have clustered, using the OPTICS algorithm, 88 SAGE libraries derived from cancerous and normal tissues as well as cell line material. Our results produced eight dense clusters representing ovarian cancer cell line, brain cancer cell line, brain cancer bulk tissue, prostate tissue, pancreatic cancer, breast cancer cell line, normal brain, and normal breast bulk tissue. The ovarian cancer and brain cancer cell lines clustered closely together, leading to a further investigation on possible associations between these two cancer types. We also investigated the utility of gene expression data in the classification between normal and cancerous tissues. Our results indicate that brain and breast cancer libraries have strong identities allowing robust discrimination from their normal counterparts. However, the SAGE expression data provide poor predictive accuracy in discriminating between prostate and ovarian cancers and their respective normal tissues. Jörg Sander 0001, Raymond T. Ng, Monica C. Sleumer, Macaire Man Saint Yuen, Steven J. M. Jones |
ACM Trans. Inf. Syst. | 5 |
| 2004 | Structural characterization of genomes by large scale sequence-structure threadingabstractBACKGROUND: Using sequence-structure threading we have conducted structural characterization of complete proteomes of 37 archaeal, bacterial and eukaryotic organisms (including worm, fly, mouse and human) totaling 167,888 genes. RESULTS: The reported data represent first rather general evaluation of performance of full sequence-structure threading on multiple genomes providing opportunity to evaluate its general applicability for large scale studies. According to the estimated results the sequence-structure threading has assigned protein folds to more then 60% of eukaryotic, 68% of archaeal and 70% of bacterial proteomes.The repertoires of protein classes, architectures, topologies and homologous superfamilies (according to the CATH 2.4 classification) have been established for distant organisms and superkingdoms. It has been found that the average abundance of CATH classes decreases from "alpha and beta" to "mainly beta", followed by "mainly alpha" and "few secondary structures".3-Layer (aba) Sandwich has been characterized as the most abundant protein architecture and Rossman fold as the most common topology. CONCLUSION: The analysis of genomic occurrences of CATH 2.4 protein homologous superfamilies and topologies has revealed the power-law character of their distributions. The corresponding double logarithmic "frequency - genomic occurrence" dependences characteristic of scale-free systems have been established for individual organisms and for three superkingdoms. Artem Cherkasov, Steven J. M. Jones |
BMC Bioinform. | 2 |
| 2004 | An approach to large scale identification of non-obvious structural similarities between proteinsabstractBACKGROUND: A new sequence independent bioinformatics approach allowing genome-wide search for proteins with similar three dimensional structures has been developed. By utilizing the numerical output of the sequence threading it establishes putative non-obvious structural similarities between proteins. When applied to the testing set of proteins with known three dimensional structures the developed approach was able to recognize structurally similar proteins with high accuracy. RESULTS: The method has been developed to identify pathogenic proteins with low sequence identity and high structural similarity to host analogues. Such protein structure relationships would be hypothesized to arise through convergent evolution or through ancient horizontal gene transfer events, now undetectable using current sequence alignment techniques. The pathogen proteins, which could mimic or interfere with host activities, would represent candidate virulence factors. The developed approach utilizes the numerical outputs from the sequence-structure threading. It identifies the potential structural similarity between a pair of proteins by correlating the threading scores of the corresponding two primary sequences against the library of the standard folds. This approach allowed up to 64% sensitivity and 99.9% specificity in distinguishing protein pairs with high structural similarity. CONCLUSION: Preliminary results obtained by comparison of the genomes of Homo sapiens and several strains of Chlamydia trachomatis have demonstrated the potential usefulness of the method in the identification of bacterial proteins with known or potential roles in virulence. Artem Cherkasov, Steven J. M. Jones |
BMC Bioinform. | 2 |
| 2004 | Structural characterization of genomes by large scale sequence-structure threading: application of reliability analysis in structural genomicsabstractBACKGROUND: We establish that the occurrence of protein folds among genomes can be accurately described with a Weibull function. Systems which exhibit Weibull character can be interpreted with reliability theory commonly used in engineering analysis. For instance, Weibull distributions are widely used in reliability, maintainability and safety work to model time-to-failure of mechanical devices, mechanisms, building constructions and equipment. RESULTS: We have found that the Weibull function describes protein fold distribution within and among genomes more accurately than conventional power functions which have been used in a number of structural genomic studies reported to date. It has also been found that the Weibull reliability parameter beta for protein fold distributions varies between genomes and may reflect differences in rates of gene duplication in evolutionary history of organisms. CONCLUSIONS: The results of this work demonstrate that reliability analysis can provide useful insights and testable predictions in the fields of comparative and structural genomics. Artem Cherkasov, Shannan J. Ho Sui, Robert C. Brunham, Steven J. M. Jones |
BMC Bioinform. | 4 |
| 2003 | IslandPath: aiding detection of genomic islands in prokaryotesabstractUNLABELLED: Genomic islands (clusters of genes of potential horizontal origin in a prokaryotic genome) are frequently associated with a particular adaptation of a microbe that is of medical, agricultural or environmental importance, such as antibiotic resistance, pathogen virulence, or metal resistance. While many sequence features associated with such islands have been adopted separately in applications for analysis of genomic islands, including pathogenicity islands, there is no single application that integrates multiple features for island detection. IslandPath is a network service which incorporates multiple DNA signals and genome annotation features into a graphical display of a bacterial or archaeal genome, to aid the detection of genomic islands. AVAILABILITY: This application is available at http://www.pathogenomics.sfu.ca/islandpath and the source code is freely available, under GNU public licence, from the authors. SUPPLEMENTARY INFORMATION: An online help file, which includes analyses of the utility of IslandPath, can be found at http://www.pathogenomics.sfu.ca/islandpath/current/islandhelp.html William W. L. Hsiao, Ivan Wan, Steven J. M. Jones, Fiona S. L. Brinkman |
Bioinform. | 3 |
| 2003 | A knowledge discovery object model API for JavaabstractBACKGROUND: Biological data resources have become heterogeneous and derive from multiple sources. This introduces challenges in the management and utilization of this data in software development. Although efforts are underway to create a standard format for the transmission and storage of biological data, this objective has yet to be fully realized. RESULTS: This work describes an application programming interface (API) that provides a framework for developing an effective biological knowledge ontology for Java-based software projects. The API provides a robust framework for the data acquisition and management needs of an ontology implementation. In addition, the API contains classes to assist in creating GUIs to represent this data visually. CONCLUSIONS: The Knowledge Discovery Object Model (KDOM) API is particularly useful for medium to large applications, or for a number of smaller software projects with common characteristics or objectives. KDOM can be coupled effectively with other biologically relevant APIs and classes. Source code, libraries, documentation and examples are available at http://www.bcgsc.ca/bioinfo/software. Scott D. Zuyderduyn, Steven J. M. Jones |
BMC Bioinform. | 2 |
| 2002 | AcePrimer: automation of PCR primer design based on gene structureabstractAbstract Summary: AcePrimer is an internet-accessed application based on CGI/Perl programming that designs PCR primers to search for deletion alleles in Caenorhabditis elegans gene knockout experiments and uses electronic PCR to search the entire genomic DNA sequence for potential false priming or multiple PCR amplification targets. Features such as the ability to target specific exons with the ‘poison primer’ approach and evaluation of primers with electronic PCR provide a flexible, web-based approach to design effective primers whilst minimizing the need for empirical optimization of PCR experiments. Availability: Web access to this program is provided at http://elegans.bcgsc.bc.ca/gko/aceprimer.shtml. * To whom correspondence should be addressed. Sheldon J. McKay, Steven J. M. Jones |
Bioinform. | 2 |
| 2002 | Assembly of fingerprint contigs: parallelized FPCabstractSUMMARY: One of the more common uses of the program FingerPrint Contigs (FPC) is to assemble random restriction digest 'fingerprints' of overlapping genomic clones into contigs. To improve the rate of assembling contigs from large fingerprint databases we have adapted FPC so that it can be run in parallel on multiple processors and servers. The current version of 'parallelized FPC' has been used in our laboratory to assemble mammalian BAC fingerprint databases, each containing more than 300000 BAC fingerprints. AVAILABILITY: This parallelized version of FPC is available under the GNU GPL licence, and can be downloaded from ftp://ftp.bcgsc.bc.ca/pub/fpcd. Steven R. Ness, W. Terpstra, Martin Krzywinski, Marco A. Marra, Steven J. M. Jones |
Bioinform. | 5 |
| 2001 | PhyloBLAST: facilitating phylogenetic analysis of BLAST resultsabstractAbstract Summary: PhyloBLAST is an internet-accessed application based on CGI/Perl programming that compares a users protein sequence to a SwissProt/TREMBL database using BLAST2 and then allows phylogenetic analyses to be performed on selected sequences from the BLAST output. Flexible features such as ability to input your own multiple sequence alignment and use PHYLIP program options provide additional web-based phylogenetic analysis functionality beyond the analysis of a BLAST result. Availability: This program is available from http://www.pathogenomics.bc.ca/phyloBLAST/ and the source code is freely available from the authors. Contact: [email protected] * To whom correspondence should be addressed. Fiona S. L. Brinkman, Ivan Wan, Robert E. W. Hancock, Ann M. Rose, Steven J. M. Jones |
Bioinform. | 5 |