Michael Brudno

dblp:10/6530 · DBLP profile ↗
← Back
50ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-7947-2243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
28 papers
Bioinformatics and computational biology · 73% Medical and health informatics · 27%
Artificial intelligence
4 papers
Transfer learning and domain adaptation · 28% Representation and self-supervised learning · 28% Information extraction and text analysis · 21%
Computer graphics and multimedia
6 papers
Visualization and visual analytics · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 90% Parallel and multicore computing · 6% Distributed systems · 4%
Human-computer interaction and pervasive computing
2 papers
User interface design and tools · 100%

Topics — the 30 heaviest of 70, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics
text visualization
1.022023
: Navigating large collections of text notes in electronic health records for clinical chart review · IEEE Trans. Vis. Comput. Graph. 2023
Doccurate: A Curation-Based Approach for Clinical Text Visualization · IEEE Trans. Vis. Comput. Graph. 2019
Machine learning › Representation and self-supervised learning › hierarchical representation › hierarchical representation learning
hierarchical latent representation
0.912025
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding · ICCV 2025
Computer vision › 3D vision
implicit neural representation
0.912025
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding · ICCV 2025
Machine learning › Transfer learning and domain adaptation
meta-learning
0.912025
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding · ICCV 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.512021
Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation · NeurIPS 2021
Natural language and speech › Information extraction and text analysis › text classification › low-resource text classification
few-shot text classification
0.512021
Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation · NeurIPS 2021
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning › multi-task representation learning
task representation
0.512021
Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation · NeurIPS 2021
Bioinformatics and computational biology › cancer genomics
gene fusion detection
0.512021
MetaFusion: a high-confidence metacaller for filtering and prioritizing RNA-seq gene fusion candidates · Bioinform. 2021
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis
0.512021
MetaFusion: a high-confidence metacaller for filtering and prioritizing RNA-seq gene fusion candidates · Bioinform. 2021
Bioinformatics and computational biology
sequence alignment
0.432016
deBGA: read alignment with de Bruijn graph-based seed and extension · Bioinform. 2016
PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants · Bioinform. 2012
PROBCONS: Probabilistic Consistency-Based Multiple Alignment of Amino Acid Sequences · AAAI 2004
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection
0.422016
Cell-free DNA fragment-size distribution analysis for non-invasive prenatal CNV prediction · Bioinform. 2016
Probabilistic method for detecting copy number variation in a fetal genome using maternal plasma sequencing · Bioinform. 2014
Medical and health informatics › clinical genomics
non-invasive prenatal testing
0.422016
Cell-free DNA fragment-size distribution analysis for non-invasive prenatal CNV prediction · Bioinform. 2016
Probabilistic method for detecting copy number variation in a fetal genome using maternal plasma sequencing · Bioinform. 2014
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.442013
SCARPA: scaffolding reads with practical algorithms · Bioinform. 2013
Hapsembler: An Assembler for Highly Polymorphic Genomes · RECOMB 2011
Ab Initio Whole Genome Shotgun Assembly with Mated Short Reads · RECOMB 2008
Bioinformatics and computational biology
genomics
0.432013
SCARPA: scaffolding reads with practical algorithms · Bioinform. 2013
Savant: genome browser for high-throughput sequencing data · Bioinform. 2010
MoGUL: Detecting Common Insertions and Deletions in a Population · RECOMB 2010
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412019
Identifying Clinical Terms in Free-Text Notes Using Ontology-Guided Machine Learning · RECOMB 2019
Medical and health informatics
clinical text processing
0.412019
Identifying Clinical Terms in Free-Text Notes Using Ontology-Guided Machine Learning · RECOMB 2019
Medical and health informatics › clinical text processing
clinical text summarization
0.412019
Doccurate: A Curation-Based Approach for Clinical Text Visualization · IEEE Trans. Vis. Comput. Graph. 2019
Bioinformatics and computational biology › sequence analysis
read mapping
0.322016
deBGA: read alignment with de Bruijn graph-based seed and extension · Bioinform. 2016
A report on the 2009 SIG on short read sequencing and algorithms (Short-SIG) · Bioinform. 2009
Bioinformatics and computational biology
comparative genomics
0.332014
GenomeVISTA - an integrated software package for whole-genome alignment and visualization · Bioinform. 2014
Phylo-VISTA: interactive visualization of multiple DNA sequence alignments · Bioinform. 2004
VISTA : visualizing global DNA sequence alignments of arbitrary length · Bioinform. 2000
Bioinformatics and computational biology › genomics › genomic data analysis
cell-free DNA analysis
0.212016
Cell-free DNA fragment-size distribution analysis for non-invasive prenatal CNV prediction · Bioinform. 2016
Bioinformatics and computational biology › genomics › structural variation
structural variant detection
0.222012
PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants · Bioinform. 2012
A robust framework for detecting structural variations in a genome · ISMB 2008
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine cloning
0.222011
SnowFlock: Virtual Machine Cloning as a First-Class Cloud Primitive · ACM Trans. Comput. Syst. 2011
SnowFlock: rapid virtual machine cloning for cloud computing · EuroSys 2009
Bioinformatics and computational biology › sequence analysis
high-throughput sequencing data analysis
0.222010
Savant: genome browser for high-throughput sequencing data · Bioinform. 2010
A report on the 2009 SIG on short read sequencing and algorithms (Short-SIG) · Bioinform. 2009
Bioinformatics and computational biology › sequence alignment › genome alignment
whole-genome alignment
0.212014
GenomeVISTA - an integrated software package for whole-genome alignment and visualization · Bioinform. 2014
Medical and health informatics
clinical genomics
0.212013
Identification of deleterious synonymous variants in human genomes · Bioinform. 2013
Bioinformatics and computational biology › sequence analysis › sequence assembly › genome assembly
scaffolding
0.212013
SCARPA: scaffolding reads with practical algorithms · Bioinform. 2013
Bioinformatics and computational biology
sequence analysis
0.222011
SHRiMP2: Sensitive yet Practical Short Read Mapping · Bioinform. 2011
VISTA : visualizing global DNA sequence alignments of arbitrary length · Bioinform. 2000
Natural language and speech › Information extraction and text analysis
text classification
0.112021
Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation · NeurIPS 2021
Bioinformatics and computational biology › genomics
breakpoint detection
0.112012
PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants · Bioinform. 2012
Bioinformatics and computational biology › sequence analysis › read mapping
split read mapping
0.112012
PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants · Bioinform. 2012

Methods — techniques the papers use, named apart from their topics

qualitative evaluation · 3.1iterative design · 3.0interviews · 2.0meta-learning · 1.4information retrieval · 1.1field study · 1.0evaluation · 1.0residual connections · 0.9frequency encoding · 0.9ontology-guided machine learning · 0.8topic model · 0.7human phenotype ontology · 0.7graph clustering · 0.5gradient-based task representation · 0.5conditional neural process · 0.5ontology topology simplification · 0.3microbenchmarking · 0.1implementation · 0.1
YearPublicationVenuePosition
2026 When It's Hard to Explain: Strategies for Reducing Prompt Uncertainty In Multimodal Generative Systems
abstract
While multimodal generative AI can support creative activities, users often struggle to prompt models to achieve desired aesthetic, acoustic, or stylistic characteristics. Besides, existing generative models are predominantly driven by text-based prompts regardless of their output modality, i.e., using text-based prompts for creating images, videos, and sounds, which often leads to high prompt uncertainty. In response, recent multimodal AI systems introduce interaction strategies to better align model interpretations with users’ creative intent, but the growing variety of strategies makes it hard to judge what works for a given use case. We address this gap with a systematic literature review (n=71) that categorizes prompt-uncertainty-reduction strategies into six types: Guiding Prompt Construction, System Refining of the Prompt, Direct Manipulations of Output Elements, Explaining Reasoning about Prompt Interpretation, Displaying Multiple Outputs, and Controlling Modifier Contribution. For each type, we summarize mechanisms, benefits, and challenges, enabling more efficient navigation of the prompt-support design space.
Nazar Ponochevnyi, Young-Ho Kim, Michael Brudno, Anastasia Kuzminykh
DIS3
2026 SilverLining: Data-First Mitigation of Spatial and Spectral Shortcuts Without Introducing New Confounders
abstract
Deep neural networks exploit shortcuts—spurious correlations like laterality markers (spatial) or scanner-specific noise (spectral)—that severely compromise generalization. Many healthcare applications face multiple concurrent shortcuts that are both spatial and spectral, which existing methods struggle to handle. We present SilverLining, an attention-based preprocessing framework that simultaneously identifies and mitigates both spatial and spectral shortcuts without introducing new spurious correlations. Our key insight is that naive removal of shortcut features can itself create new shortcuts, where models learn to exploit the removal patterns as new spurious correlations. We address this through a principled confounder-free correction strategy that maintains consistent preprocessing patterns across all classes in both spatial and frequency domains, preventing new confounders. Extensive experiments demonstrate SilverLining’s effectiveness: achieving 0.87 AUC on controlled vision tasks and 0.94 AUC on counter-shortcut medical imaging evaluation where shortcuts are reversed; improving cross-institutional chest X-ray classification from 0.72 to 0.77 AUC; and 0.54 mAP on polyp detection despite natural spurious correlations from surgical overlays. Our data-centric approach provides an effective solution for reducing multiple types of data shortcuts without architectural modifications, creating preprocessed datasets that improve model robustness across both classification and detection tasks. Our codebase is available at https://github.com/theidentity/SilverLining_WACV2026/.
Balagopal Unnikrishnan, Michael Brudno, Chris McIntosh
WACV2
2025 LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
abstract
Implicit Neural Representations (INRs) are proving to be a powerful paradigm in unifying task modeling across diverse data domains, offering key advantages such as memory efficiency and resolution independence. Conventional deep learning models are typically modality-dependent, often requiring custom architectures and objectives for different types of signals. However, existing INR frameworks frequently rely on global latent vectors or exhibit computational inefficiencies that limit their broader applicability. We introduce LIFT, a novel, high-performance framework that addresses these challenges by capturing multiscale information through meta-learning. LIFT leverages multiple parallel localized implicit functions alongside a hierarchical latent generator to produce unified latent representations that span local, intermediate, and global features. This architecture facilitates smooth transitions across local regions, enhancing expressivity while maintaining inference efficiency. Additionally, we introduce ReLIFT, an enhanced variant of LIFT that incorporates residual connections and expressive frequency encodings. With this straightforward approach, ReLIFT effectively addresses the convergence-capacity gap found in comparable methods, providing an efficient yet powerful solution to improve capacity and speed up convergence. Empirical results show that LIFT achieves state-of-the-art (SOTA) performance in generative modeling and classification tasks, with notable reductions in computational costs. Moreover, in single-task settings, the streamlined ReLIFT architecture proves effective in signal representations and inverse problem tasks.
Amirhossein Kazerouni, Soroush Mehraban, Michael Brudno, Babak Taati
ICCV3
2025 SUM: Saliency Unification Through Mamba for Visual Attention Modeling
abstract
Visual attention modeling, important for interpreting and prioritizing visual stimuli, plays a significant role in applications such as marketing, multimedia, and robotics. Traditional saliency prediction models, especially those based on Convolutional Neural Networks (CNNs) or Transformers, achieve notable success by leveraging large-scale annotated datasets. However, the current state-of-the-art (SOTA) models that use Transformers are computationally expensive. Additionally, separate models are often required for each image type, lacking a unified approach. In this paper, we propose Saliency Unification through Mamba (SUM), a novel approach that integrates the efficient long-range dependency modeling of Mamba with U-Net to provide a unified model for diverse image types. Using a novel Conditional Visual State Space (C- VSS) block, SUM dynamically adapts to various image types, including natural scenes, web pages, and commercial imagery, ensuring universal applicability across different data types. Our comprehen-sive evaluations across five benchmarks demonstrate that SUM seamlessly adapts to different visual characteristics and consistently outperforms existing models. These results position SUM as a versatile and powerful tool for advancing visual attention modeling, offering a robust solution universally applicable across different types of visual content. Our codebase and pretrained models are publicly accessible on the https://arhosseini77.github.io/sum_page/.
Alireza Hosseini, Amirhossein Kazerouni, Saeed Akhavan, Michael Brudno, Babak Taati
WACV4
2023 Evaluation of single-cell RNAseq labelling algorithms using cancer datasets
abstract
Single-cell RNA sequencing (scRNA-seq) clustering and labelling methods are used to determine precise cellular composition of tissue samples. Automated labelling methods rely on either unsupervised, cluster-based approaches or supervised, cell-based approaches to identify cell types. The high complexity of cancer poses a unique challenge, as tumor microenvironments are often composed of diverse cell subpopulations with unique functional effects that may lead to disease progression, metastasis and treatment resistance. Here, we assess 17 cell-based and 9 cluster-based scRNA-seq labelling algorithms using 8 cancer datasets, providing a comprehensive large-scale assessment of such methods in a cancer-specific context. Using several performance metrics, we show that cell-based methods generally achieved higher performance and were faster compared to cluster-based methods. Cluster-based methods more successfully labelled non-malignant cell types, likely because of a lack of gene signatures for relevant malignant cell subpopulations. Larger cell numbers present in some cell types in training data positively impacted prediction scores for cell-based methods. Finally, we examined which methods performed favorably when trained and tested on separate patient cohorts in scenarios similar to clinical applications, and which were able to accurately label particularly small or under-represented cell populations in the given datasets. We conclude that scPred and SVM show the best overall performances with cancer-specific data and provide further suggestions for algorithm selection. Our analysis pipeline for assessing the performance of cell type labelling algorithms is available in https://github.com/shooshtarilab/scRNAseq-Automated-Cell-Type-Labelling.
Erik Christensen, Andrei L. Turinsky, Mia Husic, Alaina Mahalanabis, Alaine Naidas, Juan Javier Díaz-Mejía, Michael Brudno, Trevor J. Pugh, Arun K. Ramani, Parisa Shooshtari
Briefings Bioinform.8
2023 : Navigating large collections of text notes in electronic health records for clinical chart review
abstract
Before seeing a patient for the first time, healthcare workers will typically conduct a comprehensive clinical chart review of the patient's electronic health record (EHR). Within the diverse documentation pieces included there, text notes are among the most important and thoroughly perused segments for this task; and yet they are among the least supported medium in terms of content navigation and overview. In this work, we delve deeper into the task of clinical chart review from a data visualization perspective and propose a hybrid graphics+text approach via ChartWalk, an interactive tool to support the review of text notes in EHRs. We report on our iterative design process grounded in input provided by a diverse range of healthcare professionals, with steps including: (a) initial requirements distilled from interviews and the literature, (b) an interim evaluation to validate design decisions, and (c) a task-based qualitative evaluation of our final design. We contribute lessons learned to better support the design of tools not only for clinical chart reviews but also other healthcare-related tasks around medical text analysis.
Nicole Sultanum, Farooq Naeem, Michael Brudno, Fanny Chevalier
IEEE Trans. Vis. Comput. Graph.3
2021 Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation
abstract
Large pretrained language models (LMs) like BERT have improved performance in many disparate natural language processing (NLP) tasks. However, fine tuning such models requires a large number of training examples for each target task. Simultaneously, many realistic NLP problems are "few shot", without a sufficiently large training set. In this work, we propose a novel conditional neural process-based approach for few-shot text classification that learns to transfer from other diverse tasks with rich annotation. Our key idea is to represent each task using gradient information from a base model and to train an adaptation network that modulates a text classifier conditioned on the task representation. While previous task-aware few-shot learners represent tasks by input encoding, our novel task representation is more powerful, as the gradient captures input-output relationships of a task. Experimental results show that our approach outperforms traditional fine-tuning, sequential transfer learning, and state-of-the-art meta learning approaches on a collection of diverse few-shot tasks. We further conducted analysis and ablations to justify our design choices.
Jixuan Wang, Kuan-Chieh Wang, Frank Rudzicz, Michael Brudno
NeurIPS4
2021 MetaFusion: a high-confidence metacaller for filtering and prioritizing RNA-seq gene fusion candidates
abstract
MOTIVATION: Current fusion detection tools use diverse calling approaches and provide varying results, making selection of the appropriate tool challenging. Ensemble fusion calling techniques appear promising; however, current options have limited accessibility and function. RESULTS: MetaFusion is a flexible metacalling tool that amalgamates outputs from any number of fusion callers. Individual caller results are standardized by conversion into the new file type Common Fusion Format. Calls are annotated, merged using graph clustering, filtered and ranked to provide a final output of high-confidence candidates. MetaFusion consistently achieves higher precision and recall than individual callers on real and simulated datasets, and reaches up to 100% precision, indicating that ensemble calling is imperative for high-confidence results. MetaFusion uses FusionAnnotator to annotate calls with information from cancer fusion databases and is provided with a Benchmarking Toolkit to calibrate new callers. AVAILABILITY AND IMPLEMENTATION: MetaFusion is freely available at https://github.com/ccmbioinfo/MetaFusion. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michael Apostolides, Mia Husic, Robert Siddaway, Cynthia Hawkins, Andrei L. Turinsky, Michael Brudno, Arun K. Ramani
Bioinform.7
2020 Speaker Diarization with Session-Level Speaker Embedding Refinement Using Graph Neural Networks
abstract
Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be suboptimal for distinguishing speakers locally in a specific meeting session. In this work we present the first use of graph neural networks (GNNs) for the speaker diarization problem, utilizing a GNN to refine speaker embeddings locally using the structural information between speech segments inside each session. The speaker embeddings extracted by a pre-trained model are remapped into a new embedding space, in which the different speakers within a single session are better separated. The model is trained for linkage prediction in a supervised manner by minimizing the difference between the affinity matrix constructed by the refined embeddings and the ground-truth adjacency matrix. Spectral clustering is then applied on top of the refined embeddings. We show that the clustering performance of the refined speaker embeddings outperforms the original embeddings significantly on both simulated and real meeting data, and our system achieves the state-of-the-art result on the NIST SRE 2000 CALLHOME database.
Jixuan Wang, Jian Wu 0027, Ranjani Ramamurthy, Frank Rudzicz, Michael Brudno
ICASSP6
2020 Speaker Attribution with Voice Profiles by Graph-Based Semi-Supervised Learning
abstract
Speaker attribution is required in many real-world applications, such as meeting transcription, where speaker identity is assigned to each utterance according to speaker voice profiles. In this paper, we propose to solve the speaker attribution problem by using graph-based semi-supervised learning methods. A graph of speech segments is built for each session, on which segments from voice profiles are represented by labeled nodes while segments from test utterances are unlabeled nodes. The weight of edges between nodes is evaluated by the similarities between the pretrained speaker embeddings of speech segments. Speaker attribution then becomes a semi-supervised learning problem on graphs, on which two graph-based methods are applied: label propagation (LP) and graph neural networks (GNNs). The proposed approaches are able to utilize the structural information of the graph to improve speaker attribution performance. Experimental results on real meeting data show that the graph based approaches reduce speaker attribution error by up to 68% compared to a baseline speaker identification approach that processes each utterance independently.
Jixuan Wang, Jian Wu 0027, Ranjani Ramamurthy, Frank Rudzicz, Michael Brudno
INTERSPEECH6
2020 Predicting Obstructive Hydronephrosis Based on Ultrasound Alone
Lauren Erdman, Marta Skreta, Mandy Rickard, Carson McLean, Aziz Mezlini, Daniel T. Keefe, Anne-Sophie Blais, Michael Brudno, Armando J. Lorenzo, Anna Goldenberg
MICCAI (3)8
2019 Centroid-based Deep Metric Learning for Speaker Recognition
abstract
Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set and unseen speakers. The latter case corresponds to the few-shot learning task, where a trained model is evaluated on unseen classes. Here, we optimize a speaker embedding model with prototypical network loss (PNL), a state-of-the-art approach for the few-shot image classification task. The resulting embedding model outperforms the state-of-the-art triplet loss based models in both speaker verification and identification tasks, for both seen and unseen speakers.
Jixuan Wang, Kuan-Chieh Wang, Marc T. Law, Frank Rudzicz, Michael Brudno
ICASSP5
2019 Identifying Clinical Terms in Free-Text Notes Using Ontology-Guided Machine Learning
Aryan Arbabi, David R. Adams, Sanja Fidler, Michael Brudno
RECOMB4
2019 Doccurate: A Curation-Based Approach for Clinical Text Visualization
abstract
Before seeing a patient, physicians seek to obtain an overview of the patient's medical history. Text plays a major role in this activity since it represents the bulk of the clinical documentation, but reviewing it quickly becomes onerous when patient charts grow too large. Text visualization methods have been widely explored to manage this large scale through visual summaries that rely on information retrieval algorithms to structure text and make it amenable to visualization. However, the integration with such automated approaches comes with a number of limitations, including significant error rates and the need for healthcare providers to fine-tune algorithms without expert knowledge of their inner mechanics. In addition, several of these approaches obscure or substitute the original clinical text and therefore fail to leverage qualitative and rhetorical flavours of the clinical notes. These drawbacks have limited the adoption of text visualization and other summarization technologies in clinical practice. In this work we present Doccurate, a novel system embodying a curation-based approach for the visualization of large clinical text datasets. Our approach offers automation auditing and customizability to physicians while also preserving and extensively linking to the original text. We discuss findings of a formal qualitative evaluation conducted with 6 domain experts, shedding light onto physicians' information needs, perceived strengths and limitations of automated tools, and the importance of customization while balancing efficiency. We also present use case scenarios to showcase Doccurate's envisioned usage in practice.
Nicole Sultanum, Devin Singh, Michael Brudno, Fanny Chevalier
IEEE Trans. Vis. Comput. Graph.3
2018 More Text Please! Understanding and Supporting the Use of Visualization for Clinical Text Overview
abstract
Clinical practice is heavily reliant on the use of unstructured text to document patient stories due to its expressive and flexible nature. However, a physician's capacity to recover information from text for clinical overview is severely affected when records get longer and time pressure increases. Data visualization strategies have been explored to aid in information retrieval by replacing text with graphical summaries, though often at the cost of omitting important text features. This causes physician mistrust and limits real-world adoption. This work presents our investigation into the role and use of text in clinical practice, and reports on efforts to assess the best of both worlds---text and visualization---to facilitate clinical overview. We report on insights garnered from a field study, and the lessons learned from an iterative design process and evaluation of a text-visualization prototype, MedStory, with 14 medical professionals. The results led to a number of grounded design recommendations to guide visualization design to support clinical text overview.
Nicole Sultanum, Michael Brudno, Daniel J. Wigdor, Fanny Chevalier
CHI2
2018 PhenoLines: Phenotype Comparison Visualizations for Disease Subtyping via Topic Models
abstract
PhenoLines is a visual analysis tool for the interpretation of disease subtypes, derived from the application of topic models to clinical data. Topic models enable one to mine cross-sectional patient comorbidity data (e.g., electronic health records) and construct disease subtypes-each with its own temporally evolving prevalence and co-occurrence of phenotypes-without requiring aligned longitudinal phenotype data for all patients. However, the dimensionality of topic models makes interpretation challenging, and de facto analyses provide little intuition regarding phenotype relevance or phenotype interrelationships. PhenoLines enables one to compare phenotype prevalence within and across disease subtype topics, thus supporting subtype characterization, a task that involves identifying a proposed subtype's dominant phenotypes, ages of effect, and clinical validity. We contribute a data transformation workflow that employs the Human Phenotype Ontology to hierarchically organize phenotypes and aggregate the evolving probabilities produced by topic models. We introduce a novel measure of phenotype relevance that can be used to simplify the resulting topology. The design of PhenoLines was motivated by formative interviews with machine learning and clinical experts. We describe the collaborative design process, distill high-level tasks, and report on initial evaluations with machine learning experts and a medical domain expert. These results suggest that PhenoLines demonstrates promising approaches to support the characterization and optimization of topic models.
Michael Glueck, Mahdi Pakdaman Naeini, Finale Doshi-Velez, Fanny Chevalier, Azam Khan, Daniel J. Wigdor, Michael Brudno
IEEE Trans. Vis. Comput. Graph.7
2017 Size and Texture-Based Classification of Lung Tumors with 3D CNNs
abstract
In this paper, we explore the use of current deep learning methods in the field of computer-aided diagnosis (CAD). Specifically we propose the use of 3D convolutional neural nets (CNN) in classifying lung nodules based off of their appearance in CT scans. We explore the choices of network architectures, learning parameters and problem formulations. Comparing these results to other methods we show that the proposed method has close to perfect performance on the publicly available LIDC dataset, achieving an AUC of 0:9685 and a false positive rate of 0:46% with a true positive rate of 90% where the ground truth is the expert opinion of a radiologist.
Marcus A. Brubaker, Michael Brudno
WACV3
2017 PhenoStacks: Cross-Sectional Cohort Phenotype Comparison Visualizations
abstract
Cross-sectional phenotype studies are used by genetics researchers to better understand how phenotypes vary across patients with genetic diseases, both within and between cohorts. Analyses within cohorts identify patterns between phenotypes and patients (e.g., co-occurrence) and isolate special cases (e.g., potential outliers). Comparing the variation of phenotypes between two cohorts can help distinguish how different factors affect disease manifestation (e.g., causal genes, age of onset, etc.). PhenoStacks is a novel visual analytics tool that supports the exploration of phenotype variation within and between cross-sectional patient cohorts. By leveraging the semantic hierarchy of the Human Phenotype Ontology, phenotypes are presented in context, can be grouped and clustered, and are summarized via overviews to support the exploration of phenotype distributions. The design of PhenoStacks was motivated by formative interviews with genetics researchers: we distil high-level tasks, present an algorithm for simplifying ontology topologies for visualization, and report the results of a deployment evaluation with four expert genetics researchers. The results suggest that PhenoStacks can help identify phenotype patterns, investigate data quality issues, and inform data collection design.
Michael Glueck, Alina Gvozdik, Fanny Chevalier, Azam Khan, Michael Brudno, Daniel J. Wigdor
IEEE Trans. Vis. Comput. Graph.5
2016 Cell-free DNA fragment-size distribution analysis for non-invasive prenatal CNV prediction
abstract
BACKGROUND: Non-invasive detection of aneuploidies in a fetal genome through analysis of cell-free DNA circulating in the maternal plasma is becoming a routine clinical test. Such tests, which rely on analyzing the read coverage or the allelic ratios at single-nucleotide polymorphism (SNP) loci, are not sensitive enough for smaller sub-chromosomal abnormalities due to sequencing biases and paucity of SNPs in a genome. RESULTS: We have developed an alternative framework for identifying sub-chromosomal copy number variations in a fetal genome. This framework relies on the size distribution of fragments in a sample, as fetal-origin fragments tend to be smaller than those of maternal origin. By analyzing the local distribution of the cell-free DNA fragment sizes in each region, our method allows for the identification of sub-megabase CNVs, even in the absence of SNP positions. To evaluate the accuracy of our method, we used a plasma sample with the fetal fraction of 13%, down-sampled it to samples with coverage of 10X-40X and simulated samples with CNVs based on it. Our method had a perfect accuracy (both specificity and sensitivity) for detecting 5 Mb CNVs, and after reducing the fetal fraction (to 11%, 9% and 7%), it could correctly identify 98.82-100% of the 5 Mb CNVs and had a true-negative rate of 95.29-99.76%. AVAILABILITY AND IMPLEMENTATION: Our source code is available on GitHub at https://github.com/compbio-UofT/FSDA CONTACT: : [email protected].
Aryan Arbabi, Ladislav Rampásek, Michael Brudno
Bioinform.3
2016 deBGA: read alignment with de Bruijn graph-based seed and extension
abstract
MOTIVATION: As high-throughput sequencing (HTS) technology becomes ubiquitous and the volume of data continues to rise, HTS read alignment is becoming increasingly rate-limiting, which keeps pressing the development of novel read alignment approaches. Moreover, promising novel applications of HTS technology require aligning reads to multiple genomes instead of a single reference; however, it is still not viable for the state-of-the-art aligners to align large numbers of reads to multiple genomes. RESULTS: We propose de Bruijn Graph-based Aligner (deBGA), an innovative graph-based seed-and-extension algorithm to align HTS reads to a reference genome that is organized and indexed using a de Bruijn graph. With its well-handling of repeats, deBGA is substantially faster than state-of-the-art approaches while maintaining similar or higher sensitivity and accuracy. This makes it particularly well-suited to handle the rapidly growing volumes of sequencing data. Furthermore, it provides a promising solution for aligning reads to multiple genomes and graph-based references in HTS applications. AVAILABILITY AND IMPLEMENTATION: deBGA is available at: https://github.com/hitbc/deBGA CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Bo Liu 0023, Hongzhe Guo, Michael Brudno, Yadong Wang 0001
Bioinform.3
2016 PhenoBlocks: Phenotype Comparison Visualizations
abstract
The differential diagnosis of hereditary disorders is a challenging task for clinicians due to the heterogeneity of phenotypes that can be observed in patients. Existing clinical tools are often text-based and do not emphasize consistency, completeness, or granularity of phenotype reporting. This can impede clinical diagnosis and limit their utility to genetics researchers. Herein, we present PhenoBlocks, a novel visual analytics tool that supports the comparison of phenotypes between patients, or between a patient and the hallmark features of a disorder. An informal evaluation of PhenoBlocks with expert clinicians suggested that the visualization effectively guides the process of differential diagnosis and could reinforce the importance of complete, granular phenotypic reporting.
Michael Glueck, Peter Hamilton, Fanny Chevalier, Simon Breslav, Azam Khan, Daniel J. Wigdor, Michael Brudno
IEEE Trans. Vis. Comput. Graph.7
2015 Identification of deleterious synonymous variants in human genomes
abstract
Bioinformatics (2014); 29(15), 1843–1850 doi: 10.1093/bioinformatics/btt308 In section 3.6 of the above article, an estimate of the human effective population size is listed as Ne ≈104 due to a typographical error. The text should read Ne ≈104.
Orion J. Buske, AshokKumar Manickaraj, Seema Mital, Peter N. Ray, Michael Brudno
Bioinform.5
2014 GenomeVISTA - an integrated software package for whole-genome alignment and visualization
abstract
UNLABELLED: With the ubiquitous generation of complete genome assemblies for a variety of species, efficient tools for whole-genome alignment along with user-friendly visualization are critically important. Our VISTA family of tools for comparative genomics, based on algorithms for pairwise and multiple alignments of genomic sequences and whole-genome assemblies, has become one of the standard techniques for comparative analysis. Most of the VISTA programs have been implemented as Web-accessible servers and are extensively used by the biomedical community. In this manuscript, we introduce GenomeVISTA: a novel implementation that incorporates most features of the VISTA family--fast and accurate alignment, visualization capabilities, GUI and analytical tools within a stand-alone software package. GenomeVISTA thus provides flexibility and security for users who need to conduct whole-genome comparisons on their own computers. AVAILABILITY AND IMPLEMENTATION: Implemented in Perl, C/C++ and Java, the source code is freely available for download at the VISTA Web site: http://genome.lbl.gov/vista/.
Alexander Poliakov, Justin Foong, Michael Brudno, Inna Dubchak
Bioinform.3
2014 Probabilistic method for detecting copy number variation in a fetal genome using maternal plasma sequencing
abstract
MOTIVATION: The past several years have seen the development of methodologies to identify genomic variation within a fetus through the non-invasive sequencing of maternal blood plasma. These methods are based on the observation that maternal plasma contains a fraction of DNA (typically 5-15%) originating from the fetus, and such methodologies have already been used for the detection of whole-chromosome events (aneuploidies), and to a more limited extent for smaller (typically several megabases long) copy number variants (CNVs). RESULTS: Here we present a probabilistic method for non-invasive analysis of de novo CNVs in fetal genome based on maternal plasma sequencing. Our novel method combines three types of information within a unified Hidden Markov Model: the imbalance of allelic ratios at SNP positions, the use of parental genotypes to phase nearby SNPs and depth of coverage to better differentiate between various types of CNVs and improve precision. Our simulation results, based on in silico introduction of novel CNVs into plasma samples with 13% fetal DNA concentration, demonstrate a sensitivity of 90% for CNVs >400 kb (with 13 calls in an unaffected genome), and 40% for 50-400 kb CNVs (with 108 calls in an unaffected genome). AVAILABILITY AND IMPLEMENTATION: Implementation of our model and data simulation method is available at http://github.com/compbio-UofT/fCNV.
Ladislav Rampásek, Aryan Arbabi, Michael Brudno
Bioinform.3
2013 Identification of deleterious synonymous variants in human genomes
abstract
MOTIVATION: The prioritization and identification of disease-causing mutations is one of the most significant challenges in medical genomics. Currently available methods address this problem for non-synonymous single nucleotide variants (SNVs) and variation in promoters/enhancers; however, recent research has implicated synonymous (silent) exonic mutations in a number of disorders. RESULTS: We have curated 33 such variants from literature and developed the Silent Variant Analyzer (SilVA), a machine-learning approach to separate these from among a large set of rare polymorphisms. We evaluate SilVA's performance on in silico 'infection' experiments, in which we implant known disease-causing mutations into a human genome, and show that for 15 of 33 disorders, we rank the implanted mutation among the top five most deleterious ones. Furthermore, we apply the SilVA method to two additional datasets: synonymous variants associated with Meckel syndrome, and a collection of silent variants clinically observed and stratified by a molecular diagnostics laboratory, and show that SilVA is able to accurately predict the harmfulness of silent variants in these datasets. AVAILABILITY: SilVA is open source and is freely available from the project website: http://compbio.cs.toronto.edu/silva CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Orion J. Buske, AshokKumar Manickaraj, Seema Mital, Peter N. Ray, Michael Brudno
Bioinform.5
2013 SCARPA: scaffolding reads with practical algorithms
abstract
MOTIVATION: Scaffolding is the process of ordering and orienting contigs produced during genome assembly. Accurate scaffolding is essential for finishing draft assemblies, as it facilitates the costly and laborious procedures needed to fill in the gaps between contigs. Conventional formulations of the scaffolding problem are intractable, and most scaffolding programs rely on heuristic or approximate solutions, with potentially exponential running time. RESULTS: We present SCARPA, a novel scaffolder, which combines fixed-parameter tractable and bounded algorithms with Linear Programming to produce near-optimal scaffolds. We test SCARPA on real datasets in addition to a simulated diploid genome and compare its performance with several state-of-the-art scaffolders. We show that SCARPA produces longer or similar length scaffolds that are highly accurate compared with other scaffolders. SCARPA is also capable of detecting misassembled contigs and reports them during scaffolding. AVAILABILITY: SCARPA is open source and available from http://compbio.cs.toronto.edu/scarpa.
Nilgun Donmez, Michael Brudno
Bioinform.2
2012 PRISM: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants
abstract
MOTIVATION: The development of high-throughput sequencing technologies has enabled novel methods for detecting structural variants (SVs). Current methods are typically based on depth of coverage or pair-end mapping clusters. However, most of these only report an approximate location for each SV, rather than exact breakpoints. RESULTS: We have developed pair-read informed split mapping (PRISM), a method that identifies SVs and their precise breakpoints from whole-genome resequencing data. PRISM uses a split-alignment approach informed by the mapping of paired-end reads, hence enabling breakpoint identification of multiple SV types, including arbitrary-sized inversions, deletions and tandem duplications. Comparisons to previous datasets and simulation experiments illustrate PRISM's high sensitivity, while PCR validations of PRISM results, including previously uncharacterized variants, indicate an overall precision of ~90%. AVAILABILITY: PRISM is freely available at http://compbio.cs.toronto.edu/prism.
Yadong Wang 0001, Michael Brudno
Bioinform.3
2011 Hapsembler: An Assembler for Highly Polymorphic Genomes
Nilgun Donmez, Michael Brudno
RECOMB2
2011 Clustering with Overlap for Genetic Interaction Networks via Local Search Optimization
Joseph Andrew Whitney, Judice L. Y. Koh, Michael Costanzo, Grant Brown, Charles Boone, Michael Brudno
WABI6
2011 SHRiMP2: Sensitive yet Practical Short Read Mapping
abstract
Abstract Summary: We report on a major update (version 2) of the original SHort Read Mapping Program (SHRiMP). SHRiMP2 primarily targets mapping sensitivity, and is able to achieve high accuracy at a very reasonable speed. SHRiMP2 supports both letter space and color space (AB/SOLiD) reads, enables for direct alignment of paired reads and uses parallel computation to fully utilize multi-core architectures. Availability: SHRiMP2 executables and source code are freely available at: http://compbio.cs.toronto.edu/shrimp/. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Matei David, Misko Dzamba, Dan Lister, Lucian Ilie, Michael Brudno
Bioinform.5
2011 Variant detection and the Autism sequencing project
abstract
Background Early detection of autism can improve the quality of life of affected individuals [1]. Qualitative screening methods continue to improve, but still suffer from low sensitivity despite increasing specificity [2,3]. In collaboration with the Hospital for Sick Children, we are sequencing the exomes of 1 000 individuals with autism in order to discover genetic variants associated with the disorder. Discovery of associated variants can lead to earlier diagnosis and treatment.
Orion J. Buske, Misko Dzamba, Justin Foong, Lynette Lau, Marc Fiume, Christian Marshall, Susan Walker, Aparna Prasad, Michael Brudno
BMC Bioinform.9
2011 SnowFlock: Virtual Machine Cloning as a First-Class Cloud Primitive
abstract
A basic building block of cloud computing is virtualization. Virtual machines (VMs) encapsulate a user’s computing environment and efficiently isolate it from that of other users. VMs, however, are large entities, and no clear APIs exist yet to provide users with programatic, fine-grained control on short time scales. We present SnowFlock, a paradigm and system for cloud computing that introduces VM cloning as a first-class cloud abstraction. VM cloning exploits the well-understood and effective semantics of UNIX fork. We demonstrate multiple usage models of VM cloning: users can incorporate the primitive in their code, can wrap around existing toolchains via scripting, can encapsulate the API within a parallel programming framework, or can use it to load-balance and self-scale clustered servers. VM cloning needs to be efficient to be usable. It must efficiently transmit VM state in order to avoid cloud I/O bottlenecks. We demonstrate how the semantics of cloning aid us in realizing its efficiency: state is propagated in parallel to multiple VM clones, and is transmitted during runtime, allowing for optimizations that substantially reduce the I/O load. We show detailed microbenchmark results highlighting the efficiency of our optimizations, and macrobenchmark numbers demonstrating the effectiveness of the different usage models of SnowFlock.
H. Andrés Lagar-Cavilla, Joseph Andrew Whitney, Roy Bryant, Philip Patchin, Michael Brudno, Eyal de Lara, Stephen M. Rumble, Mahadev Satyanarayanan, Adin Scannell
ACM Trans. Comput. Syst.5
2010 MoGUL: Detecting Common Insertions and Deletions in a Population
Seunghak Lee, Eric P. Xing, Michael Brudno
RECOMB3
2010 Genome variation discovery with high-throughput sequencing data
abstract
The advent of high-throughput sequencing (HTS) technologies is enabling sequencing of human genomes at a significantly lower cost. The availability of these genomes is hoped to enable novel medical diagnostics and treatment, specific to the individual, thus launching the era of personalized medicine. The data currently generated by HTS machines require extensive computational analysis in order to identify genomic variants present in the sequenced individual. In this paper, we overview HTS technologies and discuss several of the plethora of algorithms and tools designed to analyze HTS data, including algorithms for read mapping, as well as methods for identification of single-nucleotide polymorphisms, insertions/deletions and large-scale structural variants and copy-number variants from these mappings.
Adrian V. Dalca, Michael Brudno
Briefings Bioinform.2
2010 VARiD: A variation detection framework for color-space and letter-space platforms
abstract
MOTIVATION: High-throughput sequencing (HTS) technologies are transforming the study of genomic variation. The various HTS technologies have different sequencing biases and error rates, and while most HTS technologies sequence the residues of the genome directly, generating base calls for each position, the Applied Biosystem's SOLiD platform generates dibase-coded (color space) sequences. While combining data from the various platforms should increase the accuracy of variation detection, to date there are only a few tools that can identify variants from color space data, and none that can analyze color space and regular (letter space) data together. RESULTS: We present VARiD--a probabilistic method for variation detection from both letter- and color-space reads simultaneously. VARiD is based on a hidden Markov model and uses the forward-backward algorithm to accurately identify heterozygous, homozygous and tri-allelic SNPs, as well as micro-indels. Our analysis shows that VARiD performs better than the AB SOLiD toolset at detecting variants from color-space data alone, and improves the calls dramatically when letter- and color-space reads are combined. AVAILABILITY: The toolset is freely available at http://compbio.cs.utoronto.ca/varid.
Adrian V. Dalca, Stephen M. Rumble, Samuel Levy, Michael Brudno
Bioinform.4
2010 Savant: genome browser for high-throughput sequencing data
abstract
MOTIVATION: The advent of high-throughput sequencing (HTS) technologies has made it affordable to sequence many individuals' genomes. Simultaneously the computational analysis of the large volumes of data generated by the new sequencing machines remains a challenge. While a plethora of tools are available to map the resulting reads to a reference genome, and to conduct primary analysis of the mappings, it is often necessary to visually examine the results and underlying data to confirm predictions and understand the functional effects, especially in the context of other datasets. RESULTS: We introduce Savant, the Sequence Annotation, Visualization and ANalysis Tool, a desktop visualization and analysis browser for genomic data. Savant was developed for visualizing and analyzing HTS data, with special care taken to enable dynamic visualization in the presence of gigabases of genomic reads and references the size of the human genome. Savant supports the visualization of genome-based sequence, point, interval and continuous datasets, and multiple visualization modes that enable easy identification of genomic variants (including single nucleotide polymorphisms, structural and copy number variants), and functional genomic information (e.g. peaks in ChIP-seq data) in the context of genomic annotations. AVAILABILITY: Savant is freely available at http://compbio.cs.toronto.edu/savant.
Marc Fiume, Vanessa Williams, Andrew Brook, Michael Brudno
Bioinform.4
2009 SnowFlock: rapid virtual machine cloning for cloud computing
abstract
Virtual Machine (VM) fork is a new cloud computing abstraction that instantaneously clones a VM into multiple replicas running on different hosts. All replicas share the same initial state, matching the intuitive semantics of stateful worker creation. VM fork thus enables the straightforward creation and efficient deployment of many tasks demanding swift instantiation of stateful workers in a cloud environment, e.g. excess load handling, opportunistic job placement, or parallel computing. Lack of instantaneous stateful cloning forces users of cloud computing into ad hoc practices to manage application state and cycle provisioning. We present SnowFlock, our implementation of the VM fork abstraction. To evaluate SnowFlock, we focus on the demanding scenario of services requiring on-the-fly creation of hundreds of parallel workers in order to solve computationally-intensive queries in seconds. These services are prominent in fields such as bioinformatics, finance, and rendering. SnowFlock provides sub-second VM cloning, scales to hundreds of workers, consumes few cloud I/O resources, and has negligible runtime overhead.
H. Andrés Lagar-Cavilla, Joseph Andrew Whitney, Adin Scannell, Philip Patchin, Stephen M. Rumble, Eyal de Lara, Michael Brudno, Mahadev Satyanarayanan
EuroSys7
2009 A report on the 2009 SIG on short read sequencing and algorithms (Short-SIG)
abstract
High-throughput sequencing (HTS) technologies are revolutionizing the way biologists acquire and analyze genomic data. HTS instruments, such as the Illumina Genomic Analyzer and the Applied Biosystems SOLiD System, are currently able to sequence tens of gigabases per week, at a cost of 200-fold less than previous methods, potentially enabling the routine sequencing of human and other genomes. Over the last few years the promise of HTS technologies has become a reality, however, realizing that the full promise of these technologies requires the development of computational methods that can analyze the resulting datasets to infer biological meaning. HTS can be used to study many biological problems, including assembling genomes of new organisms, identifying genome variation within a population, discovering novel transcripts, analyzing gene expression, discerning the regulatory mechanisms behind the expression levels and profiling the metagenome of a community. While many HTS datasets are readily available, the main bottleneck in the analysis is the dearth of computational methods that are able to directly answer biologists' questions from these datasets. The Special Interest Group on Short Read Sequencing and Algorithms (Short-SIG), held in conjunction with the Intelligent Systems in Molecular Biology (ISMB) conference, is a meeting that brings together computational biologists interested in analyzing these HTS datasets. The first Short-SIG, held in Toronto in 2008, brought together over 120 attendees, and featured 18 podium presentations, with many of them addressing the computational problems of read mapping—the alignment of reads to a larger reference genome—and assembly—the de novo generation of the genome of an organism from short read data. During the year that followed, significant progress has been made in these fields, and the topic of the 2009 meeting, held in Stockholm on June 28, concentrated on the development of methods that can analyze the resulting read mappings and assemblies to infer biological meaning. The meeting brought together over 200 researchers, and featured 17 platform presentations selected from 27 abstracts and 17 full paper submissions. The keynote address at the SIG was delivered by Dr Edwin Cuppen of the Hubrecht Laboratory (Utrecht, The Netherlands). The paper submissions were handled in coordination with the Bioinformatics journal, and a physical copy of the Bioinformatics ‘virtual issue’ on HTS, featuring papers published on this topic in Bioinformatics over the past year, was presented to all meeting attendees. Bioinformatics also sponsored a best paper award for the conference, given to Kai Ye and his co-authors for the paper ‘Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads’, as well as an award for a paper chosen from the ‘virtual issue’, that was given to Cole Trapnell and colleagues for ‘TopHat: discovering splice junctions with RNA-Seq’. One of the most prominent applications of HTS is the resequencing of human genomes. Genomic variants are discovered by mapping reads from a donor genome to a reference human genome (typically the NCBI assembly), and the resulting mappings are then analyzed to identify differences between the donor and the reference. The SIG saw eight presentations on variation identification, including variants of all sizes—from SNPs to larger, structural variants. From the SNP discovery perspective, Adrian Dalca presented VARiD, a generalized framework for calling SNPs from both regular (letter-space) reads and di-base encoded (color-space) data. Sohrab Shah presented SNVmix, a Bayesian mixture model-based method for discovering single nucleotide variants from somatic tissues, where the observed alleles and their ratios may vary due to adjacent tissues being present in a biopsy. There were also presentations on variation discovery methods from the two leading HTS technology manufacturers, Illumina and Life Technologies (Applied Biosystems). Dirk Evers of Illumina spoke of recent improvements to the CASAVA framework that allows for more accurate discovery of variants from long reads, especially indel and copy number variants. He presented the results of CASAVA on recently sequenced paired tumor/normal genomes from a melanoma cell line. Fiona C. L. Hyland, from Life Technologies, described diBayes, a Bayesian framework for SNP discovery from color-space data, and presented an extensive analysis of a SOLiD human dataset, including discovered SNPs, small indels and larger structural variants. The advent of high-throughput sequencing has for the first time allowed large, cost-effective studies to detect larger, structural variants. Such variation has been associated with numerous diseases, including autism, schizophrenia and cancer, making their discovery an important challenge for computational biologists. Several talks at the SIG focused on the development of novel methods for discovery of such variants. Kai Ye presented a method called Pindel, which, by anchoring the mates of nonmapping reads to a genomic location, was able to use split-mapping to detect deletion events as large as 10 kb with base-level precision (Ye et al., 2009). This paper was the winner of the best SIG paper award, sponsored by Bioinformatics. Seunghak Lee and Weldon Whitener showed how to detect smaller indels from pair-end data by using the distribution of insert sizes of all matepairs that span each genomic location. Paul Medvedev described a way that the depth-of-coverage signal can be combined with pair-end mapping-based techniques to detect copy number variants within segmental duplications. Overall, this year has seen the detection of structural variation come to the forefront of algorithmic research, and the next year will hopefully bring about more fully developed biologist-friendly tools. Another exciting application of HTS technologies is RNA sequencing. RNA sequencing is currently used for several applications, including RNA expression, de novo transcriptome sequencing for nonmodel organisms and novel transcript discovery; however, computational methods for the analysis of this data are in their infancy. For RNA and microRNA expression profiling, HTS has significant advantages compared with microarray methods in that it is better able to identify quantities of very common and very rare transcripts. Short-SIG featured five talks addressing various computational problems in RNA sequencing. Cole Trapnell presented his work on BowTie (Trapnell et al., 2009), a tool to map reads from RNA sequencing to a reference genome, while allowing for split-reads where the two ends of a read map in different locations (due to exon splicing). While the original paper was published as part of the ‘virtual issue’, the presentation included new improvements to the tool. Jan Prins presented MapSplice, a RNA mapping tool that is similar to TopHat, but includes the ability to consider noncanonical splice sites. Inanc Birol presented a version of the ABySS assembler for de novo mRNA assembly. ABySS was the first tool to attempt de novo assembly of the human genome, and in their presentation they presented the first results on de novo assembly of human transcriptome data. Finally, three presentations demonstrated methods to mine RNA-seq data for specific biomedical phenomena: Regina Bohnert presented an algorithm for identifying alternative transcripts and their expression levels, Chol-Hee Jung presented an analysis of combining multiple Drosophila RNA-seq datasets in order to discover novel noncoding RNAs, and Gerald Quon showed that using the ISOLATE framework (Quon and Morris, 2009), mRNA expression levels can be used to identify the tissue of origin in metastasized tumors. The final session of the SIG was devoted to a variety of classical and newly upcoming HTS applications. Bas Dutilh presented a method for mapping metagenomic reads to a reference genome, where the reference is changed during the mapping process to more accurately represent the community consensus genome, thus allowing a larger fraction of reads to map (Dutilh et al., 2009). Juliane Klein presented LOCAS, an assembler for short read data that is targeted toward low-coverage datasets, and significantly outperforms previous methods in this context. The last two presentations addressed the statistical issues underlying HTS. Su Yeon Kim described statistical foundation of designing association studies with HTS, specifically the use of a combination of pooled and unpooled samples from a number of individuals to design association studies. Adam Kowalczyk showed that it is possible to develop univariate statistical tests to compute the likelihood that two distributions of short read datasets are identical (P-values) based on the Poisson approximation to the binomial distribution. The SIG ended with a keynote address by Dr Edwin Cuppen of the Hubrecht Laboratory, in Utrecht, The Netherlands. His presentation demonstrated both some interesting advantage of short read sequencing, such as the ability of CHiP-seq experiments to identify which genes are regulated by specific distal enhancers, and some key limitations, for example, that RNA-seq, while capable of profiling relative transcript levels in different conditions, is unable to reconstruct actual transcript levels due to biases introduced during sample preparation. In addition to the podium presentations, many Short-SIG attendees used the opportunity to discuss collaborations and the general direction of the field. Clearly, the increase in read length (only 25–35 bp 2 years ago and 50–100 bp today) is making it difficult to develop timely tools, as the problems associated with different length reads are quite dissimilar. Illumina and SOLiD reads will very soon be as long as 454 reads were a few years ago, and this dynamics is forcing bioinformaticians to rethink algorithms developed only a year ago. Similarly, the increasing throughput of the sequencing platforms is requiring the scaling of the algorithms to larger datasets. Bioinformatics remains one of the key bottlenecks in HTS data analysis, with datasets created at a faster rate than can be effectively analyzed, and few tools providing ‘one stop shopping’ for the complete analysis of a single dataset. Addressing these shortcomings is a key step to realizing the full promise of HTS technologies. Conflict of Interest: none declared.
Michael Brudno, Paul Medvedev, Jens Stoye, Francisco M. de la Vega
Bioinform.1
2009 SHRiMP: Accurate Mapping of Short Color-space Reads
abstract
The development of Next Generation Sequencing technologies, capable of sequencing hundreds of millions of short reads (25-70 bp each) in a single run, is opening the door to population genomic studies of non-model species. In this paper we present SHRiMP - the SHort Read Mapping Package: a set of algorithms and methods to map short reads to a genome, even in the presence of a large amount of polymorphism. Our method is based upon a fast read mapping technique, separate thorough alignment methods for regular letter-space as well as AB SOLiD (color-space) reads, and a statistical model for false positive hits. We use SHRiMP to map reads from a newly sequenced Ciona savignyi individual to the reference genome. We demonstrate that SHRiMP can accurately map reads to this highly polymorphic genome, while confirming high heterozygosity of C. savignyi in this second individual. SHRiMP is freely available at http://compbio.cs.toronto.edu/shrimp.
Stephen M. Rumble, Phil Lacroute, Adrian V. Dalca, Marc Fiume, Arend Sidow, Michael Brudno
PLoS Comput. Biol.6
2008 A robust framework for detecting structural variations in a genome
abstract
MOTIVATION: Recently, structural genomic variants have come to the forefront as a significant source of variation in the human population, but the identification of these variants in a large genome remains a challenge. The complete sequencing of a human individual is prohibitive at current costs, while current polymorphism detection technologies, such as SNP arrays, are not able to identify many of the large scale events. One of the most promising methods to detect such variants is the computational mapping of clone-end sequences to a reference genome. RESULTS: Here, we present a probabilistic framework for the identification of structural variants using clone-end sequencing. Unlike previous methods, our approach does not rely on an a priori determined mapping of all reads to the reference. Instead, we build a framework for finding the most probable assignment of sequenced clones to potential structural variants based on the other clones. We compare our predictions with the structural variants identified in three previous studies. While there is a statistically significant correlation between the predictions, we also find a significant number of previously uncharacterized structural variants. Furthermore, we identify a number of putative cross-chromosomal events, primarily located proximally to the centromeres of the chromosomes. AVAILABILITY: Our dataset, results and source code are available at http://compbio.cs.toronto.edu/structvar/.
Seunghak Lee, Elango Cheran, Michael Brudno
ISMB3
2008 A mixture model for the evolution of gene expression in non-homogeneous datasets
abstract
We address the challenge of assessing conservation of gene expression in complex, non-homogeneous datasets. Recent studies have demonstrated the success of probabilistic models in studying the evolution of gene expression in simple eukaryotic organisms such as yeast, for which measurements are typically scalar and independent. Models capable of studying expression evolution in much more complex organisms such as vertebrates are particularly important given the medical and scientific interest in species such as human and mouse. We present a statistical model that makes a number of significant extensions to previous models to enable characterization of changes in expression among highly complex organisms. We demonstrate the efficacy of our method on a microarray dataset containing diverse tissues from multiple vertebrate species. We anticipate that the model will be invaluable in the study of gene expression patterns in other diverse organisms as well, such as worms and insects.
Gerald T. Quon, Yee Whye Teh, Esther T. Chan, Timothy R. Hughes, Michael Brudno, Quaid Morris
NIPS5
2008 Ab Initio Whole Genome Shotgun Assembly with Mated Short Reads
Paul Medvedev, Michael Brudno
RECOMB2
2008 Read Mapping Algorithms for Single Molecule Sequencing Data
Vladimir Yanovski, Stephen M. Rumble, Michael Brudno
WABI3
2007 Computability of Models for Sequence Assembly
Paul Medvedev, Konstantinos Georgiou, Eugene W. Myers, Michael Brudno
WABI4
2004 PROBCONS: Probabilistic Consistency-Based Multiple Alignment of Amino Acid Sequences
Chuong B. Do, Michael Brudno, Serafim Batzoglou
AAAI2
2004 Chaining Algorithms for Alignment of Draft Sequence
Mukund Sundararajan, Michael Brudno, Kerrin S. Small, Arend Sidow, Serafim Batzoglou
WABI2
2004 Phylo-VISTA: interactive visualization of multiple DNA sequence alignments
abstract
MOTIVATION: The power of multi-sequence comparison for biological discovery is well established. The need for new capabilities to visualize and compare cross-species alignment data is intensified by the growing number of genomic sequence datasets being generated for an ever-increasing number of organisms. To be efficient these visualization algorithms must support the ability to accommodate consistently a wide range of evolutionary distances in a comparison framework based upon phylogenetic relationships. RESULTS: We have developed Phylo-VISTA, an interactive tool for analyzing multiple alignments by visualizing a similarity measure for multiple DNA sequences. The complexity of visual presentation is effectively organized using a framework based upon interspecies phylogenetic relationships. The phylogenetic organization supports rapid, user-guided interspecies comparison. To aid in navigation through large sequence datasets, Phylo-VISTA leverages concepts from VISTA that provide a user with the ability to select and view data at varying resolutions. The combination of multiresolution data visualization and analysis, combined with the phylogenetic framework for interspecies comparison, produces a highly flexible and powerful tool for visual data analysis of multiple sequence alignments. AVAILABILITY: Phylo-VISTA is available at http://www-gsd.lbl.gov/phylovista. It requires an Internet browser with Java Plug-in 1.4.2 and it is integrated into the global alignment program LAGAN at http://lagan.stanford.edu
Nameeta Y. Shah, Olivier Couronne, Len A. Pennacchio, Michael Brudno, Serafim Batzoglou, E. Wes Bethel, Edward M. Rubin, Bernd Hamann, Inna Dubchak
Bioinform.4
2003 AGenDA: homology-based gene prediction
abstract
Abstract Summary: We present a www server for homology-based gene prediction. The user enters a pair of evolutionary related genomic sequences, for example from human and mouse. Our software system uses CHAOS and DIALIGN to calculate an alignment of the input sequences and then searches for conserved splicing signals and start/stop codons around regions of local sequence similarity. This way, candidate exons are identified that are used, in turn, to calculate optimal gene models. The server returns the constructed gene model by email, together with a graphical representation of the underlying genomic alignment. Availability: http://bibiserv.TechFak.Uni-Bielefeld.DE/agenda/ Contact: [email protected] * To whom correspondence should be addressed.
Leila Taher, Oliver Rinner, Alexander Sczyrba, Michael Brudno, Serafim Batzoglou, Burkhard Morgenstern
Bioinform.5
2003 Fast and sensitive multiple alignment of large genomic sequences
abstract
BACKGROUND: Genomic sequence alignment is a powerful method for genome analysis and annotation, as alignments are routinely used to identify functional sites such as genes or regulatory elements. With a growing number of partially or completely sequenced genomes, multiple alignment is playing an increasingly important role in these studies. In recent years, various tools for pair-wise and multiple genomic alignment have been proposed. Some of them are extremely fast, but often efficiency is achieved at the expense of sensitivity. One way of combining speed and sensitivity is to use an anchored-alignment approach. In a first step, a fast search program identifies a chain of strong local sequence similarities. In a second step, regions between these anchor points are aligned using a slower but more accurate method. RESULTS: Herein, we present CHAOS, a novel algorithm for rapid identification of chains of local pair-wise sequence similarities. Local alignments calculated by CHAOS are used as anchor points to improve the running time of DIALIGN, a slow but sensitive multiple-alignment tool. We show that this way, the running time of DIALIGN can be reduced by more than 95% for BAC-sized and longer sequences, without affecting the quality of the resulting alignments. We apply our approach to a set of five genomic sequences around the stem-cell-leukemia (SCL) gene and demonstrate that exons and small regulatory elements can be identified by our multiple-alignment procedure. CONCLUSION: We conclude that the novel CHAOS local alignment tool is an effective way to significantly speed up global alignment tools such as DIALIGN without reducing the alignment quality. We likewise demonstrate that the DIALIGN/CHAOS combination is able to accurately align short regulatory sequences in distant orthologues.
Michael Brudno, Michael Chapman, Berthold Göttgens, Serafim Batzoglou, Burkhard Morgenstern
BMC Bioinform.1
2000 VISTA : visualizing global DNA sequence alignments of arbitrary length
abstract
SUMMARY: VISTA is a program for visualizing global DNA sequence alignments of arbitrary length. It has a clean output, allowing for easy identification of similarity, and is easily configurable, enabling the visualization of alignments of various lengths at different levels of resolution. It is currently available on the web, thus allowing for easy access by all researchers. AVAILABILITY: VISTA server is available on the web at http://www-gsd.lbl.gov/vista. The source code is available upon request. CONTACT: [email protected]
Chris Mayor, Michael Brudno, Jody R. Schwartz, Alexander Poliakov, Edward M. Rubin, Kelly A. Frazer, Lior Pachter, Inna Dubchak
Bioinform.2