VLDB 2026 Research / reviewers in the wild / expert
Olga G. Troyanskaya
dblp:09/2486
· DBLP profile ↗
45ranked-venue papers
5as first author
1since 2021 · last 2021
0000-0002-5676-5737ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
19 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 77% Deep learning architectures and training · 23% |
Topics — the 30 heaviest of 37, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
cancer genomics |
0.5 | 2 | 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterations · Bioinform. 2020 Aneuploidy prediction and tumor classification with heterogeneous hidden conditional random fields · Bioinform. 2009 |
Bioinformatics and computational biology › gene regulation
gene regulatory network |
0.5 | 2 | 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterations · Bioinform. 2020 Detailing regulatory networks through large scale data integration · Bioinform. 2009 |
Bioinformatics and computational biology › cancer genomics › breast cancer
breast cancer subtype analysis |
0.4 | 1 | 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterations · Bioinform. 2020 |
Bioinformatics and computational biology
gene expression analysis |
0.2 | 3 | 2013 | Ontology-aware classification of tissue and cell-type signals in gene expression profiles across platforms and technologies · Bioinform. 2013 Nonparametric methods for identifying differentially expressed genes in microarray data · Bioinform. 2002 Missing value estimation methods for DNA microarrays · Bioinform. 2001 |
Bioinformatics and computational biology
functional genomics |
0.2 | 3 | 2008 | Assessing the functional structure of genomic data · ISMB 2008 Context-sensitive data integration and prediction of biological networks · Bioinform. 2007 Hierarchical multi-label prediction of gene function · Bioinform. 2006 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent generative model
generative stochastic networks |
0.2 | 1 | 2014 | Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction · ICML 2014 |
Bioinformatics and computational biology
protein structure prediction |
0.2 | 1 | 2014 | Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction · ICML 2014 |
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction |
0.2 | 1 | 2014 | Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction · ICML 2014 |
Bioinformatics and computational biology › gene expression analysis
sample classification |
0.2 | 1 | 2013 | Ontology-aware classification of tissue and cell-type signals in gene expression profiles across platforms and technologies · Bioinform. 2013 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.1 | 1 | 2012 | An effective statistical evaluation of ChIPseq dataset similarity · Bioinform. 2012 |
Bioinformatics and computational biology › cancer genomics
copy number analysis |
0.1 | 2 | 2009 | Aneuploidy prediction and tumor classification with heterogeneous hidden conditional random fields · Bioinform. 2009 Accurate detection of aneuploidies in array CGH and gene expression microarray data · Bioinform. 2004 |
Bioinformatics and computational biology › epigenomics
DNA methylation |
0.1 | 1 | 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterations · Bioinform. 2020 |
Bioinformatics and computational biology
epigenomics |
0.1 | 1 | 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterations · Bioinform. 2020 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.1 | 2 | 2006 | A scalable method for integration and functional analysis of multiple microarray datasets · Bioinform. 2006 Accurate detection of aneuploidies in array CGH and gene expression microarray data · Bioinform. 2004 |
Bioinformatics and computational biology › data integration
bayesian data integration |
0.1 | 2 | 2009 | Context-sensitive data integration and prediction of biological networks · Bioinform. 2007 Detailing regulatory networks through large scale data integration · Bioinform. 2009 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
array CGH analysis |
0.1 | 1 | 2009 | Aneuploidy prediction and tumor classification with heterogeneous hidden conditional random fields · Bioinform. 2009 |
Bioinformatics and computational biology › cancer genomics
cancer classification |
0.1 | 1 | 2009 | Aneuploidy prediction and tumor classification with heterogeneous hidden conditional random fields · Bioinform. 2009 |
Bioinformatics and computational biology
protein function prediction |
0.1 | 1 | 2009 | The impact of incomplete knowledge on evaluation: an experimental benchmark for protein function prediction · Bioinform. 2009 |
Bioinformatics and computational biology
gene expression |
0.1 | 1 | 2007 | Exploring the functional landscape of gene expression: directed search of large microarray compendia · Bioinform. 2007 |
Bioinformatics and computational biology › functional genomics
gene function prediction |
0.1 | 1 | 2006 | Hierarchical multi-label prediction of gene function · Bioinform. 2006 |
Bioinformatics and computational biology
hierarchical multi-label classification |
0.1 | 1 | 2006 | Hierarchical multi-label prediction of gene function · Bioinform. 2006 |
Bioinformatics and computational biology › data integration
multi-dataset integration |
0.1 | 1 | 2006 | A scalable method for integration and functional analysis of multiple microarray datasets · Bioinform. 2006 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2014 | Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure Prediction · ICML 2014 |
Bioinformatics and computational biology › clinical bioinformatics
aneuploidy detection |
0.0 | 1 | 2004 | Accurate detection of aneuploidies in array CGH and gene expression microarray data · Bioinform. 2004 |
Bioinformatics and computational biology › cancer genomics
chromosomal aberration detection |
0.0 | 1 | 2004 | Accurate detection of aneuploidies in array CGH and gene expression microarray data · Bioinform. 2004 |
Bioinformatics and computational biology › epigenomics
chromatin analysis |
0.0 | 1 | 2012 | An effective statistical evaluation of ChIPseq dataset similarity · Bioinform. 2012 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein-DNA interaction |
0.0 | 1 | 2012 | An effective statistical evaluation of ChIPseq dataset similarity · Bioinform. 2012 |
Bioinformatics and computational biology › gene expression analysis › differential expression analysis
differentially expressed gene identification |
0.0 | 1 | 2002 | Nonparametric methods for identifying differentially expressed genes in microarray data · Bioinform. 2002 |
Bioinformatics and computational biology
genomics |
0.0 | 1 | 2002 | Sequence complexity profiles of prokaryotic genomic sequences: A fast algorithm for calculating linguistic complexity · Bioinform. 2002 |
Bioinformatics and computational biology › sequence analysis
sequence complexity analysis |
0.0 | 1 | 2002 | Sequence complexity profiles of prokaryotic genomic sequences: A fast algorithm for calculating linguistic complexity · Bioinform. 2002 |
Methods — techniques the papers use, named apart from their topics
neural architecture search · 0.5multi-task learning · 0.5convolutional neural network · 0.5motif analysis · 0.4copy number analysis · 0.4ChIP-seq · 0.4markov chain · 0.4generative stochastic network · 0.4convolutional architecture · 0.4functional genomic data integration · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | CROTON: an automated and variant-aware deep learning framework for predicting CRISPR/Cas9 editing outcomesabstractMOTIVATION: CRISPR/Cas9 is a revolutionary gene-editing technology that has been widely utilized in biology, biotechnology and medicine. CRISPR/Cas9 editing outcomes depend on local DNA sequences at the target site and are thus predictable. However, existing prediction methods are dependent on both feature and model engineering, which restricts their performance to existing knowledge about CRISPR/Cas9 editing. RESULTS: Herein, deep multi-task convolutional neural networks (CNNs) and neural architecture search (NAS) were used to automate both feature and model engineering and create an end-to-end deep-learning framework, CROTON (CRISPR Outcomes Through cONvolutional neural networks). The CROTON model architecture was tuned automatically with NAS on a synthetic large-scale construct-based dataset and then tested on an independent primary T cell genomic editing dataset. CROTON outperformed existing expert-designed models and non-NAS CNNs in predicting 1 base pair insertion and deletion probability as well as deletion and frameshift frequency. Interpretation of CROTON revealed local sequence determinants for diverse editing outcomes. Finally, CROTON was utilized to assess how single nucleotide variants (SNVs) affect the genome editing outcomes of four clinically relevant target genes: the viral receptors ACE2 and CCR5 and the immune checkpoint inhibitors CTLA4 and PDCD1. Large SNV-induced differences in CROTON predictions in these target genes suggest that SNVs should be taken into consideration when designing widely applicable gRNAs. AVAILABILITY AND IMPLEMENTATION: https://github.com/vli31/CROTON. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victoria R. Li, Zijun Zhang 0004, Olga G. Troyanskaya |
Bioinform. | 3 |
| 2020 | Subtype-specific transcriptional regulators in breast tumors subjected to genetic and epigenetic alterationsabstractMOTIVATION: Breast cancer consists of multiple distinct tumor subtypes, and results from epigenetic and genetic aberrations that give rise to distinct transcriptional profiles. Despite previous efforts to understand transcriptional deregulation through transcription factor networks, the transcriptional mechanisms leading to subtypes of the disease remain poorly understood. RESULTS: We used a sophisticated computational search of thousands of expression datasets to define extended signatures of distinct breast cancer subtypes. Using ENCODE ChIP-seq data of surrogate cell lines and motif analysis we observed that these subtypes are determined by a distinct repertoire of lineage-specific transcription factors. Furthermore, specific pattern and abundance of copy number and DNA methylation changes at these TFs and targets, compared to other genes and to normal cells were observed. Overall, distinct transcriptional profiles are linked to genetic and epigenetic alterations at lineage-specific transcriptional regulators in breast cancer subtypes. AVAILABILITY AND IMPLEMENTATION: The analysis code and data are deposited at https://bitbucket.org/qzhu/breast.cancer.tf/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qian Zhu 0005, Xavier Tekpli, Olga G. Troyanskaya, Vessela N. Kristensen |
Bioinform. | 3 |
| 2018 | A loop-counting method for covariate-corrected low-rank biclustering of gene-expression and genome-wide association study dataabstractA common goal in data-analysis is to sift through a large data-matrix and detect any significant submatrices (i.e., biclusters) that have a low numerical rank. We present a simple algorithm for tackling this biclustering problem. Our algorithm accumulates information about 2-by-2 submatrices (i.e., 'loops') within the data-matrix, and focuses on rows and columns of the data-matrix that participate in an abundance of low-rank loops. We demonstrate, through analysis and numerical-experiments, that this loop-counting method performs well in a variety of scenarios, outperforming simple spectral methods in many situations of interest. Another important feature of our method is that it can easily be modified to account for aspects of experimental design which commonly arise in practice. For example, our algorithm can be modified to correct for controls, categorical- and continuous-covariates, as well as sparsity within the data. We demonstrate these practical features with two examples; the first drawn from gene-expression analysis and the second drawn from a much larger genome-wide-association-study (GWAS). Aaditya V. Rangan, Caroline C. McGrouther, John Kelsoe, Nicholas J. Schork, Eli Stahl, Qian Zhu 0005, Arjun Krishnan, Victoria Yao, Olga G. Troyanskaya, Seda Bilaloglu, Preeti Raghavan, Sarah Bergen, Anders Juréus, Mikael Landen |
PLoS Comput. Biol. | 9 |
| 2015 | Tissue-aware data integration approach for the inference of pathway interactions in metazoan organismsabstractMOTIVATION: Leveraging the large compendium of genomic data to predict biomedical pathways and specific mechanisms of protein interactions genome-wide in metazoan organisms has been challenging. In contrast to unicellular organisms, biological and technical variation originating from diverse tissues and cell-lineages is often the largest source of variation in metazoan data compendia. Therefore, a new computational strategy accounting for the tissue heterogeneity in the functional genomic data is needed to accurately translate the vast amount of human genomic data into specific interaction-level hypotheses. RESULTS: We developed an integrated, scalable strategy for inferring multiple human gene interaction types that takes advantage of data from diverse tissue and cell-lineage origins. Our approach specifically predicts both the presence of a functional association and also the most likely interaction type among human genes or its protein products on a whole-genome scale. We demonstrate that directly incorporating tissue contextual information improves the accuracy of our predictions, and further, that such genome-wide results can be used to significantly refine regulatory interactions from primary experimental datasets (e.g. ChIP-Seq, mass spectrometry). AVAILABILITY AND IMPLEMENTATION: An interactive website hosting all of our interaction predictions is publically available at http://pathwaynet.princeton.edu. Software was implemented using the open-source Sleipnir library, which is available for download at https://bitbucket.org/libsleipnir/libsleipnir.bitbucket.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Christopher Y. Park, Arjun Krishnan, Qian Zhu 0005, Aaron K. Wong, Olga G. Troyanskaya |
Bioinform. | 6 |
| 2014 | Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure PredictionabstractPredicting protein secondary structure is a fundamental problem in protein structure prediction. Here we present a new supervised generative stochastic network (GSN) based method to predict local secondary structure with deep hierarchical representations. GSN is a recently proposed deep learning technique (Bengio & Thibodeau-Laufer, 2013) to globally train deep generative model. We present the supervised extension of GSN, which learns a Markov chain to sample from a conditional distribution, and applied it to protein structure prediction. To scale the model to full-sized, high-dimensional data, like protein sequences with hundreds of amino-acids, we introduce a convolutional architecture, which allows efficient learning across multiple layers of hierarchical representations. Our architecture uniquely focuses on predicting structured low-level labels informed with both low and high-level representations learned by the model. In our application this corresponds to labeling the secondary structure state of each amino-acid residue. We trained and tested the model on separate sets of non-homologous proteins sharing less than 30% sequence identity. Our model achieves 66.4% Q8 accuracy on the CB513 dataset, better than the previously reported best performance 64.9% (Wang et al., 2011) for this challenging secondary structure prediction problem. Olga G. Troyanskaya |
ICML | 2 |
| 2014 | Global Quantitative Modeling of Chromatin Factor InteractionsabstractChromatin is the driver of gene regulation, yet understanding the molecular interactions underlying chromatin factor combinatorial patterns (or the "chromatin codes") remains a fundamental challenge in chromatin biology. Here we developed a global modeling framework that leverages chromatin profiling data to produce a systems-level view of the macromolecular complex of chromatin. Our model ultilizes maximum entropy modeling with regularization-based structure learning to statistically dissect dependencies between chromatin factors and produce an accurate probability distribution of chromatin code. Our unsupervised quantitative model, trained on genome-wide chromatin profiles of 73 histone marks and chromatin proteins from modENCODE, enabled making various data-driven inferences about chromatin profiles and interactions. We provided a highly accurate predictor of chromatin factor pairwise interactions validated by known experimental evidence, and for the first time enabled higher-order interaction prediction. Our predictions can thus help guide future experimental studies. The model can also serve as an inference engine for predicting unknown chromatin profiles--we demonstrated that with this approach we can leverage data from well-characterized cell types to help understand less-studied cell type or conditions. Olga G. Troyanskaya |
PLoS Comput. Biol. | 2 |
| 2013 | Ontology-aware classification of tissue and cell-type signals in gene expression profiles across platforms and technologiesabstractAbstract Motivation: Leveraging gene expression data through large-scale integrative analyses for multicellular organisms is challenging because most samples are not fully annotated to their tissue/cell-type of origin. A computational method to classify samples using their entire gene expression profiles is needed. Such a method must be applicable across thousands of independent studies, hundreds of gene expression technologies and hundreds of diverse human tissues and cell-types. Results: We present Unveiling RNA Sample Annotation (URSA) that leverages the complex tissue/cell-type relationships and simultaneously estimates the probabilities associated with hundreds of tissues/cell-types for any given gene expression profile. URSA provides accurate and intuitive probability values for expression profiles across independent studies and outperforms other methods, irrespective of data preprocessing techniques. Moreover, without re-training, URSA can be used to classify samples from diverse microarray platforms and even from next-generation sequencing technology. Finally, we provide a molecular interpretation for the tissue and cell-type models as the biological basis for URSA’s classifications. Availability and implementation: An interactive web interface for using URSA for gene expression analysis is available at: ursa.princeton.edu. The source code is available at https://bitbucket.org/youngl/ursa_backend. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Arjun Krishnan, Qian Zhu 0005, Olga G. Troyanskaya |
Bioinform. | 4 |
| 2013 | Functional Knowledge Transfer for High-accuracy Prediction of Under-studied Biological ProcessesabstractA key challenge in genetics is identifying the functional roles of genes in pathways. Numerous functional genomics techniques (e.g. machine learning) that predict protein function have been developed to address this question. These methods generally build from existing annotations of genes to pathways and thus are often unable to identify additional genes participating in processes that are not already well studied. Many of these processes are well studied in some organism, but not necessarily in an investigator's organism of interest. Sequence-based search methods (e.g. BLAST) have been used to transfer such annotation information between organisms. We demonstrate that functional genomics can complement traditional sequence similarity to improve the transfer of gene annotations between organisms. Our method transfers annotations only when functionally appropriate as determined by genomic data and can be used with any prediction algorithm to combine transferred gene function knowledge with organism-specific high-throughput data to enable accurate function prediction. We show that diverse state-of-art machine learning algorithms leveraging functional knowledge transfer (FKT) dramatically improve their accuracy in predicting gene-pathway membership, particularly for processes with little experimental knowledge in an organism. We also show that our method compares favorably to annotation transfer by sequence similarity. Next, we deploy FKT with state-of-the-art SVM classifier to predict novel genes to 11,000 biological processes across six diverse organisms and expand the coverage of accurate function predictions to processes that are often ignored because of a dearth of annotated genes in an organism. Finally, we perform in vivo experimental investigation in Danio rerio and confirm the regulatory role of our top predicted novel gene, wnt5b, in leftward cell migration during heart development. FKT is immediately applicable to many bioinformatics techniques and will help biologists systematically integrate prior knowledge from diverse systems to direct targeted experiments in their organism of study. Christopher Y. Park, Aaron K. Wong, Casey S. Greene, Jessica Rowland, Yuanfang Guan, Lars Ailo Bongo, Rebecca D. Burdine, Olga G. Troyanskaya |
PLoS Comput. Biol. | 8 |
| 2012 | An effective statistical evaluation of ChIPseq dataset similarityabstractMOTIVATION: ChIPseq is rapidly becoming a common technique for investigating protein-DNA interactions. However, results from individual experiments provide a limited understanding of chromatin structure, as various chromatin factors cooperate in complex ways to orchestrate transcription. In order to quantify chromtain interactions, it is thus necessary to devise a robust similarity metric applicable to ChIPseq data. Unfortunately, moving past simple overlap calculations to give statistically rigorous comparisons of ChIPseq datasets often involves arbitrary choices of distance metrics, with significance being estimated by computationally intensive permutation tests whose statistical power may be sensitive to non-biological experimental and post-processing variation. RESULTS: We show that it is in fact possible to compare ChIPseq datasets through the efficient computation of exact P-values for proximity. Our method is insensitive to non-biological variation in datasets such as peak width, and can rigorously model peak location biases by evaluating similarity conditioned on a restricted set of genomic regions (such as mappable genome or promoter regions). Applying our method to the well-studied dataset of Chen et al. (2008), we elucidate novel interactions which conform well with our biological understanding. By comparing ChIPseq data in an asymmetric way, we are able to observe clear interaction differences between cofactors such as p300 and factors that bind DNA directly. AVAILABILITY: Source code is available for download at http://sonorus.princeton.edu/IntervalStats/IntervalStats.tar.gz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Maria D. Chikina, Olga G. Troyanskaya |
Bioinform. | 2 |
| 2012 | Chapter 2: Data-Driven View of Disease BiologyabstractModern experimental strategies often generate genome-scale measurements of human tissues or cell lines in various physiological states. Investigators often use these datasets individually to help elucidate molecular mechanisms of human diseases. Here we discuss approaches that effectively weight and integrate hundreds of heterogeneous datasets to gene-gene networks that focus on a specific process or disease. Diverse and systematic genome-scale measurements provide such approaches both a great deal of power and a number of challenges. We discuss some such challenges as well as methods to address them. We also raise important considerations for the assessment and evaluation of such approaches. When carefully applied, these integrative data-driven methods can make novel high-quality predictions that can transform our understanding of the molecular-basis of human disease. Casey S. Greene, Olga G. Troyanskaya |
PLoS Comput. Biol. | 2 |
| 2012 | Tissue-Specific Functional Networks for Prioritizing Phenotype and Disease GenesabstractIntegrated analyses of functional genomics data have enormous potential for identifying phenotype-associated genes. Tissue-specificity is an important aspect of many genetic diseases, reflecting the potentially different roles of proteins and pathways in diverse cell lineages. Accounting for tissue specificity in global integration of functional genomics data is challenging, as "functionality" and "functional relationships" are often not resolved for specific tissue types. We address this challenge by generating tissue-specific functional networks, which can effectively represent the diversity of protein function for more accurate identification of phenotype-associated genes in the laboratory mouse. Specifically, we created 107 tissue-specific functional relationship networks through integration of genomic data utilizing knowledge of tissue-specific gene expression patterns. Cross-network comparison revealed significantly changed genes enriched for functions related to specific tissue development. We then utilized these tissue-specific networks to predict genes associated with different phenotypes. Our results demonstrate that prediction performance is significantly improved through using the tissue-specific networks as compared to the global functional network. We used a testis-specific functional relationship network to predict genes associated with male fertility and spermatogenesis phenotypes, and experimentally confirmed one top prediction, Mbyl1. We then focused on a less-common genetic disease, ataxia, and identified candidates uniquely predicted by the cerebellum network, which are supported by both literature and experimental evidence. Our systems-level, tissue-specific scheme advances over traditional global integration and analyses and establishes a prototype to address the tissue-specific effects of genetic perturbations, diseases and drugs. Yuanfang Guan, Dmitriy Gorenshteyn, Margit Burmeister, Aaron K. Wong, John C. Schimenti, Mary Ann Handel, Carol J. Bult, Matthew A. Hibbs, Olga G. Troyanskaya |
PLoS Comput. Biol. | 9 |
| 2011 | Accurate Quantification of Functional Analogy among Close HomologsabstractCorrectly evaluating functional similarities among homologous proteins is necessary for accurate transfer of experimental knowledge from one organism to another, and is of particular importance for the development of animal models of human disease. While the fact that sequence similarity implies functional similarity is a fundamental paradigm of molecular biology, sequence comparison does not directly assess the extent to which two proteins participate in the same biological processes, and has limited utility for analyzing families with several parologous members. Nevertheless, we show that it is possible to provide a cross-organism functional similarity measure in an unbiased way through the exclusive use of high-throughput gene-expression data. Our methodology is based on probabilistic cross-species mapping of functionally analogous proteins based on Bayesian integrative analysis of gene expression compendia. We demonstrate that even among closely related genes, our method is able to predict functionally analogous homolog pairs better than relying on sequence comparison alone. We also demonstrate that the landscape of functional similarity is often complex and that definitive "functional orthologs" do not always exist. Even in these cases, our method and the online interface we provide are designed to allow detailed exploration of sources of inferred functional similarity that can be evaluated by the user. Maria D. Chikina, Olga G. Troyanskaya |
PLoS Comput. Biol. | 2 |
| 2010 | Functional Genomics Complements Quantitative Genetics in Identifying Disease-Gene AssociationsabstractAn ultimate goal of genetic research is to understand the connection between genotype and phenotype in order to improve the diagnosis and treatment of diseases. The quantitative genetics field has developed a suite of statistical methods to associate genetic loci with diseases and phenotypes, including quantitative trait loci (QTL) linkage mapping and genome-wide association studies (GWAS). However, each of these approaches have technical and biological shortcomings. For example, the amount of heritable variation explained by GWAS is often surprisingly small and the resolution of many QTL linkage mapping studies is poor. The predictive power and interpretation of QTL and GWAS results are consequently limited. In this study, we propose a complementary approach to quantitative genetics by interrogating the vast amount of high-throughput genomic data in model organisms to functionally associate genes with phenotypes and diseases. Our algorithm combines the genome-wide functional relationship network for the laboratory mouse and a state-of-the-art machine learning method. We demonstrate the superior accuracy of this algorithm through predicting genes associated with each of 1157 diverse phenotype ontology terms. Comparison between our prediction results and a meta-analysis of quantitative genetic studies reveals both overlapping candidates and distinct, accurate predictions uniquely identified by our approach. Focusing on bone mineral density (BMD), a phenotype related to osteoporotic fracture, we experimentally validated two of our novel predictions (not observed in any previous GWAS/QTL studies) and found significant bone density defects for both Timp2 and Abcg8 deficient mice. Our results suggest that the integration of functional genomics data into networks, which itself is informative of protein function and interactions, can successfully be utilized as a complementary approach to quantitative genetics to predict disease risks. All supplementary material is available at http://cbfg.jax.org/phenotype. Yuanfang Guan, Cheryl L. Ackert-Bicknell, Braden Kell, Olga G. Troyanskaya, Matthew A. Hibbs |
PLoS Comput. Biol. | 4 |
| 2010 | Systematic Planning of Genome-Scale Experiments in Poorly Studied SpeciesabstractGenome-scale datasets have been used extensively in model organisms to screen for specific candidates or to predict functions for uncharacterized genes. However, despite the availability of extensive knowledge in model organisms, the planning of genome-scale experiments in poorly studied species is still based on the intuition of experts or heuristic trials. We propose that computational and systematic approaches can be applied to drive the experiment planning process in poorly studied species based on available data and knowledge in closely related model organisms. In this paper, we suggest a computational strategy for recommending genome-scale experiments based on their capability to interrogate diverse biological processes to enable protein function assignment. To this end, we use the data-rich functional genomics compendium of the model organism to quantify the accuracy of each dataset in predicting each specific biological process and the overlap in such coverage between different datasets. Our approach uses an optimized combination of these quantifications to recommend an ordered list of experiments for accurately annotating most proteins in the poorly studied related organisms to most biological processes, as well as a set of experiments that target each specific biological process. The effectiveness of this experiment- planning system is demonstrated for two related yeast species: the model organism Saccharomyces cerevisiae and the comparatively poorly studied Saccharomyces bayanus. Our system recommended a set of S. bayanus experiments based on an S. cerevisiae microarray data compendium. In silico evaluations estimate that less than 10% of the experiments could achieve similar functional coverage to the whole microarray compendium. This estimation was confirmed by performing the recommended experiments in S. bayanus, therefore significantly reducing the labor devoted to characterize the poorly studied genome. This experiment-planning framework could readily be adapted to the design of other types of large-scale experiments as well as other groups of organisms. Yuanfang Guan, Maitreya J. Dunham, Amy A. Caudy, Olga G. Troyanskaya |
PLoS Comput. Biol. | 4 |
| 2010 | Mapping Dynamic Histone Acetylation Patterns to Gene Expression in Nanog-Depleted Murine Embryonic Stem CellsabstractEmbryonic stem cells (ESC) have the potential to self-renew indefinitely and to differentiate into any of the three germ layers. The molecular mechanisms for self-renewal, maintenance of pluripotency and lineage specification are poorly understood, but recent results point to a key role for epigenetic mechanisms. In this study, we focus on quantifying the impact of histone 3 acetylation (H3K9,14ac) on gene expression in murine embryonic stem cells. We analyze genome-wide histone acetylation patterns and gene expression profiles measured over the first five days of cell differentiation triggered by silencing Nanog, a key transcription factor in ESC regulation. We explore the temporal and spatial dynamics of histone acetylation data and its correlation with gene expression using supervised and unsupervised statistical models. On a genome-wide scale, changes in acetylation are significantly correlated to changes in mRNA expression and, surprisingly, this coherence increases over time. We quantify the predictive power of histone acetylation for gene expression changes in a balanced cross-validation procedure. In an in-depth study we focus on genes central to the regulatory network of Mouse ESC, including those identified in a recent genome-wide RNAi screen and in the PluriNet, a computationally derived stem cell signature. We find that compared to the rest of the genome, ESC-specific genes show significantly more acetylation signal and a much stronger decrease in acetylation over time, which is often not reflected in a concordant expression change. These results shed light on the complexity of the relationship between histone acetylation and gene expression and are a step forward to dissect the multilayer regulatory mechanisms that determine stem cell fate. Florian Markowetz, Klaas W. Mulder, Edoardo M. Airoldi, Ihor Lemischka, Olga G. Troyanskaya |
PLoS Comput. Biol. | 5 |
| 2010 | Simultaneous Genome-Wide Inference of Physical, Genetic, Regulatory, and Functional Pathway ComponentsabstractBiomolecular pathways are built from diverse types of pairwise interactions, ranging from physical protein-protein interactions and modifications to indirect regulatory relationships. One goal of systems biology is to bridge three aspects of this complexity: the growing body of high-throughput data assaying these interactions; the specific interactions in which individual genes participate; and the genome-wide patterns of interactions in a system of interest. Here, we describe methodology for simultaneously predicting specific types of biomolecular interactions using high-throughput genomic data. This results in a comprehensive compendium of whole-genome networks for yeast, derived from ∼3,500 experimental conditions and describing 30 interaction types, which range from general (e.g. physical or regulatory) to specific (e.g. phosphorylation or transcriptional regulation). We used these networks to investigate molecular pathways in carbon metabolism and cellular transport, proposing a novel connection between glycogen breakdown and glucose utilization supported by recent publications. Additionally, 14 specific predicted interactions in DNA topological change and protein biosynthesis were experimentally validated. We analyzed the systems-level network features within all interactomes, verifying the presence of small-world properties and enrichment for recurring network motifs. This compendium of physical, synthetic, regulatory, and functional interaction networks has been made publicly available through an interactive web interface for investigators to utilize in future research at http://function.princeton.edu/bioweaver/. Christopher Y. Park, David C. Hess, Curtis Huttenhower, Olga G. Troyanskaya |
PLoS Comput. Biol. | 4 |
| 2009 | Aneuploidy prediction and tumor classification with heterogeneous hidden conditional random fieldsabstractMOTIVATION: The heterogeneity of cancer cannot always be recognized by tumor morphology, but may be reflected by the underlying genetic aberrations. Array comparative genome hybridization (array-CGH) methods provide high-throughput data on genetic copy numbers, but determining the clinically relevant copy number changes remains a challenge. Conventional classification methods for linking recurrent alterations to clinical outcome ignore sequential correlations in selecting relevant features. Conversely, existing sequence classification methods can only model overall copy number instability, without regard to any particular position in the genome. RESULTS: Here, we present the heterogeneous hidden conditional random field, a new integrated array-CGH analysis method for jointly classifying tumors, inferring copy numbers and identifying clinically relevant positions in recurrent alteration regions. By capturing the sequentiality as well as the locality of changes, our integrated model provides better noise reduction, and achieves more relevant gene retrieval and more accurate classification than existing methods. We provide an efficient L1-regularized discriminative training algorithm, which notably selects a small set of candidate genes most likely to be clinically relevant and driving the recurrent amplicons of importance. Our method thus provides unbiased starting points in deciding which genomic regions and which genes in particular to pursue for further examination. Our experiments on synthetic data and real genomic cancer prediction data show that our method is superior, both in prediction accuracy and relevant feature discovery, to existing methods. We also demonstrate that it can be used to generate novel biological hypotheses for breast cancer. Zafer Barutçuoglu, Edoardo M. Airoldi, Vanessa Dumeaux, Robert E. Schapire, Olga G. Troyanskaya |
Bioinform. | 5 |
| 2009 | The impact of incomplete knowledge on evaluation: an experimental benchmark for protein function predictionabstractMOTIVATION: Rapidly expanding repositories of highly informative genomic data have generated increasing interest in methods for protein function prediction and inference of biological networks. The successful application of supervised machine learning to these tasks requires a gold standard for protein function: a trusted set of correct examples, which can be used to assess performance through cross-validation or other statistical approaches. Since gene annotation is incomplete for even the best studied model organisms, the biological reliability of such evaluations may be called into question. RESULTS: We address this concern by constructing and analyzing an experimentally based gold standard through comprehensive validation of protein function predictions for mitochondrion biogenesis in Saccharomyces cerevisiae. Specifically, we determine that (i) current machine learning approaches are able to generalize and predict novel biology from an incomplete gold standard and (ii) incomplete functional annotations adversely affect the evaluation of machine learning performance. While computational approaches performed better than predicted in the face of incomplete data, relative comparison of competing approaches-even those employing the same training data-is problematic with a sparse gold standard. Incomplete knowledge causes individual methods' performances to be differentially underestimated, resulting in misleading performance evaluations. We provide a benchmark gold standard for yeast mitochondria to complement current databases and an analysis of our experimental results in the hopes of mitigating these effects in future comparative evaluations. AVAILABILITY: The mitochondrial benchmark gold standard, as well as experimental results and additional data, is available at http://function.princeton.edu/mitochondria. Curtis Huttenhower, Matthew A. Hibbs, Chad L. Myers, Amy A. Caudy, David C. Hess, Olga G. Troyanskaya |
Bioinform. | 6 |
| 2009 | Detailing regulatory networks through large scale data integrationabstractMOTIVATION: Much of a cell's regulatory response to changing environments occurs at the transcriptional level. Particularly in higher organisms, transcription factors (TFs), microRNAs and epigenetic modifications can combine to form a complex regulatory network. Part of this system can be modeled as a collection of regulatory modules: co-regulated genes, the conditions under which they are co-regulated and sequence-level regulatory motifs. RESULTS: We present the Combinatorial Algorithm for Expression and Sequence-based Cluster Extraction (COALESCE) system for regulatory module prediction. The algorithm is efficient enough to discover expression biclusters and putative regulatory motifs in metazoan genomes (>20,000 genes) and very large microarray compendia (>10,000 conditions). Using Bayesian data integration, it can also include diverse supporting data types such as evolutionary conservation or nucleosome placement. We validate its performance using a functional evaluation of co-clustered genes, known yeast and Escherichea coli TF targets, synthetic data and various metazoan data compendia. In all cases, COALESCE performs as well or better than current biclustering and motif prediction tools, with high accuracy in functional and TF/target assignments and zero false positives on synthetic data. COALESCE provides an efficient and flexible platform within which large, diverse data collections can be integrated to predict metazoan regulatory networks. AVAILABILITY: Source code (C++) is available at http://function.princeton.edu/sleipnir, and supporting data and a web interface are provided at http://function.princeton.edu/coalesce. CONTACT: [email protected]; [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Curtis Huttenhower, K. Tsheko Mutungu, Natasha Indik, Woongcheol Yang, Mark Schroeder, Joshua J. Forman, Olga G. Troyanskaya, Hilary A. Coller |
Bioinform. | 7 |
| 2009 | Papers on normalization, variable selection, classification or clustering of microarray dataabstractOver the last decade or so, there have been large numbers of methods published on approaches for normalization, variable (gene) selection, classification and clustering of microarray data. As indicated in the scope document for Bioinformatics, this requires papers describing new methods for these problems to meet a very high standard, showing important improvement in results for real biological data, as well as novelty. In this editorial, we describe some standards that need to be met for papers in these areas to be seriously considered. We ask that prospective authors consider these points carefully before submission of their papers to Bioinformatics. The role of simulation: Simulation can be useful in investigating the properties of various methods of data analysis. Yet, there are important barriers to credible use of simulation in microarray studies, largely due to what we do not know about the statistical distribution of measured gene expression levels. First, the distribution across transcripts of true expression values is dependent on the biological state of the tissue or cell, and for a given state this is unknown, even in distributional form, and may further exhibit gene- and platform-specific effects. Second, the correlation within biological replicates of true expression is unknown, and is likely unknowable in detail given that it is expressed by a correlation matrix with on the order of a billion entries. Third, the distribution of changes from one biological state to another is unknown. Fourth, the correlation in observational errors in gene expression across genes is unknown and similarly probably unknowable in detail. On the other hand, the measurement error for a given transcript has been well described by several authors (Ideker et al., 2001; Rocke and Durbin, 2001). Given this gap between knowledge and simulation specification, it is likely that any new method can be shown to be superior to some other method(s) by careful choice of simulation parameters, since simulations often include biases in the distributions selected and in other assumptions of the models. Thus, while simulation may still be worthwhile, and a useful tool for exploring robustness and parameter space of a new method, it is insufficient evidence for superiority of a new method without substantial support from significant improvement in results from analysis of real data. Normalization: Normalization necessarily involves a trade-off between its positive role in reducing variability, and its potentially negative role in increasing bias. There are a number of good image analysis, preprocessing, transformation and normalization methods extant for single- and dual-color DNA microarrays. To show that a new method is better requires comparison demonstrating that results in differential expression analysis, classification or clustering are better with the new normalization method than with previous methods. Not one but several previous methods should be chosen for comparison including the most widely used approaches. Several datasets should be used, including spike-in and dilution studies when feasible, as well as ‘real’ biological datasets. Showing that more genes are differentially expressed using a normalization method is not compelling evidence of superiority without a good estimate of the false-positive rate or a compelling biological analysis of the resulting differentially expressed genes. Variable selection: Typically, new variable selection methods are proposed as part of a classification or clustering strategy, and demonstrating superiority of the variable selection method usually means demonstrating superiority of the combined methodology. It is quite important that metrics for evaluation be used that are robust to intra-array correlations and variable selection artifacts. For example, in cross-validation studies in which variable selection is followed by a classification method, selection of variables using all the data and then cross-validating the classification accuracy introduces substantial bias, making classification methods appear more accurate than they really are (Ambroise and McLachlan, 2002). It is important that any method be compared with several of the most widely used existing methods, including baseline approaches such as filtering by t-score or forward stepwise analysis. Such comparisons should be performed on more than one biological dataset. Further, the method must demonstrate significant improvement over existing methods; incremental improvements will not be considered of sufficient interest to warrant review. Classification and prediction: New classification or prediction methods for microarray data enter a crowded arena. From long-standing techniques such as logistic regression and linear discriminant analysis to the more modern support vector machines and neural networks, most known classification methods have already been applied to microarray data. To show that a newly proposed classification method is a real advance, a substantial improvement in performance needs to be shown over a reasonable selection of existing datasets and methods, including commonly used or simple methods. This is because, consciously or subconsciously, the developer of a new method optimizes its characteristics against the datasets to be used for evaluation. Variable selection and parameter choice for all methods needs to be done strictly in the training set (whether there is one training set or many as in cross validation). Resampling methods like permuting the class labels on the arrays or the bootstrap can be used to provide robust estimates of the significance of differential expression, but do not in themselves give estimates of classification performance except to show that the performance is better than chance. Experience shows that there is considerable noise in classification accuracy experiments, so modest increases in achieved accuracy are usually not convincing. Experience also shows that classification performance in a microarray problem depends strongly on the dataset, and less on the variable selection and classification methods. More than modest differences are required to excite interest in a new method. Authors should keep in mind the ‘No Free Lunch Theorems’ of Wolpert and Macready (1997) which demonstrated that there is no optimization/classification method that outperforms all others in all circumstances (Wolpert, 1996). Clustering: Demonstrating superiority of a clustering method is in many ways more difficult than demonstrating superiority in a classification method. Usually, there is no ground truth against which to compare the clustering results. Defining a criterion (e.g. the Rand index) and showing that a clustering method achieves better scores on this criterion is often not compelling, since such criteria are easily optimized (again, consciously or subconsciously) to ensure superiority. For reasons discussed above, simulation is also not usually sufficient. Ideally, a new clustering method would demonstrate novel biological insights or some attractive statistical properties not available from previous methods, including several commonly used methods. Requiring new biological findings is a difficult standard, but a necessary one to insure that new published methods are useful and likely to be used. To conclude, microarrays remain a useful technology to address a wide array of biological problems and the optimal analysis of these data to extract meaningful results still pose many bioinformatics challenges. However, with a number of successful methods already addressing the well-established microarray data analysis problems, publication of new methods in this area requires either identification of a new challenge and formulation of a new problem or development of a substantially better methodology then those existing that can be benchmarked on a variety of datasets. We hope that suggestions provided above for evaluation and validation of such new methods would increase the likelihood of them supporting biological discoveries in the future. Funding: DMR to NIH grants P42-ES04699 and R01-HG003352. David M. Rocke, Trey Ideker, Olga G. Troyanskaya, John Quackenbush, Joaquín Dopazo |
Bioinform. | 3 |
| 2009 | Selected proceedings of the First Summit on Translational Bioinformatics 2008
Atul J. Butte, Indra Neil Sarkar, Marco Ramoni, Yves A. Lussier, Olga G. Troyanskaya |
BMC Bioinform. | 5 |
| 2009 | Graphle: Interactive exploration of large, dense graphsabstractBACKGROUND: A wide variety of biological data can be modeled as network structures, including experimental results (e.g. protein-protein interactions), computational predictions (e.g. functional interaction networks), or curated structures (e.g. the Gene Ontology). While several tools exist for visualizing large graphs at a global level or small graphs in detail, previous systems have generally not allowed interactive analysis of dense networks containing thousands of vertices at a level of detail useful for biologists. Investigators often wish to explore specific portions of such networks from a detailed, gene-specific perspective, and balancing this requirement with the networks' large size, complex structure, and rich metadata is a substantial computational challenge. RESULTS: Graphle is an online interface to large collections of arbitrary undirected, weighted graphs, each possibly containing tens of thousands of vertices (e.g. genes) and hundreds of millions of edges (e.g. interactions). These are stored on a centralized server and accessed efficiently through an interactive Java applet. The Graphle applet allows a user to examine specific portions of a graph, retrieving the relevant neighborhood around a set of query vertices (genes). This neighborhood can then be refined and modified interactively, and the results can be saved either as publication-quality images or as raw data for further analysis. The Graphle web site currently includes several hundred biological networks representing predicted functional relationships from three heterogeneous data integration systems: S. cerevisiae data from bioPIXIE, E. coli data using MEFIT, and H. sapiens data from HEFalMp. CONCLUSIONS: Graphle serves as a search and visualization engine for biological networks, which can be managed locally (simplifying collaborative data sharing) and investigated remotely. The Graphle framework is freely downloadable and easily installed on new servers, allowing any lab to quickly set up a Graphle site from which their own biological network data can be shared online. Curtis Huttenhower, Sajid O. Mehmood, Olga G. Troyanskaya |
BMC Bioinform. | 3 |
| 2009 | Predicting Cellular Growth from Gene Expression SignaturesabstractMaintaining balanced growth in a changing environment is a fundamental systems-level challenge for cellular physiology, particularly in microorganisms. While the complete set of regulatory and functional pathways supporting growth and cellular proliferation are not yet known, portions of them are well understood. In particular, cellular proliferation is governed by mechanisms that are highly conserved from unicellular to multicellular organisms, and the disruption of these processes in metazoans is a major factor in the development of cancer. In this paper, we develop statistical methodology to identify quantitative aspects of the regulatory mechanisms underlying cellular proliferation in Saccharomyces cerevisiae. We find that the expression levels of a small set of genes can be exploited to predict the instantaneous growth rate of any cellular culture with high accuracy. The predictions obtained in this fashion are robust to changing biological conditions, experimental methods, and technological platforms. The proposed model is also effective in predicting growth rates for the related yeast Saccharomyces bayanus and the highly diverged yeast Schizosaccharomyces pombe, suggesting that the underlying regulatory signature is conserved across a wide range of unicellular evolution. We investigate the biological significance of the gene expression signature that the predictions are based upon from multiple perspectives: by perturbing the regulatory network through the Ras/PKA pathway, observing strong upregulation of growth rate even in the absence of appropriate nutrients, and discovering putative transcription factor binding sites, observing enrichment in growth-correlated genes. More broadly, the proposed methodology enables biological insights about growth at an instantaneous time scale, inaccessible by direct experimental methods. Data and tools enabling others to apply our methods are available at http://function.princeton.edu/growthrate. Edoardo M. Airoldi, Curtis Huttenhower, David Gresham, Charles Lu 0004, Amy A. Caudy, Maitreya J. Dunham, James R. Broach, David Botstein, Olga G. Troyanskaya |
PLoS Comput. Biol. | 9 |
| 2009 | Coordinated Concentration Changes of Transcripts and Metabolites in Saccharomyces cerevisiaeabstractMetabolite concentrations can regulate gene expression, which can in turn regulate metabolic activity. The extent to which functionally related transcripts and metabolites show similar patterns of concentration changes, however, remains unestablished. We measure and analyze the metabolomic and transcriptional responses of Saccharomyces cerevisiae to carbon and nitrogen starvation. Our analysis demonstrates that transcripts and metabolites show coordinated response dynamics. Furthermore, metabolites and gene products whose concentration profiles are alike tend to participate in related biological processes. To identify specific, functionally related genes and metabolites, we develop an approach based on Bayesian integration of the joint metabolomic and transcriptomic data. This algorithm finds interactions by evaluating transcript-metabolite correlations in light of the experimental context in which they occur and the class of metabolite involved. It effectively predicts known enzymatic and regulatory relationships, including a gene-metabolite interaction central to the glycolytic-gluconeogenetic switch. This work provides quantitative evidence that functionally related metabolites and transcripts show coherent patterns of behavior on the genome scale and lays the groundwork for building gene-metabolite interaction networks directly from systems-level data. Patrick H. Bradley, Matthew J. Brauer, Joshua D. Rabinowitz, Olga G. Troyanskaya |
PLoS Comput. Biol. | 4 |
| 2009 | Global Prediction of Tissue-Specific Gene Expression and Context-Dependent Gene Networks in Caenorhabditis elegansabstractTissue-specific gene expression plays a fundamental role in metazoan biology and is an important aspect of many complex diseases. Nevertheless, an organism-wide map of tissue-specific expression remains elusive due to difficulty in obtaining these data experimentally. Here, we leveraged existing whole-animal Caenorhabditis elegans microarray data representing diverse conditions and developmental stages to generate accurate predictions of tissue-specific gene expression and experimentally validated these predictions. These patterns of tissue-specific expression are more accurate than existing high-throughput experimental studies for nearly all tissues; they also complement existing experiments by addressing tissue-specific expression present at particular developmental stages and in small tissues. We used these predictions to address several experimentally challenging questions, including the identification of tissue-specific transcriptional motifs and the discovery of potential miRNA regulation specific to particular tissues. We also investigate the role of tissue context in gene function through tissue-specific functional interaction networks. To our knowledge, this is the first study producing high-accuracy predictions of tissue-specific expression and interactions for a metazoan organism based on whole-animal data. Maria D. Chikina, Curtis Huttenhower, Coleen T. Murphy, Olga G. Troyanskaya |
PLoS Comput. Biol. | 4 |
| 2009 | Directing Experimental Biology: A Case Study in Mitochondrial BiogenesisabstractComputational approaches have promised to organize collections of functional genomics data into testable predictions of gene and protein involvement in biological processes and pathways. However, few such predictions have been experimentally validated on a large scale, leaving many bioinformatic methods unproven and underutilized in the biology community. Further, it remains unclear what biological concerns should be taken into account when using computational methods to drive real-world experimental efforts. To investigate these concerns and to establish the utility of computational predictions of gene function, we experimentally tested hundreds of predictions generated from an ensemble of three complementary methods for the process of mitochondrial organization and biogenesis in Saccharomyces cerevisiae. The biological data with respect to the mitochondria are presented in a companion manuscript published in PLoS Genetics (doi:10.1371/journal.pgen.1000407). Here we analyze and explore the results of this study that are broadly applicable for computationalists applying gene function prediction techniques, including a new experimental comparison with 48 genes representing the genomic background. Our study leads to several conclusions that are important to consider when driving laboratory investigations using computational prediction approaches. While most genes in yeast are already known to participate in at least one biological process, we confirm that genes with known functions can still be strong candidates for annotation of additional gene functions. We find that different analysis techniques and different underlying data can both greatly affect the types of functional predictions produced by computational methods. This diversity allows an ensemble of techniques to substantially broaden the biological scope and breadth of predictions. We also find that performing prediction and validation steps iteratively allows us to more completely characterize a biological area of interest. While this study focused on a specific functional area in yeast, many of these observations may be useful in the contexts of other processes and organisms. Matthew A. Hibbs, Chad L. Myers, Curtis Huttenhower, David C. Hess, Kai Li 0001, Amy A. Caudy, Olga G. Troyanskaya |
PLoS Comput. Biol. | 7 |
| 2008 | Assessing the functional structure of genomic dataabstractMOTIVATION: The availability of genome-scale data has enabled an abundance of novel analysis techniques for investigating a variety of systems-level biological relationships. As thousands of such datasets become available, they provide an opportunity to study high-level associations between cellular pathways and processes. This also allows the exploration of shared functional enrichments between diverse biological datasets, and it serves to direct experimenters to areas of low data coverage or with high probability of new discoveries. RESULTS: We analyze the functional structure of Saccharomyces cerevisiae datasets from over 950 publications in the context of over 140 biological processes. This includes a coverage analysis of biological processes given current high-throughput data, a data-driven map of associations between processes, and a measure of similar functional activity between genome-scale datasets. This uncovers subtle gene expression similarities in three otherwise disparate microarray datasets due to a shared strain background. We also provide several means of predicting areas of yeast biology likely to benefit from additional high-throughput experimental screens. AVAILABILITY: Predictions are provided in supplementary tables; software and additional data are available from the authors by request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Curtis Huttenhower, Olga G. Troyanskaya |
ISMB | 2 |
| 2008 | The Sleipnir library for computational functional genomicsabstractMOTIVATION: Biological data generation has accelerated to the point where hundreds or thousands of whole-genome datasets of various types are available for many model organisms. This wealth of data can lead to valuable biological insights when analyzed in an integrated manner, but the computational challenge of managing such large data collections is substantial. In order to mine these data efficiently, it is necessary to develop methods that use storage, memory and processing resources carefully. RESULTS: The Sleipnir C++ library implements a variety of machine learning and data manipulation algorithms with a focus on heterogeneous data integration and efficiency for very large biological data collections. Sleipnir allows microarray processing, functional ontology mining, clustering, Bayesian learning and inference and support vector machine tasks to be performed for heterogeneous data on scales not previously practical. In addition to the library, which can easily be integrated into new computational systems, prebuilt tools are provided to perform a variety of common tasks. Many tools are multithreaded for parallelization in desktop or high-throughput computing environments, and most tasks can be performed in minutes for hundreds of datasets using a standard personal computer. AVAILABILITY: Source code (C++) and documentation are available at http://function.princeton.edu/sleipnir and compiled binaries are available from the authors on request. Curtis Huttenhower, Mark Schroeder, Maria D. Chikina, Olga G. Troyanskaya |
Bioinform. | 4 |
| 2008 | A Genomewide Functional Network for the Laboratory MouseabstractEstablishing a functional network is invaluable to our understanding of gene function, pathways, and systems-level properties of an organism and can be a powerful resource in directing targeted experiments. In this study, we present a functional network for the laboratory mouse based on a Bayesian integration of diverse genetic and functional genomic data. The resulting network includes probabilistic functional linkages among 20,581 protein-coding genes. We show that this network can accurately predict novel functional assignments and network components and present experimental evidence for predictions related to Nanog homeobox (Nanog), a critical gene in mouse embryonic stem cell pluripotency. An analysis of the global topology of the mouse functional network reveals multiple biologically relevant systems-level features of the mouse proteome. Specifically, we identify the clustering coefficient as a critical characteristic of central modulators that affect diverse pathways as well as genes associated with different phenotype traits and diseases. In addition, a cross-species comparison of functional interactomes on a genomic scale revealed distinct functional characteristics of conserved neighborhoods as compared to subnetworks specific to higher organisms. Thus, our global functional network for the laboratory mouse provides the community with a key resource for discovering protein functions and novel pathway components as well as a tool for exploring systems-level topological and evolutionary features of cellular interactomes. To facilitate exploration of this network by the biomedical research community, we illustrate its application in function and disease gene discovery through an interactive, Web-based, publicly available interface at http://mouseNET.princeton.edu. Yuanfang Guan, Chad L. Myers, Ihor Lemischka, Carol J. Bult, Olga G. Troyanskaya |
PLoS Comput. Biol. | 6 |
| 2007 | Scalable, Dynamic Analysis and Visualization for Genomic DatasetsabstractA challenge in data analysis and visualization is to build new-generation software tools and systems to truly accelerate scientific discoveries. The recent focus of Princeton's next-generation software project is to investigate how to develop new-generation data analysis and visualization capabilities for genomic scientists to analyze high-throughput genomic datasets. This paper describes the software tools we have recently developed to enable dynamic, large-scale data analysis and visualization of multiple datasets on large-scale, high-resolution display wall systems. Our initial experience with the deployed tools at Princeton's Lewis-Sigler Institute for Integrative Genomics is very encouraging. Scientists can effectively learn new knowledge from multiple datasets, find new insights, and generate new hypotheses that are not possible with current methods. Grant Wallace, Matthew A. Hibbs, Maitreya J. Dunham, Rachel S. G. Sealfon, Olga G. Troyanskaya, Kai Li 0001 |
IPDPS | 5 |
| 2007 | Viewing the Larger Context of Genomic Data through Horizontal IntegrationabstractGenomics is an important emerging scientific field that relies on meaningful data visualization as a key step in analysis. Specifically, most investigation of gene expression microarray data is performed using visualization techniques. However, as microarrays become more ubiquitous, researchers must analyze their own data within the context of previously published work in order to gain a more complete understanding. No current method for microarray visualization and analysis enables biology researchers to observe the greater context of data that surrounds their own results, which severely limits the ability of researchers draw novel conclusions. Here we present a system, called HIDRA, that visually integrates the simultaneous display of multiple microarray datasets to identify important parallels and dissimilarities. We demonstrate the power of our approach through examples of real-world biological insights that can be observed using HIDRA that are not apparent using other techniques. Matthew A. Hibbs, Grant Wallace, Maitreya J. Dunham, Kai Li 0001, Olga G. Troyanskaya |
IV | 5 |
| 2007 | Exploring the functional landscape of gene expression: directed search of large microarray compendiaabstractMOTIVATION: The increasing availability of gene expression microarray technology has resulted in the publication of thousands of microarray gene expression datasets investigating various biological conditions. This vast repository is still underutilized due to the lack of methods for fast, accurate exploration of the entire compendium. RESULTS: We have collected Saccharomyces cerevisiae gene expression microarray data containing roughly 2400 experimental conditions. We analyzed the functional coverage of this collection and we designed a context-sensitive search algorithm for rapid exploration of the compendium. A researcher using our system provides a small set of query genes to establish a biological search context; based on this query, we weight each dataset's relevance to the context, and within these weighted datasets we identify additional genes that are co-expressed with the query set. Our method exhibits an average increase in accuracy of 273% compared to previous mega-clustering approaches when recapitulating known biology. Further, we find that our search paradigm identifies novel biological predictions that can be verified through further experimentation. Our methodology provides the ability for biological researchers to explore the totality of existing microarray data in a manner useful for drawing conclusions and formulating hypotheses, which we believe is invaluable for the research community. AVAILABILITY: Our query-driven search engine, called SPELL, is available at http://function.princeton.edu/SPELL. SUPPLEMENTARY INFORMATION: Several additional data files, figures and discussions are available at http://function.princeton.edu/SPELL/supplement. Matthew A. Hibbs, David C. Hess, Chad L. Myers, Curtis Huttenhower, Kai Li 0001, Olga G. Troyanskaya |
Bioinform. | 6 |
| 2007 | Context-sensitive data integration and prediction of biological networksabstractMOTIVATION: Several recent methods have addressed the problem of heterogeneous data integration and network prediction by modeling the noise inherent in high-throughput genomic datasets, which can dramatically improve specificity and sensitivity and allow the robust integration of datasets with heterogeneous properties. However, experimental technologies capture different biological processes with varying degrees of success, and thus, each source of genomic data can vary in relevance depending on the biological process one is interested in predicting. Accounting for this variation can significantly improve network prediction, but to our knowledge, no previous approaches have explicitly leveraged this critical information about biological context. RESULTS: We confirm the presence of context-dependent variation in functional genomic data and propose a Bayesian approach for context-sensitive integration and query-based recovery of biological process-specific networks. By applying this method to Saccharomyces cerevisiae, we demonstrate that leveraging contextual information can significantly improve the precision of network predictions, including assignment for uncharacterized genes. We expect that this general context-sensitive approach can be applied to other organisms and prediction scenarios. AVAILABILITY: A software implementation of our approach is available on request from the authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at http://avis.princeton.edu/contextPIXIE/ Chad L. Myers, Olga G. Troyanskaya |
Bioinform. | 2 |
| 2007 | Nearest Neighbor Networks: clustering expression data based on gene neighborhoodsabstractBACKGROUND: The availability of microarrays measuring thousands of genes simultaneously across hundreds of biological conditions represents an opportunity to understand both individual biological pathways and the integrated workings of the cell. However, translating this amount of data into biological insight remains a daunting task. An important initial step in the analysis of microarray data is clustering of genes with similar behavior. A number of classical techniques are commonly used to perform this task, particularly hierarchical and K-means clustering, and many novel approaches have been suggested recently. While these approaches are useful, they are not without drawbacks; these methods can find clusters in purely random data, and even clusters enriched for biological functions can be skewed towards a small number of processes (e.g. ribosomes). RESULTS: We developed Nearest Neighbor Networks (NNN), a graph-based algorithm to generate clusters of genes with similar expression profiles. This method produces clusters based on overlapping cliques within an interaction network generated from mutual nearest neighborhoods. This focus on nearest neighbors rather than on absolute distance measures allows us to capture clusters with high connectivity even when they are spatially separated, and requiring mutual nearest neighbors allows genes with no sufficiently similar partners to remain unclustered. We compared the clusters generated by NNN with those generated by eight other clustering methods. NNN was particularly successful at generating functionally coherent clusters with high precision, and these clusters generally represented a much broader selection of biological processes than those recovered by other methods. CONCLUSION: The Nearest Neighbor Networks algorithm is a valuable clustering method that effectively groups genes that are likely to be functionally related. It is particularly attractive due to its simplicity, its success in the analysis of large datasets, and its ability to span a wide range of biological functions with high precision. Curtis Huttenhower, Avi I. Flamholz, Jessica N. Landis, Sauhard Sahi, Chad L. Myers, Kellen L. Olszewski, Matthew A. Hibbs, Nathan O. Siemers, Olga G. Troyanskaya, Hilary A. Coller |
BMC Bioinform. | 9 |
| 2007 | "Getting Started In...": A Series Not to MissabstractIn recent decades, computational biology has established itself as a critical part of biomedical research and as an ever-growing and highly interdisciplinary field. It seems that every year, new areas of research appear—sequence analysis now includes SNP detection and whole genome alignments, gene expression microarray data can be analyzed in concert with interaction networks, and protein structures can be used to predict potential drug targets. This diversity and development of bioinformatics has led to an increasing number of approaches and techniques that require understanding of areas as diverse as biology and computer science, medicine and statistics, chemistry and applied mathematics. With the knowledge required to explore each area of computational biology, the barrier grows for those trying to enter it. So how can one start learning about tiling array analysis or text mining in biology? What should a new graduate student read before analyzing his first microarray? How can a machine learning researcher find out what interesting problems in bioinformatics she may address with probabilistic graphical models? We hope that help is on the way.
This month, PLoS Computational Biology and the International Society for Computational Biology begin a series of short, practical articles for students and active researchers who want to learn more about new areas of computational biology and are unsure where or how to start. The aim of each article in the “Getting Started in…” series is to introduce the essentials: define the area and what it is about, highlight the debates and issues of relevance, and provide directions to the most relevant books, articles, or Web sites to find out more. The series will not include review articles or detailed tutorials; these are available in the Education section of the Journal. Rather, each “Getting Started in…” article will aim to be a cache of “go to” information for someone for whom the field is completely new. We hope each part of the series, written by experts in areas as diverse as data integration and phylogeny reconstruction, will be as invaluable as receiving an e-mail from a colleague who takes time and thought to offer the best advice and the essential introduction to his or her area of research.
The first expert to inform, motivate, and inspire readers to consider a new direction is Dr. Xiaole Shirley Liu, who introduces tiling microarrays. As the series progresses, we can look forward to learning about text mining and probabilistic graphical models, to name a few topics. We hope you find this new series useful and enjoyable. Olga G. Troyanskaya |
PLoS Comput. Biol. | 1 |
| 2006 | Hierarchical multi-label prediction of gene functionabstractMOTIVATION: Assigning functions for unknown genes based on diverse large-scale data is a key task in functional genomics. Previous work on gene function prediction has addressed this problem using independent classifiers for each function. However, such an approach ignores the structure of functional class taxonomies, such as the Gene Ontology (GO). Over a hierarchy of functional classes, a group of independent classifiers where each one predicts gene membership to a particular class can produce a hierarchically inconsistent set of predictions, where for a given gene a specific class may be predicted positive while its inclusive parent class is predicted negative. Taking the hierarchical structure into account resolves such inconsistencies and provides an opportunity for leveraging all classifiers in the hierarchy to achieve higher specificity of predictions. RESULTS: We developed a Bayesian framework for combining multiple classifiers based on the functional taxonomy constraints. Using a hierarchy of support vector machine (SVM) classifiers trained on multiple data types, we combined predictions in our Bayesian framework to obtain the most probable consistent set of predictions. Experiments show that over a 105-node subhierarchy of the GO, our Bayesian framework improves predictions for 93 nodes. As an additional benefit, our method also provides implicit calibration of SVM margin outputs to probabilities. Using this method, we make function predictions for multiple proteins, and experimentally confirm predictions for proteins involved in mitosis. SUPPLEMENTARY INFORMATION: Results for the 105 selected GO classes and predictions for 1059 unknown genes are available at: http://function.princeton.edu/genesite/ CONTACT: [email protected]. Zafer Barutçuoglu, Robert E. Schapire, Olga G. Troyanskaya |
Bioinform. | 3 |
| 2006 | A scalable method for integration and functional analysis of multiple microarray datasetsabstractMOTIVATION: The diverse microarray datasets that have become available over the past several years represent a rich opportunity and challenge for biological data mining. Many supervised and unsupervised methods have been developed for the analysis of individual microarray datasets. However, integrated analysis of multiple datasets can provide a broader insight into genetic regulation of specific biological pathways under a variety of conditions. RESULTS: To aid in the analysis of such large compendia of microarray experiments, we present Microarray Experiment Functional Integration Technology (MEFIT), a scalable Bayesian framework for predicting functional relationships from integrated microarray datasets. Furthermore, MEFIT predicts these functional relationships within the context of specific biological processes. All results are provided in the context of one or more specific biological functions, which can be provided by a biologist or drawn automatically from catalogs such as the Gene Ontology (GO). Using MEFIT, we integrated 40 Saccharomyces cerevisiae microarray datasets spanning 712 unique conditions. In tests based on 110 biological functions drawn from the GO biological process ontology, MEFIT provided a 5% or greater performance increase for 54 functions, with a 5% or more decrease in performance in only two functions. Curtis Huttenhower, Matthew A. Hibbs, Chad L. Myers, Olga G. Troyanskaya |
Bioinform. | 4 |
| 2006 | GOLEM: an interactive graph-based gene-ontology navigation and analysis toolabstractBACKGROUND: The Gene Ontology has become an extremely useful tool for the analysis of genomic data and structuring of biological knowledge. Several excellent software tools for navigating the gene ontology have been developed. However, no existing system provides an interactively expandable graph-based view of the gene ontology hierarchy. Furthermore, most existing tools are web-based or require an Internet connection, will not load local annotations files, and provide either analysis or visualization functionality, but not both. RESULTS: To address the above limitations, we have developed GOLEM (Gene Ontology Local Exploration Map), a visualization and analysis tool for focused exploration of the gene ontology graph. GOLEM allows the user to dynamically expand and focus the local graph structure of the gene ontology hierarchy in the neighborhood of any chosen term. It also supports rapid analysis of an input list of genes to find enriched gene ontology terms. The GOLEM application permits the user either to utilize local gene ontology and annotations files in the absence of an Internet connection, or to access the most recent ontology and annotation information from the gene ontology webpage. GOLEM supports global and organism-specific searches by gene ontology term name, gene ontology id and gene name. CONCLUSION: GOLEM is a useful software tool for biologists interested in visualizing the local directed acyclic graph structure of the gene ontology hierarchy and searching for gene ontology terms enriched in genes of interest. It is freely available both as an application and as an applet at http://function.princeton.edu/GOLEM. Rachel S. G. Sealfon, Matthew A. Hibbs, Curtis Huttenhower, Chad L. Myers, Olga G. Troyanskaya |
BMC Bioinform. | 5 |
| 2005 | Putting microarrays in a context: Integrated analysis of diverse biological dataabstractIn recent years, multiple types of high-throughput functional genomic data that facilitate rapid functional annotation of sequenced genomes have become available. Gene expression microarrays are the most commonly available source of such data. However, genomic data often sacrifice specificity for scale, yielding very large quantities of relatively lower-quality data than traditional experimental methods. Thus sophisticated analysis methods are necessary to make accurate functional interpretation of these large-scale data sets. This review presents an overview of recently developed methods that integrate the analysis of microarray data with sequence, interaction, localisation and literature data, and further outlines current challenges in the field. The focus of this review is on the use of such methods for gene function prediction, understanding of protein regulation and modelling of biological networks. Olga G. Troyanskaya |
Briefings Bioinform. | 1 |
| 2005 | Visualization methods for statistical analysis of microarray clustersabstractBACKGROUND: The most common method of identifying groups of functionally related genes in microarray data is to apply a clustering algorithm. However, it is impossible to determine which clustering algorithm is most appropriate to apply, and it is difficult to verify the results of any algorithm due to the lack of a gold-standard. Appropriate data visualization tools can aid this analysis process, but existing visualization methods do not specifically address this issue. RESULTS: We present several visualization techniques that incorporate meaningful statistics that are noise-robust for the purpose of analyzing the results of clustering algorithms on microarray data. This includes a rank-based visualization method that is more robust to noise, a difference display method to aid assessments of cluster quality and detection of outliers, and a projection of high dimensional data into a three dimensional space in order to examine relationships between clusters. Our methods are interactive and are dynamically linked together for comprehensive analysis. Further, our approach applies to both protein and gene expression microarrays, and our architecture is scalable for use on both desktop/laptop screens and large-scale display devices. This methodology is implemented in GeneVAnD (Genomic Visual ANalysis of Datasets) and is available at http://function.princeton.edu/GeneVAnD. CONCLUSION: Incorporating relevant statistical information into data visualizations is key for analysis of large biological datasets, particularly because of high levels of noise and the lack of a gold-standard for comparisons. We developed several new visualization techniques and demonstrated their effectiveness for evaluating cluster quality and relationships between clusters. Matthew A. Hibbs, Nathaniel C. Dirksen, Kai Li 0001, Olga G. Troyanskaya |
BMC Bioinform. | 4 |
| 2005 | Visualization-based discovery and analysis of genomic aberrations in microarray dataabstractBACKGROUND: Chromosomal copy number changes (aneuploidies) play a key role in cancer progression and molecular evolution. These copy number changes can be studied using microarray-based comparative genomic hybridization (array CGH) or gene expression microarrays. However, accurate identification of amplified or deleted regions requires a combination of visual and computational analysis of these microarray data. RESULTS: We have developed ChARMView, a visualization and analysis system for guided discovery of chromosomal abnormalities from microarray data. Our system facilitates manual or automated discovery of aneuploidies through dynamic visualization and integrated statistical analysis. ChARMView can be used with array CGH and gene expression microarray data, and multiple experiments can be viewed and analyzed simultaneously. CONCLUSION: ChARMView is an effective and accurate visualization and analysis system for recognizing even small aneuploidies or subtle expression biases, identifying recurring aberrations in sets of experiments, and pinpointing functionally relevant copy number changes. ChARMView is freely available under the GNU GPL at http://function.princeton.edu/ChARMView. Chad L. Myers, Olga G. Troyanskaya |
BMC Bioinform. | 3 |
| 2004 | Accurate detection of aneuploidies in array CGH and gene expression microarray dataabstractMOTIVATION: Chromosomal copy number changes (aneuploidies) are common in cell populations that undergo multiple cell divisions including yeast strains, cell lines and tumor cells. Identification of aneuploidies is critical in evolutionary studies, where changes in copy number serve an adaptive purpose, as well as in cancer studies, where amplifications and deletions of chromosomal regions have been identified as a major pathogenetic mechanism. Aneuploidies can be studied on whole-genome level using array CGH (a microarray-based method that measures the DNA content), but their presence also affects gene expression. In gene expression microarray analysis, identification of copy number changes is especially important in preventing aberrant biological conclusions based on spurious gene expression correlation or masked phenotypes that arise due to aneuploidies. Previously suggested approaches for aneuploidy detection from microarray data mostly focus on array CGH, address only whole-chromosome or whole-arm copy number changes, and rely on thresholds or other heuristics, making them unsuitable for fully automated general application to gene expression datasets. There is a need for a general and robust method for identification of aneuploidies of any size from both array CGH and gene expression microarray data. RESULTS: We present ChARM (Chromosomal Aberration Region Miner), a robust and accurate expectation-maximization based method for identification of segmental aneuploidies (partial chromosome changes) from gene expression and array CGH microarray data. Systematic evaluation of the algorithm on synthetic and biological data shows that the method is robust to noise, aneuploidal segment size and P-value cutoff. Using our approach, we identify known chromosomal changes and predict novel potential segmental aneuploidies in commonly used yeast deletion strains and in breast cancer. ChARM can be routinely used to identify aneuploidies in array CGH datasets and to screen gene expression data for aneuploidies or array biases. Our methodology is sensitive enough to detect statistically significant and biologically relevant aneuploidies even when expression or DNA content changes are subtle as in mixed populations of cells. AVAILABILITY: Code available by request from the authors and on Web supplement at http://function.cs.princeton.edu/ChARM/ Chad L. Myers, Maitreya J. Dunham, Sun-Yuan Kung, Olga G. Troyanskaya |
Bioinform. | 4 |
| 2002 | Sequence complexity profiles of prokaryotic genomic sequences: A fast algorithm for calculating linguistic complexityabstractMOTIVATION: One of the major features of genomic DNA sequences, distinguishing them from texts in most spoken or artificial languages, is their high repetitiveness. Variation in the repetitiveness of genomic texts reflects the presence and density of different biologically important messages. Thus, deviation from an expected number of repeats in both directions indicates a possible presence of a biological signal. Linguistic complexity corresponds to repetitiveness of a genomic text, and potential regulatory sites may be discovered through construction of typical patterns of complexity distribution. RESULTS: We developed software for fast calculation of linguistic sequence complexity of DNA sequences. Our program utilizes suffix trees to compute the number of subwords present in genomic sequences, thereby allowing calculation of linguistic complexity in time linear in genome size. The measure of linguistic complexity was applied to the complete genome of Haemophilus influenzae. Maps of complexity along the entire genome were obtained using sliding windows of 40, 100, and 2000 nucleotides. This approach provided an efficient way to detect simple sequence repeats in this genome. In addition, local profiles of complexity distribution around the starts of translation were constructed for 21 complete prokaryotic genomes. We hypothesize that complexity profiles correspond to evolutionary relationships between organisms. We found principal differences in profiles of the GC-rich and other (non-GC-rich) genomes. We also found characteristic differences in profiles of AT genomes, which probably reflect individual species variations in translational regulation. AVAILABILITY: The program is available upon request from Alexander Bolshoy or at http://csweb.haifa.ac.il/library/#complex. Olga G. Troyanskaya, Ora Arbell, Yair Koren, Gad M. Landau, Alexander Bolshoy |
Bioinform. | 1 |
| 2002 | Nonparametric methods for identifying differentially expressed genes in microarray dataabstractMOTIVATION: Gene expression experiments provide a fast and systematic way to identify disease markers relevant to clinical care. In this study, we address the problem of robust identification of differentially expressed genes from microarray data. Differentially expressed genes, or discriminator genes, are genes with significantly different expression in two user-defined groups of microarray experiments. We compare three model-free approaches: (1). nonparametric t-test, (2). Wilcoxon (or Mann-Whitney) rank sum test, and (3). a heuristic method based on high Pearson correlation to a perfectly differentiating gene ('ideal discriminator method'). We systematically assess the performance of each method based on simulated and biological data under varying noise levels and p-value cutoffs. RESULTS: All methods exhibit very low false positive rates and identify a large fraction of the differentially expressed genes in simulated data sets with noise level similar to that of actual data. Overall, the rank sum test appears most conservative, which may be advantageous when the computationally identified genes need to be tested biologically. However, if a more inclusive list of markers is desired, a higher p-value cutoff or the nonparametric t-test may be appropriate. When applied to data from lung tumor and lymphoma data sets, the methods identify biologically relevant differentially expressed genes that allow clear separation of groups in question. Thus the methods described and evaluated here provide a convenient and robust way to identify differentially expressed genes for further biological and clinical analysis. Olga G. Troyanskaya, Mitchell E. Garber, Patrick O. Brown, David Botstein, Russ B. Altman |
Bioinform. | 1 |
| 2001 | Missing value estimation methods for DNA microarraysabstractMOTIVATION: Gene expression microarray experiments can generate data sets with multiple missing expression values. Unfortunately, many algorithms for gene expression analysis require a complete matrix of gene array values as input. For example, methods such as hierarchical clustering and K-means clustering are not robust to missing data, and may lose effectiveness even with a few missing values. Methods for imputing missing data are needed, therefore, to minimize the effect of incomplete data sets on analyses, and to increase the range of data sets to which these algorithms can be applied. In this report, we investigate automated methods for estimating missing data. RESULTS: We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1--20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method (as well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions. Olga G. Troyanskaya, Michael N. Cantor, Gavin Sherlock, Patrick O. Brown, Trevor J. Hastie, Robert Tibshirani, David Botstein, Russ B. Altman |
Bioinform. | 1 |