Trey Ideker

dblp:77/1450 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Harmonization and integration of pharmacogenomics screens
abstract
MOTIVATION: Large pharmacogenomics screens have generated a wealth of information cataloguing the responses of more than a thousand tumor cell-line models to FDA-approved and exploratory drugs. Although centralized repositories have consolidated data access, the diversity of experimental platforms and response metrics used in these screens have made it challenging to integrate and compare their measured drug responses. Towards better pharmacogenomic data harmonization, we surveyed a range of data analysis protocols based on different curve-fitting functions (sigmoid, piecewise linear), different response metrics (IC50, EC50, integrated AUC), and different drug concentration windows (full range or truncated). RESULTS: We found that an AUC derived from a sigmoidal curve fitted to a truncated dose range yields the strongest agreement between screening platforms, significantly bettering other protocols surveyed. This harmonization procedure also best aligns drug responses across successive iterations of the same platform. These findings broadly inform efforts to integrate drug response data in large-scale analyses. AVAILABILITY: The source code to generate drug response profiles and correlations are available at https://github.com/digitaltumors/Pharmacogenomics_Screens_Harmonization.git.
Aleysha T. Chen, Marcus R. Kelly, Trey Ideker, Nicole M. Mattson
Bioinform.3
2025 An Adversarial Scheme for Integrating Multi-modal Data on Protein Function
Rami Nasser, Leah V. Schaffer, Trey Ideker, Roded Sharan
RECOMB3
2025 Cell Mapping Toolkit: an end-to-end pipeline for mapping subcellular organization
abstract
SUMMARY: Cells are organized as a hierarchy of macromolecular assemblies, ranging from small protein complexes to entire organelles. Various technologies have been developed to elucidate subcellular architecture at different scales, such as mass spectrometry approaches for mapping protein biophysical interactions and immunofluorescence imaging for mapping protein localization. We present the Cell Mapping Toolkit, which is designed to systematically integrate data from different modalities into unified hierarchical maps of subcellular organization. The toolkit facilitates an end-to-end pipeline including processing datasets, integrating modalities, and visualizing the final cell map with rich metadata including provenance documentation at each step. The Cell Mapping Toolkit provides researchers with tools for analyzing, integrating, and visualizing diverse protein datasets in a robust and reproducible framework. AVAILABILITY AND IMPLEMENTATION: The code is freely available and is hosted on GitHub at https://github.com/idekerlab/cellmaps_pipeline. Comprehensive documentation and practical examples are provided at https://cellmaps-pipeline.readthedocs.io/.
Joanna Lenkiewicz, Christopher Churas, Mengzhou Hu, Gege Qian, Maxwell Adam Levinson, Sadnan Al Manir, Dylan Fong, Keiichiro Ono, Chengzhan Gao, Dexter Pratt, Jillian Parker, Tim Clark, Trey Ideker, Leah V. Schaffer
Bioinform.16
2024 Multi-modal contrastive learning of subcellular organization using DICE
abstract
The data deluge in biology calls for computational approaches that can integrate multiple datasets of different types to build a holistic view of biological processes or structures of interest. An emerging paradigm in this domain is the unsupervised learning of data embeddings that can be used for downstream clustering and classification tasks. While such approaches for integrating data of similar types are becoming common, there is scarcer work on consolidating different data modalities such as network and image information. Here, we introduce DICE (Data Integration through Contrastive Embedding), a contrastive learning model for multi-modal data integration. We apply this model to study the subcellular organization of proteins by integrating protein-protein interaction data and protein image data measured in HEK293 cells. We demonstrate the advantage of data integration over any single modality and show that our framework outperforms previous integration approaches. Availability: https://github.com/raminass/protein-contrastive Contact: [email protected].
Rami Nasser, Leah V. Schaffer, Trey Ideker, Roded Sharan
Bioinform.3
2024 Representing mutations for predicting cancer drug response
abstract
MOTIVATION: Predicting cancer drug response requires a comprehensive assessment of many mutations present across a tumor genome. While current drug response models generally use a binary mutated/unmutated indicator for each gene, not all mutations in a gene are equivalent. RESULTS: Here, we construct and evaluate a series of predictive models based on leading methods for quantitative mutation scoring. Such methods include VEST4 and CADD, which score the impact of a mutation on gene function, and CHASMplus, which scores the likelihood a mutation drives cancer. The resulting predictive models capture cellular responses to dabrafenib, which targets BRAF-V600 mutations, whereas models based on binary mutation status do not. Performance improvements generalize to other drugs, extending genetic indications for PIK3CA, ERBB2, EGFR, PARP1, and ABL1 inhibitors. Introducing quantitative mutation features in drug response models increases performance and mechanistic understanding. AVAILABILITY AND IMPLEMENTATION: Code and example datasets are available at https://github.com/pgwall/qms.
Patrick Wall, Trey Ideker
Bioinform.2
2023 NDEx IQuery: a multi-method network gene set analysis leveraging the Network Data Exchange
abstract
MOTIVATION: The investigation of sets of genes using biological pathways is a common task for researchers and is supported by a wide variety of software tools. This type of analysis generates hypotheses about the biological processes that are active or modulated in a specific experimental context. RESULTS: The Network Data Exchange Integrated Query (NDEx IQuery) is a new tool for network and pathway-based gene set interpretation that complements or extends existing resources. It combines novel sources of pathways, integration with Cytoscape, and the ability to store and share analysis results. The NDEx IQuery web application performs multiple gene set analyses based on diverse pathways and networks stored in NDEx. These include curated pathways from WikiPathways and SIGNOR, published pathway figures from the last 27 years, machine-assembled networks using the INDRA system, and the new NCI-PID v2.0, an updated version of the popular NCI Pathway Interaction Database. NDEx IQuery's integration with MSigDB and cBioPortal now provides pathway analysis in the context of these two resources. AVAILABILITY AND IMPLEMENTATION: NDEx IQuery is available at https://www.ndexbio.org/iquery and is implemented in Javascript and Java.
Rudolf T. Pillich, Christopher Churas, Dylan Fong, Benjamin M. Gyori, Trey Ideker, Klas Karis, Sophie N. Liu, Keiichiro Ono, Alexander R. Pico, Dexter Pratt
Bioinform.6
2021 Genetic dissection of complex traits using hierarchical biological knowledge
abstract
Despite the growing constellation of genetic loci linked to common traits, these loci have yet to account for most heritable variation, and most act through poorly understood mechanisms. Recent machine learning (ML) systems have used hierarchical biological knowledge to associate genetic mutations with phenotypic outcomes, yielding substantial predictive power and mechanistic insight. Here, we use an ontology-guided ML system to map single nucleotide variants (SNVs) focusing on 6 classic phenotypic traits in natural yeast populations. The 29 identified loci are largely novel and account for ~17% of the phenotypic variance, versus <3% for standard genetic analysis. Representative results show that sensitivity to hydroxyurea is linked to SNVs in two alternative purine biosynthesis pathways, and that sensitivity to copper arises through failure to detoxify reactive oxygen species in fatty acid metabolism. This work demonstrates a knowledge-based approach to amplifying and interpreting signals in population genetic studies.
Hidenori Tanaka, Jason F. Kreisberg, Trey Ideker
PLoS Comput. Biol.3
2020 Multiscale community detection in Cytoscape
abstract
Detection of community structure has become a fundamental step in the analysis of biological networks with application to protein function annotation, disease gene prediction, and drug discovery. This recent impact creates a need to make these techniques and their accompanying visualization schemes available to a broad range of biologists. Here we present a service-oriented, end-to-end software framework, CDAPS (Community Detection APplication and Service), that integrates the identification, annotation, visualization, and interrogation of multiscale network communities, accessible within the popular Cytoscape network analysis platform. With novel design principles, CDAPS addresses unmet new challenges, such as identifying hierarchical community structures, comparison of outputs generated from diverse network resources, and easy deployment of new algorithms, to facilitate community-sourced science. We demonstrate that the CDAPS framework can be applied to high-throughput protein-protein interaction networks to gain novel insights, such as the identification of putative new members of known protein complexes.
Akshat Singhal, Christopher Churas, Dexter Pratt, Santo Fortunato, Trey Ideker
PLoS Comput. Biol.7
2019 Mitigating Data Scarcity in Protein Binding Prediction Using Meta-Learning
Yunan Luo, Jianzhu Ma, Xiaoming Zhao 0001, Yufeng Su, Yang Liu 0097, Trey Ideker, Jian Peng 0001
RECOMB6
2019 On entropy and information in gene interaction networks
abstract
MOTIVATION: Modern biological experiments often produce candidate lists of genes presumably related to the studied phenotype. One can ask if the gene list as a whole makes sense in the context of existing knowledge: Are the genes in the list reasonably related to each other or do they look like a random assembly? There are also situations when one wants to know if two or more gene sets are closely related. Gene enrichment tests based on counting the number of genes two sets have in common are adequate if we presume that two genes are related only when they are in fact identical. If by related we mean well connected in the interaction network space, we need a new measure of relatedness for gene sets. RESULTS: We derive entropy, interaction information and mutual information for gene sets on interaction networks, starting from a simple phenomenological model of a living cell. Formally, the model describes a set of interacting linear harmonic oscillators in thermal equilibrium. Because the energy function is a quadratic form of the degrees of freedom, entropy and all other derived information quantities can be calculated exactly. We apply these concepts to estimate the probability that genes from several independent genome-wide association studies are not mutually informative; to estimate the probability that two disjoint canonical metabolic pathways are not mutually informative; and to infer relationships among human diseases based on their gene signatures. We show that the present approach is able to predict observationally validated relationships not detectable by gene enrichment methods. The converse is also true; the two methods are therefore complementary. AVAILABILITY AND IMPLEMENTATION: The functions defined in this paper are available in an R package, gsia, available for download at https://github.com/ucsd-ccbb/gsia.
Z. S. Wallace, Sara Brin Rosenthal, Kathleen M. Fisch, Trey Ideker, Roman Sásik
Bioinform.4
2019 Classifying tumors by supervised network propagation
abstract
Bioinformatics (2018) doi: 10.1093/bioinformatics/bty247
Wei Zhang 0067, Jianzhu Ma, Trey Ideker
Bioinform.3
2019 Rare variant phasing using paired tumor: normal sequence data
abstract
BACKGROUND: In standard high throughput sequencing analysis, genetic variants are not assigned to a homologous chromosome of origin. This process, called haplotype phasing, can reveal information important for understanding the relationship between genetic variants and biological phenotypes. For example, in genes that carry multiple heterozygous missense variants, phasing resolves whether one or both gene copies are altered. Here, we present a novel approach to phasing variants that takes advantage of unique properties of paired tumor:normal sequencing data from cancer studies. RESULTS: VAF phasing uses changes in variant allele frequency (VAF) between tumor and normal samples in regions of somatic chromosomal gain or loss to phase germline variants. We apply VAF phasing to 6180 samples from the Cancer Genome Atlas (TCGA) and demonstrate that our method is highly concordant with other standard phasing methods, and can phase an average of 33% more variants than other read-backed phasing methods. Using variant annotation tools designed to score gene haplotypes, we find a suggestive association between carrying multiple missense variants in a single copy of a cancer predisposition gene and earlier age of cancer diagnosis. CONCLUSIONS: VAF phasing exploits unique properties of tumor genomes to increase the number of germline variants that can be phased over standard read-backed methods in paired tumor:normal samples. Our phase-informed association testing results call attention to the need to develop more tools for assessing the joint effect of multiple genetic variants.
Alexandra R. Buckley, Trey Ideker, Hannah Carter, Nicholas J. Schork
BMC Bioinform.2
2018 Deciphering Signaling Specificity with Deep Neural Networks
Yunan Luo, Jianzhu Ma, Yang Liu 0097, Trey Ideker, Jian Peng 0001
RECOMB5
2018 ndexr - an R package to interface with the network data exchange
abstract
Motivation: Seamless exchange of biological network data enables bioinformatic algorithms to integrate networks as prior knowledge input as well as to document resulting network output. However, the interoperability between pathway databases and various methods and platforms for analysis is currently lacking. The Network Data Exchange (NDEx) is an open-source data commons that facilitates the user-centered sharing and publication of networks of many types and formats. Results: Here, we present a software package that allows users to programmatically connect to and interface with NDEx servers from within R. The network repository can be searched and networks can be retrieved and converted into igraph-compatible objects. These networks can be modified and extended within R and uploaded back to the NDEx servers. Availability and implementation: ndexr is a free and open-source R package, available via GitHub (https://github.com/frankkramer-lab/ndexr) and Bioconductor (http://bioconductor.org/packages/ndexr/). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Florian Auer, Zaynab Hammoud, Alexandr Ishkin, Dexter Pratt, Trey Ideker, Frank Kramer 0001
Bioinform.5
2018 pyNBS: a Python implementation for network-based stratification of tumor mutations
abstract
Summary: We present pyNBS: a modularized Python 2.7 implementation of the network-based stratification (NBS) algorithm for stratifying tumor somatic mutation profiles into molecularly and clinically relevant subtypes. In addition to release of the software, we benchmark its key parameters and provide a compact cancer reference network that increases the significance of tumor stratification using the NBS algorithm. The structure of the code exposes key steps of the algorithm to foster further collaborative development. Availability and implementation: The package, along with examples and data, can be downloaded and installed from the URL https://github.com/idekerlab/pyNBS.
Justin K. Huang, Tongqiu Jia, Daniel E. Carlin, Trey Ideker
Bioinform.4
2018 Classifying tumors by supervised network propagation
abstract
Motivation: Network propagation has been widely used to aggregate and amplify the effects of tumor mutations using knowledge of molecular interaction networks. However, propagating mutations through interactions irrelevant to cancer leads to erosion of pathway signals and complicates the identification of cancer subtypes. Results: To address this problem we introduce a propagation algorithm, Network-Based Supervised Stratification (NBS2), which learns the mutated subnetworks underlying tumor subtypes using a supervised approach. Given an annotated molecular network and reference tumor mutation profiles for which subtypes have been predefined, NBS2 is trained by adjusting the weights on interaction features such that network propagation best recovers the provided subtypes. After training, weights are fixed such that mutation profiles of new tumors can be accurately classified. We evaluate NBS2 on breast and glioblastoma tumors, demonstrating that it outperforms the best network-based approaches in classifying tumors to known subtypes for these diseases. By interpreting the interaction weights, we highlight characteristic molecular pathways driving selected subtypes. Availability and implementation: The NBS2 package is freely available at: https://github.com/wzhang1984/NBSS. Supplementary information: Supplementary data are available at Bioinformatics online.
Wei Zhang 0067, Jianzhu Ma, Trey Ideker
Bioinform.3
2017 Network propagation in the cytoscape cyberinfrastructure
abstract
Network propagation is an important and widely used algorithm in systems biology, with applications in protein function prediction, disease gene prioritization, and patient stratification. However, up to this point it has required significant expertise to run. Here we extend the popular network analysis program Cytoscape to perform network propagation as an integrated function. Such integration greatly increases the access to network propagation by putting it in the hands of biologists and linking it to the many other types of network analysis and visualization available through Cytoscape. We demonstrate the power and utility of the algorithm by identifying mutations conferring resistance to Vemurafenib.
Daniel E. Carlin, Barry Demchak, Dexter Pratt, Eric Sage, Trey Ideker
PLoS Comput. Biol.5
2017 Network approaches and applications in biology
abstract
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
Trey Ideker, Ruth Nussinov
PLoS Comput. Biol.1
2014 Inferring gene ontologies from pairwise similarity data
abstract
MOTIVATION: While the manually curated Gene Ontology (GO) is widely used, inferring a GO directly from -omics data is a compelling new problem. Recognizing that ontologies are a directed acyclic graph (DAG) of terms and hierarchical relations, algorithms are needed that: analyze a full matrix of gene-gene pairwise similarities from -omics data; infer true hierarchical structure in these data rather than enforcing hierarchy as a computational artifact; and respect biological pleiotropy, by which a term in the hierarchy can relate to multiple higher level terms. Methods addressing these requirements are just beginning to emerge-none has been evaluated for GO inference. METHODS: We consider two algorithms [Clique Extracted Ontology (CliXO), LocalFitness] that uniquely satisfy these requirements, compared with methods including standard clustering. CliXO is a new approach that finds maximal cliques in a network induced by progressive thresholding of a similarity matrix. We evaluate each method's ability to reconstruct the GO biological process ontology from a similarity matrix based on (a) semantic similarities for GO itself or (b) three -omics datasets for yeast. RESULTS: For task (a) using semantic similarity, CliXO accurately reconstructs GO (>99% precision, recall) and outperforms other approaches (<20% precision, <20% recall). For task (b) using -omics data, CliXO outperforms other methods using two -omics datasets and achieves ∼30% precision and recall using YeastNet v3, similar to an earlier approach (Network Extracted Ontology) and better than LocalFitness or standard clustering (20-25% precision, recall). CONCLUSION: This study provides algorithmic foundation for building gene ontologies by capturing hierarchical and pleiotropic structure embedded in biomolecular data.
Michael Kramer, Janusz Dutkowski, Michael Yu, Vineet Bafna, Trey Ideker
Bioinform.5
2013 Improving Breast Cancer Survival Analysis through Competition-Based Multidimensional Modeling
abstract
Breast cancer is the most common malignancy in women and is responsible for hundreds of thousands of deaths annually. As with most cancers, it is a heterogeneous disease and different breast cancer subtypes are treated differently. Understanding the difference in prognosis for breast cancer based on its molecular and phenotypic features is one avenue for improving treatment by matching the proper treatment with molecular subtypes of the disease. In this work, we employed a competition-based approach to modeling breast cancer prognosis using large datasets containing genomic and clinical information and an online real-time leaderboard program used to speed feedback to the modeling team and to encourage each modeler to work towards achieving a higher ranked submission. We find that machine learning methods combined with molecular features selected based on expert prior knowledge can improve survival predictions compared to current best-in-class methodologies and that ensemble models trained across multiple user submissions systematically outperform individual models within the ensemble. We also find that model scores are highly consistent across multiple independent evaluations. This study serves as the pilot phase of a much larger competition open to the whole research community, with the goal of understanding general strategies for model optimization using clinical and molecular profiling data and providing an objective, transparent system for assessing prognostic models.
Erhan Bilal, Janusz Dutkowski, Justin Guinney, In Sock Jang, Benjamin A. Logsdon, Gaurav Pandey 0002, Benjamin A. Sauerwine, Yishai Shimoni, Hans Kristian Moen Vollan, Brigham H. Mecham, Oscar M. Rueda, Jorg Tost, Christina Curtis, Mariano J. Alvarez, Vessela N. Kristensen, Samuel Aparicio, Anne-Lise Børresen-Dale, Carlos Caldas, Andrea Califano, Stephen H. Friend, Trey Ideker, Eric E. Schadt, Gustavo Stolovitzky, Adam A. Margolin
PLoS Comput. Biol.21
2011 PiNGO: a Cytoscape plugin to find candidate genes in biological networks
abstract
UNLABELLED: PiNGO is a tool to screen biological networks for candidate genes, i.e. genes predicted to be involved in a biological process of interest. The user can narrow the search to genes with particular known functions or exclude genes belonging to particular functional classes. PiNGO provides support for a wide range of organisms and Gene Ontology classification schemes, and it can easily be customized for other organisms and functional classifications. PiNGO is implemented as a plugin for Cytoscape, a popular network visualization platform. AVAILABILITY: PiNGO is distributed as an open-source Java package under the GNU General Public License (http://www.gnu.org/), and can be downloaded via the Cytoscape plugin manager. A detailed user guide and tutorial are available on the PiNGO website (http://www.psb.ugent.be/esb/PiNGO.
Michael E. Smoot, Keiichiro Ono, Trey Ideker, Steven Maere
Bioinform.3
2011 Cytoscape 2.8: new features for data integration and network visualization
abstract
UNLABELLED: Cytoscape is a popular bioinformatics package for biological network visualization and data integration. Version 2.8 introduces two powerful new features--Custom Node Graphics and Attribute Equations--which can be used jointly to greatly enhance Cytoscape's data integration and visualization capabilities. Custom Node Graphics allow an image to be projected onto a node, including images generated dynamically or at remote locations. Attribute Equations provide Cytoscape with spreadsheet-like functionality in which the value of an attribute is computed dynamically as a function of other attributes and network properties. AVAILABILITY AND IMPLEMENTATION: Cytoscape is a desktop Java application released under the Library Gnu Public License (LGPL). Binary install bundles and source code for Cytoscape 2.8 are available for download from http://cytoscape.org.
Michael E. Smoot, Keiichiro Ono, Johannes Ruscheinski, Peng-Liang Wang, Trey Ideker
Bioinform.5
2011 Protein Networks as Logic Functions in Development and Cancer
abstract
Many biological and clinical outcomes are based not on single proteins, but on modules of proteins embedded in protein networks. A fundamental question is how the proteins within each module contribute to the overall module activity. Here, we study the modules underlying three representative biological programs related to tissue development, breast cancer metastasis, or progression of brain cancer, respectively. For each case we apply a new method, called Network-Guided Forests, to identify predictive modules together with logic functions which tie the activity of each module to the activity of its component genes. The resulting modules implement a diverse repertoire of decision logic which cannot be captured using the simple approximations suggested in previous work such as gene summation or subtraction. We show that in cancer, certain combinations of oncogenes and tumor suppressors exert competing forces on the system, suggesting that medical genetics should move beyond cataloguing individual cancer genes to cataloguing their combinatorial logic.
Janusz Dutkowski, Trey Ideker
PLoS Comput. Biol.2
2010 Nonlinear dimension reduction and clustering by Minimum Curvilinearity unfold neuropathic pain and tissue embryological classes
abstract
MOTIVATION: Nonlinear small datasets, which are characterized by low numbers of samples and very high numbers of measures, occur frequently in computational biology, and pose problems in their investigation. Unsupervised hybrid-two-phase (H2P) procedures-specifically dimension reduction (DR), coupled with clustering-provide valuable assistance, not only for unsupervised data classification, but also for visualization of the patterns hidden in high-dimensional feature space. METHODS: 'Minimum Curvilinearity' (MC) is a principle that-for small datasets-suggests the approximation of curvilinear sample distances in the feature space by pair-wise distances over their minimum spanning tree (MST), and thus avoids the introduction of any tuning parameter. MC is used to design two novel forms of nonlinear machine learning (NML): Minimum Curvilinear embedding (MCE) for DR, and Minimum Curvilinear affinity propagation (MCAP) for clustering. RESULTS: Compared with several other unsupervised and supervised algorithms, MCE and MCAP, whether individually or combined in H2P, overcome the limits of classical approaches. High performance was attained in the visualization and classification of: (i) pain patients (proteomic measurements) in peripheral neuropathy; (ii) human organ tissues (genomic transcription factor measurements) on the basis of their embryological origin. CONCLUSION: MC provides a valuable framework to estimate nonlinear distances in small datasets. Its extension to large datasets is prefigured for novel NMLs. Classification of neuropathic pain by proteomic profiles offers new insights for future molecular and systems biology characterization of pain. Improvements in tissue embryological classification refine results obtained in an earlier study, and suggest a possible reinterpretation of skin attribution as mesodermal. AVAILABILITY: https://sites.google.com/site/carlovittoriocannistraci/home.
Carlo V. Cannistraci, Timothy Ravasi, Franco Maria Montevecchi, Trey Ideker, Massimo Alessio
Bioinform.4
2010 Evidence mining and novelty assessment of protein-protein interactions with the ConsensusPathDB plugin for Cytoscape
abstract
SUMMARY: Protein-protein interaction detection methods are applied on a daily basis by molecular biologists worldwide. After generating a set of potential interactions, biologists face the problem of highlighting the ones that are novel and collecting evidence with respect to literature and annotation. This task can be as tedious as searching for every predicted interaction in several interaction data repositories, or manually screening the scientific literature. To facilitate the task of evidence mining and novelty assessment of protein-protein interactions, we have developed a Cytoscape plugin that automatically mines publication references, database references, interaction detection method descriptions and pathway annotation for a user-supplied network of interactions. The basis for the annotation is ConsensusPathDB-a meta-database that integrates numerous protein-protein, signaling, metabolic and gene regulatory interaction repositories for currently three species: Homo sapiens, Saccharomyces cerevisiae and Mus musculus. AVAILABILITY: The ConsensusPathDB plugin for Cytoscape (version 2.7.0 or later) can be installed within Cytoscape on a major operating system (Windows, Mac OS, Unix/Linux) with Sun Java 1.5 or later installed through Cytoscape's Plugin manager (category 'Network and Attribute I/O'). The plugin is freely available for download on the ConsensusPathDB web site (http://cpdb.molgen.mpg.de). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Konstantin Pentchev, Keiichiro Ono, Ralf Herwig, Trey Ideker, Atanas Kamburov
Bioinform.4
2009 Papers on normalization, variable selection, classification or clustering of microarray data
abstract
Over the last decade or so, there have been large numbers of methods published on approaches for normalization, variable (gene) selection, classification and clustering of microarray data. As indicated in the scope document for Bioinformatics, this requires papers describing new methods for these problems to meet a very high standard, showing important improvement in results for real biological data, as well as novelty. In this editorial, we describe some standards that need to be met for papers in these areas to be seriously considered. We ask that prospective authors consider these points carefully before submission of their papers to Bioinformatics. The role of simulation: Simulation can be useful in investigating the properties of various methods of data analysis. Yet, there are important barriers to credible use of simulation in microarray studies, largely due to what we do not know about the statistical distribution of measured gene expression levels. First, the distribution across transcripts of true expression values is dependent on the biological state of the tissue or cell, and for a given state this is unknown, even in distributional form, and may further exhibit gene- and platform-specific effects. Second, the correlation within biological replicates of true expression is unknown, and is likely unknowable in detail given that it is expressed by a correlation matrix with on the order of a billion entries. Third, the distribution of changes from one biological state to another is unknown. Fourth, the correlation in observational errors in gene expression across genes is unknown and similarly probably unknowable in detail. On the other hand, the measurement error for a given transcript has been well described by several authors (Ideker et al., 2001; Rocke and Durbin, 2001). Given this gap between knowledge and simulation specification, it is likely that any new method can be shown to be superior to some other method(s) by careful choice of simulation parameters, since simulations often include biases in the distributions selected and in other assumptions of the models. Thus, while simulation may still be worthwhile, and a useful tool for exploring robustness and parameter space of a new method, it is insufficient evidence for superiority of a new method without substantial support from significant improvement in results from analysis of real data. Normalization: Normalization necessarily involves a trade-off between its positive role in reducing variability, and its potentially negative role in increasing bias. There are a number of good image analysis, preprocessing, transformation and normalization methods extant for single- and dual-color DNA microarrays. To show that a new method is better requires comparison demonstrating that results in differential expression analysis, classification or clustering are better with the new normalization method than with previous methods. Not one but several previous methods should be chosen for comparison including the most widely used approaches. Several datasets should be used, including spike-in and dilution studies when feasible, as well as ‘real’ biological datasets. Showing that more genes are differentially expressed using a normalization method is not compelling evidence of superiority without a good estimate of the false-positive rate or a compelling biological analysis of the resulting differentially expressed genes. Variable selection: Typically, new variable selection methods are proposed as part of a classification or clustering strategy, and demonstrating superiority of the variable selection method usually means demonstrating superiority of the combined methodology. It is quite important that metrics for evaluation be used that are robust to intra-array correlations and variable selection artifacts. For example, in cross-validation studies in which variable selection is followed by a classification method, selection of variables using all the data and then cross-validating the classification accuracy introduces substantial bias, making classification methods appear more accurate than they really are (Ambroise and McLachlan, 2002). It is important that any method be compared with several of the most widely used existing methods, including baseline approaches such as filtering by t-score or forward stepwise analysis. Such comparisons should be performed on more than one biological dataset. Further, the method must demonstrate significant improvement over existing methods; incremental improvements will not be considered of sufficient interest to warrant review. Classification and prediction: New classification or prediction methods for microarray data enter a crowded arena. From long-standing techniques such as logistic regression and linear discriminant analysis to the more modern support vector machines and neural networks, most known classification methods have already been applied to microarray data. To show that a newly proposed classification method is a real advance, a substantial improvement in performance needs to be shown over a reasonable selection of existing datasets and methods, including commonly used or simple methods. This is because, consciously or subconsciously, the developer of a new method optimizes its characteristics against the datasets to be used for evaluation. Variable selection and parameter choice for all methods needs to be done strictly in the training set (whether there is one training set or many as in cross validation). Resampling methods like permuting the class labels on the arrays or the bootstrap can be used to provide robust estimates of the significance of differential expression, but do not in themselves give estimates of classification performance except to show that the performance is better than chance. Experience shows that there is considerable noise in classification accuracy experiments, so modest increases in achieved accuracy are usually not convincing. Experience also shows that classification performance in a microarray problem depends strongly on the dataset, and less on the variable selection and classification methods. More than modest differences are required to excite interest in a new method. Authors should keep in mind the ‘No Free Lunch Theorems’ of Wolpert and Macready (1997) which demonstrated that there is no optimization/classification method that outperforms all others in all circumstances (Wolpert, 1996). Clustering: Demonstrating superiority of a clustering method is in many ways more difficult than demonstrating superiority in a classification method. Usually, there is no ground truth against which to compare the clustering results. Defining a criterion (e.g. the Rand index) and showing that a clustering method achieves better scores on this criterion is often not compelling, since such criteria are easily optimized (again, consciously or subconsciously) to ensure superiority. For reasons discussed above, simulation is also not usually sufficient. Ideally, a new clustering method would demonstrate novel biological insights or some attractive statistical properties not available from previous methods, including several commonly used methods. Requiring new biological findings is a difficult standard, but a necessary one to insure that new published methods are useful and likely to be used. To conclude, microarrays remain a useful technology to address a wide array of biological problems and the optimal analysis of these data to extract meaningful results still pose many bioinformatics challenges. However, with a number of successful methods already addressing the well-established microarray data analysis problems, publication of new methods in this area requires either identification of a new challenge and formulation of a new problem or development of a substantially better methodology then those existing that can be benchmarked on a variety of datasets. We hope that suggestions provided above for evaluation and validation of such new methods would increase the likelihood of them supporting biological discoveries in the future. Funding: DMR to NIH grants P42-ES04699 and R01-HG003352.
David M. Rocke, Trey Ideker, Olga G. Troyanskaya, John Quackenbush, Joaquín Dopazo
Bioinform.2
2008 NetworkBLAST: comparative analysis of protein networks
abstract
UNLABELLED: The identification of protein complexes is a fundamental challenge in interpreting protein-protein interaction data. Cross-species analysis allows coping with the high levels of noise that are typical to these data. The NetworkBLAST web-server provides a platform for identifying protein complexes in protein-protein interaction networks. It can analyze a single network or two networks from different species. In the latter case, NetworkBLAST outputs a set of putative complexes that are evolutionarily conserved across the two networks. AVAILABILITY: NetworkBLAST is available as web-server at: www.cs.tau.ac.il/~roded/networkblast.htm.
Maxim Kalaev, Michael E. Smoot, Trey Ideker, Roded Sharan
Bioinform.3
2008 Correcting for gene-specific dye bias in DNA microarrays using the method of maximum likelihood
abstract
MOTIVATION: In two-color microarray experiments, well-known differences exist in the labeling and hybridization efficiency of Cy3 and Cy5 dyes. Previous reports have revealed that these differences can vary on a gene-by-gene basis, an effect termed gene-specific dye bias. If uncorrected, this bias can influence the determination of differentially expressed genes. RESULTS: We show that the magnitude of the bias scales multiplicatively with signal intensity and is dependent on which nucleotide has been conjugated to the fluorescent dye. A method is proposed to account for gene-specific dye bias within a maximum-likelihood error modeling framework. Using two different labeling schemes, we show that correcting for gene-specific dye bias results in the superior identification of differentially expressed genes within this framework. Improvement is also possible in related ANOVA approaches. AVAILABILITY: A software implementation of this procedure is freely available at http://cellcircuits.org/VERA
Ryan M. Kelley, Hoda Feizi, Trey Ideker
Bioinform.3
2008 Functional Maps of Protein Complexes from Quantitative Genetic Interaction Data
abstract
Recently, a number of advanced screening technologies have allowed for the comprehensive quantification of aggravating and alleviating genetic interactions among gene pairs. In parallel, TAP-MS studies (tandem affinity purification followed by mass spectroscopy) have been successful at identifying physical protein interactions that can indicate proteins participating in the same molecular complex. Here, we propose a method for the joint learning of protein complexes and their functional relationships by integration of quantitative genetic interactions and TAP-MS data. Using 3 independent benchmark datasets, we demonstrate that this method is >50% more accurate at identifying functionally related protein pairs than previous approaches. Application to genes involved in yeast chromosome organization identifies a functional map of 91 multimeric complexes, a number of which are novel or have been substantially expanded by addition of new subunits. Interestingly, we find that complexes that are enriched for aggravating genetic interactions (i.e., synthetic lethality) are more likely to contain essential genes, linking each of these interactions to an underlying mechanism. These results demonstrate the importance of both large-scale genetic and physical interaction data in mapping pathway architecture and function.
Sourav Bandyopadhyay, Ryan M. Kelley, Nevan J. Krogan, Trey Ideker
PLoS Comput. Biol.4
2008 Inferring Pathway Activity toward Precise Disease Classification
abstract
The advent of microarray technology has made it possible to classify disease states based on gene expression profiles of patients. Typically, marker genes are selected by measuring the power of their expression profiles to discriminate among patients of different disease states. However, expression-based classification can be challenging in complex diseases due to factors such as cellular heterogeneity within a tissue sample and genetic heterogeneity across patients. A promising technique for coping with these challenges is to incorporate pathway information into the disease classification procedure in order to classify disease based on the activity of entire signaling pathways or protein complexes rather than on the expression levels of individual genes or proteins. We propose a new classification method based on pathway activities inferred for each patient. For each pathway, an activity level is summarized from the gene expression levels of its condition-responsive genes (CORGs), defined as the subset of genes in the pathway whose combined expression delivers optimal discriminative power for the disease phenotype. We show that classifiers using pathway activity achieve better performance than classifiers based on individual gene expression, for both simple and complex case-control studies including differentiation of perturbed from non-perturbed cells and subtyping of several different kinds of cancer. Moreover, the new method outperforms several previous approaches that use a static (i.e., non-conditional) definition of pathways. Within a pathway, the identified CORGs may facilitate the development of better diagnostic markers and the discovery of core alterations in human disease.
Han-Yu Chuang, Jong-Won Kim 0001, Trey Ideker, Doheon Lee
PLoS Comput. Biol.4
2006 Bioinformatics in the human interactome project
abstract
‘In the early days of the Human Interactome Project, a meeting was organized…’. Perhaps, a few years from now, newspapers will describe in those terms how straightforward it was to plan the large-scale mapping of protein interactions in human and other model organisms. Scientists attending the second Cold Spring Harbor Laboratory/Wellcome Trust symposium on ‘Interactome Networks’1 know well that things are not that easy. Important scientific, technical and sociological issues remain before the ‘Human Interactome Project’ can be considered on its way. But things are definitely moving. At the meeting, Marc Vidal proposed some concrete goals for such a project: ‘To produce hundreds of different sets of cloned ORFs (ORFeomes) and 100 million interactions with a 1–5% false positive rate. To add directionality and signs to the interactions (i.e. activation or inhibition), and to study the variation of the interactions associated with diseases’. Nonetheless, the community still has much to decide. It will still have to agree on these goals, to subdivide the project into recognizable milestones, to set a time line for achieving these milestones, to associate cost to each of the operations, and perhaps most importantly, to obtain funding for such an ambitious endeavor. But there is little doubt that having clear goals will help strengthen the ties between researchers in this already very active community, as well as to engage new partners and grant agencies. As with the Human Genome Project, bioinformatics and computational biology will be of profound importance to any protein interaction mapping effort. Following the presentations during the meeting (for a recent review see Sharan and Ideker, 2006) an early and essential bioinformatics task will be the creation of database standards (i.e. the IMEX interaction database standard and the emerging Biopax bio-pathways standard) and analysis/visualization platforms, such as the one provided by the Cytoscape project. It was also interesting to realize the progress that has already been made in the analysis of the structure, function and properties of protein interaction networks (as well as gene control and metabolic networks), even while the number of reliable datasets is still small. It is also rewarding to see also how the first large-scale simulations based on protein interaction data are becoming a reality. These efforts in analysis and simulation are proceeding in parallel with those dedicated to the prediction of new interactions, modules, motifs and functional properties (phenotypes, diseases and others), in most cases by integrating complementary sources of information on functional and structural interactions. All this domain of emerging ‘Network Biology’ offers a direct connection between computational and the experimental analysis. In this respect, two questions emerge from the meeting as critical for the future: (1) the development of methods able to distinguish physical from functional interactions (and/or different types of physical interaction) and (2) the mapping of the details of the physical interactions (i.e. interacting residues) and other information that is necessary for the interpretation of the variation data (SNPs) and for experimental manipulation of interaction networks. Finally, an interesting controversy arose during the meeting that might have consequences for our bioinformatics community. Some argue that it will be more effective to concentrate all efforts into scale-up of the experimental proteomics technology, postponing the bioinformatics analysis to a second phase once the underlying data are fully (or at least mostly) complete. On the contrary, we think that, it is essential to continuously support the development of the methods that will be required for the interpretation of the Human Interactome Project, including network alignments, annotation, analysis and others. The analogy with the Human Genome Project can be useful here. In that case even if the basic alignment techniques were ready since the 70's when the genome sequencing emerged basic bioinformatic technologies were not available (c.f. just remember the challenging analysis of the first bacterial genome in 1995 (Casari et al., 1995), or the struggle to assemble and represent the first draft of the human genome (Istrail et al., 2004). We are convinced that by pushing in parallel experimental and computational developments we can prepare in a more effective way the future of this area of research. Indeed, much of the current interest in large-scale proteomics is related with the impact that the early computational analysis of the first (and imperfect) datasets have had (i.e. the first ‘scale free’ and ‘motif discovery’ papers of Barabasi (Jeong et al., 2000) and Alon (Shen-Orr et al., 2002) teams have captured the imagination of biologist, physicists and theoreticians like few other problems in molecular biology have). Moreover, integrative and computational approaches have already been indispensable for assessing data quality and scoring confidence in specific interactions as well as whole interaction datasets. Finally, at a practical level, what biologists see as a result of large-scale proteomics are computational representations based on the data provided by databases. Therefore, a successful Human Proteome Project depends intimately on ongoing developments in bioinformatics, as they proceed in parallel with the large-scale experiments.
Trey Ideker, Alfonso Valencia
Bioinform.1
2006 A direct comparison of protein interaction confidence assignment schemes
abstract
BACKGROUND: Recent technological advances have enabled high-throughput measurements of protein-protein interactions in the cell, producing large protein interaction networks for various species at an ever-growing pace. However, common technologies like yeast two-hybrid may experience high rates of false positive detection. To combat false positive discoveries, a number of different methods have been recently developed that associate confidence scores with protein interactions. Here, we perform a rigorous comparative analysis and performance assessment among these different methods. RESULTS: We measure the extent to which each set of confidence scores correlates with similarity of the interacting proteins in terms of function, expression, pattern of sequence conservation, and homology to interacting proteins in other species. We also employ a new metric, the Signal-to-Noise Ratio of protein complexes embedded in each network, to assess the power of the different methods. Seven confidence assignment schemes, including those of Bader et al., Deane et al., Deng et al., Sharan et al., and Qi et al., are compared in this work. CONCLUSION: Although the performance of each assignment scheme varies depending on the particular metric used for assessment, we observe that Deng et al. yields the best performance overall (in three out of four viable measures). Importantly, we also find that utilizing any of the probability assignment schemes is always more beneficial than assuming all observed interactions to be true or equally likely.
Silpa Suthram, Tomer Shlomi, Eytan Ruppin, Roded Sharan, Trey Ideker
BMC Bioinform.5
2006 Integrated Assessment and Prediction of Transcription Factor Binding
abstract
Systematic chromatin immunoprecipitation (chIP-chip) experiments have become a central technique for mapping transcriptional interactions in model organisms and humans. However, measurement of chromatin binding does not necessarily imply regulation, and binding may be difficult to detect if it is condition or cofactor dependent. To address these challenges, we present an approach for reliably assigning transcription factors (TFs) to target genes that integrates many lines of direct and indirect evidence into a single probabilistic model. Using this approach, we analyze publicly available chIP-chip binding profiles measured for yeast TFs in standard conditions, showing that our model interprets these data with significantly higher accuracy than previous methods. Pooling the high-confidence interactions reveals a large network containing 363 significant sets of factors (TF modules) that cooperate to regulate common target genes. In addition, the method predicts 980 novel binding interactions with high confidence that are likely to occur in so-far untested conditions. Indeed, using new chIP-chip experiments we show that predicted interactions for the factors Rpn4p and Pdr1p are observed only after treatment of cells with methyl-methanesulfonate, a DNA-damaging agent. We outline the first approach for consistently integrating all available evidences for TF-target interactions and we comprehensively identify the resulting TF module hierarchy. Prioritizing experimental conditions for each factor will be especially important as increasing numbers of chIP-chip assays are performed in complex organisms such as humans, for which "standard conditions" are ill defined.
Andreas Beyer, Christopher T. Workman, Jens Hollunder, Dörte Radke, Ulrich Möller, Thomas Wilhelm 0003, Trey Ideker
PLoS Comput. Biol.7
2005 Efficient Algorithms for Detecting Signaling Pathways in Protein Interaction Networks
Trey Ideker, Richard M. Karp, Roded Sharan
RECOMB2
2004 Identification of protein complexes by comparative analysis of yeast and bacterial protein interaction data
abstract
Mounting evidence shows that many protein complexes are conserved in evolution. Here we use conservation to find complexes that are common to yeast S. Cerevisiae and bacteria H. pylori. Our analysis combines protein interaction data, that are available for each of the two species, and orthology information based on protein sequence comparison. We develop a detailed probabilistic model for protein complexes in a single species, and a model for the conservation of complexes between two species. Using these models, one can recast the question of finding conserved complexes as a problem of searching for heavy subgraphs in an edge- and node-weighted graph, whose nodes are orthologous protein pairs.We tested this approach on the data currently available for yeast and bacteria and detected 11 significantly conserved complexes. Several of these complexes match very well with prior experimental knowledge on complexes in yeast only, and serve for validation of our methodology. The complexes suggest new functions for a variety of uncharacterized proteins. By identifying a conserved complex whose yeast proteins function predominantly in the nuclear pore complex, we propose that the corresponding bacterial proteins function as a coherent cellular membrane transport system. We also compare our results to two alternative methods for detecting complexes, and demonstrate that our methodology obtains a much higher specificity.
Roded Sharan, Trey Ideker, Brian P. Kelley, Ron Shamir, Richard M. Karp
RECOMB2
2002 Discovering regulatory and signalling circuits in molecular interaction networks
abstract
Abstract Motivation: In model organisms such as yeast, large databases of protein–protein and protein-DNA interactions have become an extremely important resource for the study of protein function, evolution, and gene regulatory dynamics. In this paper we demonstrate that by integrating these interactions with widely-available mRNA expression data, it is possible to generate concrete hypotheses for the underlying mechanisms governing the observed changes in gene expression. To perform this integration systematically and at large scale, we introduce an approach for screening a molecular interaction network to identify active subnetworks, i.e., connected regions of the network that show significant changes in expression over particular subsets of conditions. The method we present here combines a rigorous statistical measure for scoring subnetworks with a search algorithm for identifying subnetworks with high score. Results: We evaluated our procedure on a small network of 332 genes and 362 interactions and a large network of 4160 genes containing all 7462 protein–protein and protein-DNA interactions in the yeast public databases. In the case of the small network, we identified five significant subnetworks that covered 41 out of 77 (53%) of all significant changes in expression. Both network analyses returned several top-scoring subnetworks with good correspondence to known regulatory mechanisms in the literature. These results demonstrate how large-scale genomic approaches may be used to uncover signalling and regulatory pathways in a systematic, integrative fashion. Availability: The methods presented in this paper are implemented in the Cytoscape software package which is available to the academic community at http://www.cytoscape.org. Contact: [email protected] Keywords: molecular interactions; gene expression; data integration; simulated annealing; Monte carlo methods. *To whom correspondence should be addressed.
Trey Ideker, Owen Ozier, Benno Schwikowski, Andrew F. Siegel
ISMB1