VLDB 2026 Research / reviewers in the wild / expert
John Quackenbush
dblp:68/1152
· DBLP profile ↗
46ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-2702-5879ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FACT: Gated Fusion-Augmented Causal Mask Transformer for Pseudotime Analysis
Chengjie Zheng, Iris Shen, John Quackenbush, Viola Fanfani, Wei Ding 0003, Ping Chen 0001 |
IEEE Big Data | 4 |
| 2024 | Pruning neural network models for gene regulatory dynamics using data and domain knowledgeabstractThe practical utility of machine learning models in the sciences often hinges on their interpretability. It is common to assess a model's merit for scientific discovery, and thus novel insights, by how well it aligns with already available domain knowledge - a dimension that is currently largely disregarded in the comparison of neural network models. While pruning can simplify deep neural network architectures and excels in identifying sparse models, as we show in the context of gene regulatory network inference, state-of-the-art techniques struggle with biologically meaningful structure learning. To address this issue, we propose DASH, a generalizable framework that guides network pruning by using domain-specific structural information in model fitting and leads to sparser, better interpretable models that are more robust to noise. Using both synthetic data with ground truth information, as well as real-world gene expression data, we show that DASH, using knowledge about gene interaction partners within the putative regulatory network, outperforms general pruning methods by a large margin and yields deeper insights into the biological systems being studied. Intekhab Hossain, Jonas Fischer, Rebekka Burkholz, John Quackenbush |
NeurIPS | 4 |
| 2024 | BONOBO: Bayesian Optimized Sample-Specific Networks Obtained by Omics Data
Enakshi Saha, Viola Fanfani, Panagiotis Mandros, Marouen Ben Guebila, Jonas Fischer, Katherine H. Shutta, Kimberly Glass, Dawn L. DeMeo, Camila Miranda Lopes-Ramos, John Quackenbush |
RECOMB | 10 |
| 2024 | SpaCeNet: Spatial Cellular Networks from Omics Data
Stefan Schrod, Niklas Lück, Robert Lohmayer, Stefan Solbrig, Tina Wipfler, Katherine H. Shutta, Marouen Ben Guebila, Andreas Schäfer 0005, Tim Beißbarth, Helena U. Zacharias, Peter J. Oefner, John Quackenbush, Michael Altenbuchinger |
RECOMB | 12 |
| 2024 | Higher-order correction of persistent batch effects in correlation networksabstractMOTIVATION: Systems biology analyses often use correlations in gene expression profiles to infer co-expression networks that are then used as input for gene regulatory network inference or to identify functional modules of co-expressed or putatively co-regulated genes. While systematic biases, including batch effects, are known to induce spurious associations and confound differential gene expression analyses (DE), the impact of batch effects on gene co-expression has not been fully explored. Methods have been developed to adjust expression values, ensuring conditional independence of mean and variance from batch or other covariates for each gene, resulting in improved fidelity of DE analysis. However, such adjustments do not address the potential for spurious differential co-expression (DC) between groups. Consequently, uncorrected, artifactual DC can skew the correlation structure, leading to the identification of false, non-biological associations, even when the input data are corrected using standard batch correction. RESULTS: In this work, we demonstrate the persistence of confounders in covariance after standard batch correction using synthetic and real-world gene expression data examples. We then introduce Co-expression Batch Reduction Adjustment (COBRA), a method for computing a batch-corrected gene co-expression matrix based on estimating a conditional covariance matrix. COBRA estimates a reduced set of parameters expressing the co-expression matrix as a function of the sample covariates, allowing control for continuous and categorical covariates. COBRA is computationally efficient, leveraging the inherently modular structure of genomic data to estimate accurate gene regulatory associations and facilitate functional analysis for high-dimensional genomic data. AVAILABILITY AND IMPLEMENTATION: COBRA is available under the GLP3 open source license in R and Python in netZoo (https://netzoo.github.io). Soel Micheletti, Daniel Schlauch, John Quackenbush, Marouen Ben Guebila |
Bioinform. | 3 |
| 2021 | Cascade Size Distributions: Why They Matter and How to Compute Them EfficientlyabstractCascade models are central to understanding, predicting, and controlling epidemic spreading and information propagation. Related optimization, including influence maximization, model parameter inference, or the development of vaccination strategies, relies heavily on sampling from a model. This is either inefficient or inaccurate. As alternative, we present an efficient message passing algorithm that computes the probability distribution of the cascade size for the Independent Cascade Model on weighted directed networks and generalizations. Our approach is exact on trees but can be applied to any network topology. It approximates locally tree-like networks well, scales to large networks, and can lead to surprisingly good performance on more dense networks, as we also exemplify on real world data. Rebekka Burkholz, John Quackenbush |
AAAI | 2 |
| 2021 | Gene Regulatory Network Inference as Relaxed Graph MatchingabstractBipartite network inference is a ubiquitous problem across disciplines. One important example in the field molecular biology is gene regulatory network inference. Gene regulatory networks are an instrumental tool aiding in the discovery of the molecular mechanisms driving diverse diseases, including cancer. However, only noisy observations of the projections of these regulatory networks are typically assayed. In an effort to better estimate regulatory networks from their noisy projections, we formulate a non-convex but analytically tractable optimization problem called OTTER. This problem can be interpreted as relaxed graph matching between the two projections of the bipartite network. OTTER's solutions can be derived explicitly and inspire a spectral algorithm, for which we provide network recovery guarantees. We also provide an alternative approach based on gradient descent that is more robust to noise compared to the spectral algorithm. Interestingly, this gradient descent approach resembles the message passing equations of an established gene regulatory network inference method, PANDA. Using three cancer-related data sets, we show that OTTER outperforms state-of-the-art inference methods in predicting transcription factor binding to gene regulatory regions. To encourage new graph matching applications to this problem, we have made all networks and validation data publicly available. Deborah A. Weighill, Marouen Ben Guebila, Camila Miranda Lopes-Ramos, Kimberly Glass, John Quackenbush, John Platig, Rebekka Burkholz |
AAAI | 5 |
| 2021 | Scaling up Continuous-Time Markov Chains Helps Resolve UnderspecificationabstractModeling the time evolution of discrete sets of items (e.g., genetic mutations) is a fundamental problem in many biomedical applications. We approach this problem through the lens of continuous-time Markov chains, and show that the resulting learning task is generally underspecified in the usual setting of cross-sectional data. We explore a perhaps surprising remedy: including a number of additional independent items can help determine time order, and hence resolve underspecification. This is in sharp contrast to the common practice of limiting the analysis to a small subset of relevant items, which is followed largely due to poor scaling of existing methods. To put our theoretical insight into practice, we develop an approximate likelihood maximization method for learning continuous-time Markov chains, which can scale to hundreds of items and is orders of magnitude faster than previous methods. We demonstrate the effectiveness of our approach on synthetic and real cancer data. Alkis Gotovos, Rebekka Burkholz, John Quackenbush, Stefanie Jegelka |
NeurIPS | 3 |
| 2020 | Catalysis Clustering with GAN by Incorporating Domain KnowledgeabstractClustering is an important unsupervised learning method with serious challenges when data is sparse and high-dimensional. Generated clusters are often evaluated with general measures, which may not be meaningful or useful for practical applications and domains. Using a distance metric, a clustering algorithm searches through the data space, groups close items into one cluster, and assigns far away samples to different clusters. In many real-world applications, the number of dimensions is high and data space becomes very sparse. Selection of a suitable distance metric is very difficult and becomes even harder when categorical data is involved. Moreover, existing distance metrics are mostly generic, and clusters created based on them will not necessarily make sense to domain-specific applications. One option to address these challenges is to integrate domain-defined rules and guidelines into the clustering process. In this work we propose a GAN-based approach called Catalysis Clustering to incorporate domain knowledge into the clustering process. With GANs we generate catalysts, which are special synthetic points drawn from the original data distribution and verified to improve clustering quality when measured by a domain-specific metric. We then perform clustering analysis using both catalysts and real data. Final clusters are produced after catalyst points are removed. Experiments on two challenging real-world datasets clearly show that our approach is effective and can generate clusters that are meaningful and useful for real-world applications. Olga Andreeva, Wei Li 0121, Wei Ding 0003, Marieke L. Kuijjer, John Quackenbush, Ping Chen 0001 |
KDD | 5 |
| 2020 | A Novel Deep Learning Model by Stacking Conditional Restricted Boltzmann Machine and Deep Neural NetworkabstractA real-world system often exhibits complex dynamics arising from interaction among its subunits. In machine learning and data mining, these interactions are usually formulated as dependency and correlation among system variables. Similar to Convolution Neural Network dealing with spatially correlated features and Recurrent Neural Network with temporally correlated features, in this paper we present a novel deep learning model to tackle functionally interactive features by stacking a Conditional Restricted Boltzmann Machine and a Deep Neural Network (CRBM-DNN). Variables with their dependency relationships are organized into a bipartite graph, which is further converted into a Restricted Boltzmann Machine conditioned by domain knowledge. We integrate this CRBM and a DNN into one deep learning model constrained by one overall cost function. CRBM-DNN can solve both supervised and unsupervised learning problems. Compared to a regular neural network of the same size, CRBM-DNN has fewer parameters so they require fewer training samples. We perform extensive comparative studies with a large number of supervised learning and unsupervised learning methods using several challenging real-world datasets, and achieve significant superior performance. Tianyu Kang, Ping Chen 0001, John Quackenbush, Wei Ding 0003 |
KDD | 3 |
| 2020 | PUMA: PANDA Using MicroRNA AssociationsabstractMOTIVATION: Conventional methods to analyze genomic data do not make use of the interplay between multiple factors, such as between microRNAs (miRNAs) and the messenger RNA (mRNA) transcripts they regulate, and thereby often fail to identify the cellular processes that are unique to specific tissues. We developed PUMA (PANDA Using MicroRNA Associations), a computational tool that uses message passing to integrate a prior network of miRNA target predictions with target gene co-expression information to model genome-wide gene regulation by miRNAs. We applied PUMA to 38 tissues from the Genotype-Tissue Expression project, integrating RNA-Seq data with two different miRNA target predictions priors, built on predictions from TargetScan and miRanda, respectively. We found that while target predictions obtained from these two different resources are considerably different, PUMA captures similar tissue-specific miRNA-target regulatory interactions in the different network models. Furthermore, the tissue-specific functions of miRNAs we identified based on regulatory profiles (available at: https://kuijjer.shinyapps.io/puma_gtex/) are highly similar between networks modeled on the two target prediction resources. This indicates that PUMA consistently captures important tissue-specific miRNA regulatory processes. In addition, using PUMA we identified miRNAs regulating important tissue-specific processes that, when mutated, may result in disease development in the same tissue. AVAILABILITY AND IMPLEMENTATION: PUMA is available in C++, MATLAB and Python on GitHub (https://github.com/kuijjerlab and https://netzoo.github.io/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Marieke L. Kuijjer, Maud Fagny, Alessandro Marin, John Quackenbush, Kimberly Glass |
Bioinform. | 4 |
| 2019 | Identification of differentially expressed gene sets using the Generalized Berk-Jones statisticabstractMOTIVATION: Cancer genomics studies frequently aim to identify genes that are differentially expressed between clinically distinct patient subgroups, generally by testing single genes one at a time. However, the results of any individual transcriptomic study are often not fully reproducible. A particular challenge impeding statistical analysis is the difficulty of distinguishing between differential expression comprising part of the genomic disease etiology and that induced by downstream effects. More robust analytical approaches that are well-powered to detect potentially causative genes, are less prone to discovering spurious associations, and can deliver reproducible findings across different studies are needed. RESULTS: We propose a set-based procedure for testing of differential expression and show that this set-based approach can produce more robust results by aggregating information across multiple, correlated genomic markers. Specifically, we adapt the Generalized Berk-Jones statistic to test for the transcription factors that may contribute to the progression of estrogen receptor positive breast cancer. We demonstrate the ability of our method to produce reproducible findings by applying the same analysis to 21 publicly available datasets, producing a similar list of significant transcription factors across most studies. Our Generalized Berk-Jones approach produces results that show improved consistency over three set-based testing algorithms: Generalized Higher Criticism, Gene Set Analysis and Gene Set Enrichment Analysis. AVAILABILITY AND IMPLEMENTATION: Data are in the MetaGxBreast R package. Code is available at github.com/ryanrsun/gaynor_sun_GBJ_breast_cancer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sheila M. Gaynor, Ryan Sun, Xihong Lin, John Quackenbush |
Bioinform. | 4 |
| 2018 | Clustering on Sparse Data in Non-overlapping Feature Space with Applications to Cancer SubtypingabstractThis paper presents a new algorithm, Reinforced and Informed Network-based Clustering(RINC), for finding unknown groups of similar data objects in sparse and largely non-overlapping feature space where a network structure among features can be observed. Sparse and non-overlapping unlabeled data become increasingly common and available especially in text mining and biomedical data mining. RINC inserts a domain informed model into a modelless neural network. In particular, our approach integrates physically meaningful feature dependencies into the neural network architecture and soft computational constraint. Our learning algorithm efficiently clusters sparse data through integrated smoothing and sparse auto-encoder learning. The informed design requires fewer samples for training and at least part of the model becomes explainable. The architecture of the reinforced network layers smooths sparse data over the network dependency in the feature space. Most importantly, through back-propagation, the weights of the reinforced smoothing layers are simultaneously constrained by the remaining sparse auto-encoder layers that set the target values to be equal to the raw inputs. Empirical results demonstrate that RINC achieves improved accuracy and renders physically meaningful clustering results. Tianyu Kang, Kourosh Zarringhalam, Marieke L. Kuijjer, Ping Chen 0001, John Quackenbush, Wei Ding 0003 |
ICDM | 5 |
| 2017 | Estimating gene regulatory networks with pandaRabstractAbstract PANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that begins with a model of transcription factor–target gene interactions and uses message passing to update the network model given available transcriptomic and protein–protein interaction data. PANDA is used to estimate networks for each experimental group and the network models are then compared between groups to explore transcriptional processes that distinguish the groups. We present pandaR (bioconductor.org/packages/pandaR), a Bioconductor package that implements PANDA and provides a framework for exploratory data analysis on gene regulatory networks. Availability and Implementation: PandaR is provided as a Bioconductor R Package and is available at bioconductor.org/packages/pandaR. Daniel Schlauch, Joseph N. Paulson, Albert Young, Kimberly Glass, John Quackenbush |
Bioinform. | 5 |
| 2017 | Biomarker correlation network in colorectal carcinoma by tumor anatomic locationabstractBACKGROUND: Colorectal carcinoma evolves through a multitude of molecular events including somatic mutations, epigenetic alterations, and aberrant protein expression, influenced by host immune reactions. One way to interrogate the complex carcinogenic process and interactions between aberrant events is to model a biomarker correlation network. Such a network analysis integrates multidimensional tumor biomarker data to identify key molecular events and pathways that are central to an underlying biological process. Due to embryological, physiological, and microbial differences, proximal and distal colorectal cancers have distinct sets of molecular pathological signatures. Given these differences, we hypothesized that a biomarker correlation network might vary by tumor location. RESULTS: We performed network analyses of 54 biomarkers, including major mutational events, microsatellite instability (MSI), epigenetic features, protein expression status, and immune reactions using data from 1380 colorectal cancer cases: 690 cases with proximal colon cancer and 690 cases with distal colorectal cancer matched by age and sex. Edges were defined by statistically significant correlations between biomarkers using Spearman correlation analyses. We found that the proximal colon cancer network formed a denser network (total number of edges, n = 173) than the distal colorectal cancer network (n = 95) (P < 0.0001 in permutation tests). The value of the average clustering coefficient was 0.50 in the proximal colon cancer network and 0.30 in the distal colorectal cancer network, indicating the greater clustering tendency of the proximal colon cancer network. In particular, MSI was a key hub, highly connected with other biomarkers in proximal colon cancer, but not in distal colorectal cancer. Among patients with non-MSI-high cancer, BRAF mutation status emerged as a distinct marker with higher connectivity in the network of proximal colon cancer, but not in distal colorectal cancer. CONCLUSION: In proximal colon cancer, tumor biomarkers tended to be correlated with each other, and MSI and BRAF mutation functioned as key molecular characteristics during the carcinogenesis. Our findings highlight the importance of considering multiple correlated pathways for therapeutic targets especially in proximal colon cancer. Reiko Nishihara, Kimberly Glass, Kosuke Mima, Tsuyoshi Hamada, Jonathan A. Nowak, Zhi Rong Qian, Peter Kraft, Edward L. Giovannucci, Charles S. Fuchs, Andrew T. Chan, John Quackenbush, Shuji Ogino, Jukka-Pekka Onnela |
BMC Bioinform. | 11 |
| 2017 | Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse dataabstractBACKGROUND: Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. RESULTS: We find that analysis of large RNA-Seq data sets requires both careful quality control and the need to account for sparsity due to the heterogeneity intrinsic in multi-group studies. We developed Yet Another RNA Normalization software pipeline (YARN), that includes quality control and preprocessing, gene filtering, and normalization steps designed to facilitate downstream analysis of large, heterogeneous RNA-Seq data sets and we demonstrate its use with data from the Genotype-Tissue Expression (GTEx) project. CONCLUSIONS: An R package instantiating YARN is available at http://bioconductor.org/packages/yarn . Joseph N. Paulson, Cho-Yi Chen, Camila Miranda Lopes-Ramos, Marieke L. Kuijjer, John Platig, Abhijeet R. Sonawane, Maud Fagny, Kimberly Glass, John Quackenbush |
BMC Bioinform. | 9 |
| 2016 | PyPanda: a Python package for gene regulatory network reconstructionabstractPANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that uses message-passing to integrate multiple sources of 'omics data. PANDA was originally coded in C ++. In this application note we describe PyPanda, the Python version of PANDA. PyPanda runs considerably faster than the C ++ version and includes additional features for network analysis. AVAILABILITY AND IMPLEMENTATION: The open source PyPanda Python package is freely available at http://github.com/davidvi/pypanda CONTACT: [email protected] or [email protected]. David G. P. van IJzendoorn, Kimberly Glass, John Quackenbush, Marieke L. Kuijjer |
Bioinform. | 3 |
| 2016 | BatchQC: interactive software for evaluating sample and batch effects in genomic dataabstractSequencing and microarray samples often are collected or processed in multiple batches or at different times. This often produces technical biases that can lead to incorrect results in the downstream analysis. There are several existing batch adjustment tools for '-omics' data, but they do not indicate a priori whether adjustment needs to be conducted or how correction should be applied. We present a software pipeline, BatchQC, which addresses these issues using interactive visualizations and statistics that evaluate the impact of batch effects in a genomic dataset. BatchQC can also apply existing adjustment tools and allow users to evaluate their benefits interactively. We used the BatchQC pipeline on both simulated and real data to demonstrate the effectiveness of this software toolkit. AVAILABILITY AND IMPLEMENTATION: BatchQC is available through Bioconductor: http://bioconductor.org/packages/BatchQC and GitHub: https://github.com/mani2012/BatchQC CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Solaiappan Manimaran, Heather Marie Selby, Kwame Okrah, Claire Ruberman, Jeffrey T. Leek, John Quackenbush, Benjamin Haibe-Kains, Héctor Corrada Bravo, W. Evan Johnson |
Bioinform. | 6 |
| 2016 | Bipartite Community Structure of eQTLsabstractGenome Wide Association Studies (GWAS) and expression quantitative trait locus (eQTL) analyses have identified genetic associations with a wide range of human phenotypes. However, many of these variants have weak effects and understanding their combined effect remains a challenge. One hypothesis is that multiple SNPs interact in complex networks to influence functional processes that ultimately lead to complex phenotypes, including disease states. Here we present CONDOR, a method that represents both cis- and trans-acting SNPs and the genes with which they are associated as a bipartite graph and then uses the modular structure of that graph to place SNPs into a functional context. In applying CONDOR to eQTLs in chronic obstructive pulmonary disease (COPD), we found the global network "hub" SNPs were devoid of disease associations through GWAS. However, the network was organized into 52 communities of SNPs and genes, many of which were enriched for genes in specific functional classes. We identified local hubs within each community ("core SNPs") and these were enriched for GWAS SNPs for COPD and many other diseases. These results speak to our intuition: rather than single SNPs influencing single genes, we see groups of SNPs associated with the expression of families of functionally related genes and that disease SNPs are associated with the perturbation of those functions. These methods are not limited in their application to COPD and can be used in the analysis of a wide variety of disease processes and other phenotypic traits. John Platig, Peter J. Castaldi, Dawn L. DeMeo, John Quackenbush |
PLoS Comput. Biol. | 4 |
| 2015 | A network model for angiogenesis in ovarian cancerabstractBACKGROUND: We recently identified two robust ovarian cancer subtypes, defined by the expression of genes involved in angiogenesis, with significant differences in clinical outcome. To identify potential regulatory mechanisms that distinguish the subtypes we applied PANDA, a method that uses an integrative approach to model information flow in gene regulatory networks. RESULTS: We find distinct differences between networks that are active in the angiogenic and non-angiogenic subtypes, largely defined by a set of key transcription factors that, although previously reported to play a role in angiogenesis, are not strongly differentially-expressed between the subtypes. Our network analysis indicates that these factors are involved in the activation (or repression) of different genes in the two subtypes, resulting in differential expression of their network targets. Mechanisms mediating differences between subtypes include a previously unrecognized pro-angiogenic role for increased genome-wide DNA methylation and complex patterns of combinatorial regulation. CONCLUSIONS: The models we develop require a shift in our interpretation of the driving factors in biological networks away from the genes themselves and toward their interactions. The observed regulatory changes between subtypes suggest therapeutic interventions that may help in the treatment of ovarian cancer. Kimberly Glass, John Quackenbush, Dimitrios Spentzos, Benjamin Haibe-Kains, Guo-Cheng Yuan |
BMC Bioinform. | 2 |
| 2013 | RamiGO: an R/Bioconductor package providing an AmiGO Visualize interfaceabstractSUMMARY: The R/Bioconductor package RamiGO is an R interface to AmiGO that enables visualization of Gene Ontology (GO) trees. Given a list of GO terms, RamiGO uses the AmiGO visualize API to import Graphviz-DOT format files into R, and export these either as images (SVG, PNG) or into Cytoscape for extended network analyses. RamiGO provides easy customization of annotation, highlighting of specific GO terms, colouring of terms by P-value or export of a simplified summary GO tree. We illustrate RamiGO functionalities in a genome-wide gene set analysis of prognostic genes in breast cancer. AVAILABILITY AND IMPLEMENTATION: RamiGO is provided in R/Bioconductor, is open source under the Artistic-2.0 License and is available with a user manual containing installation, operating instructions and tutorials. It requires R version 2.15.0 or higher. URL: http://bioconductor.org/packages/release/bioc/html/RamiGO.html Markus S. Schröder, Daniel Gusenleitner, John Quackenbush, Aedín C. Culhane, Benjamin Haibe-Kains |
Bioinform. | 3 |
| 2013 | Research and applications: Comparison and validation of genomic predictors for anticancer drug sensitivityabstractBACKGROUND: An enduring challenge in personalized medicine lies in selecting the right drug for each individual patient. While testing of drugs on patients in large trials is the only way to assess their clinical efficacy and toxicity, we dramatically lack resources to test the hundreds of drugs currently under development. Therefore the use of preclinical model systems has been intensively investigated as this approach enables response to hundreds of drugs to be tested in multiple cell lines in parallel. METHODS: Two large-scale pharmacogenomic studies recently screened multiple anticancer drugs on over 1000 cell lines. We propose to combine these datasets to build and robustly validate genomic predictors of drug response. We compared five different approaches for building predictors of increasing complexity. We assessed their performance in cross-validation and in two large validation sets, one containing the same cell lines present in the training set and another dataset composed of cell lines that have never been used during the training phase. RESULTS: Sixteen drugs were found in common between the datasets. We were able to validate multivariate predictors for three out of the 16 tested drugs, namely irinotecan, PD-0325901, and PLX4720. Moreover, we observed that response to 17-AAG, an inhibitor of Hsp90, could be efficiently predicted by the expression level of a single gene, NQO1. CONCLUSION: These results suggest that genomic predictors could be robustly validated for specific drugs. If successfully validated in patients' tumor cells, and subsequently in clinical trials, they could act as companion tests for the corresponding drugs and play an important role in personalized medicine. Simon Papillon-Cavanagh, Nicolas De Jay, Nehme Hachem, Catharina Olsen, Gianluca Bontempi, Hugo J. W. L. Aerts, John Quackenbush, Benjamin Haibe-Kains |
J. Am. Medical Informatics Assoc. | 7 |
| 2013 | Significance Analysis of Prognostic SignaturesabstractA major goal in translational cancer research is to identify biological signatures driving cancer progression and metastasis. A common technique applied in genomics research is to cluster patients using gene expression data from a candidate prognostic gene set, and if the resulting clusters show statistically significant outcome stratification, to associate the gene set with prognosis, suggesting its biological and clinical importance. Recent work has questioned the validity of this approach by showing in several breast cancer data sets that "random" gene sets tend to cluster patients into prognostically variable subgroups. This work suggests that new rigorous statistical methods are needed to identify biologically informative prognostic gene sets. To address this problem, we developed Significance Analysis of Prognostic Signatures (SAPS) which integrates standard prognostic tests with a new prognostic significance test based on stratifying patients into prognostic subtypes with random gene sets. SAPS ensures that a significant gene set is not only able to stratify patients into prognostically variable groups, but is also enriched for genes showing strong univariate associations with patient prognosis, and performs significantly better than random gene sets. We use SAPS to perform a large meta-analysis (the largest completed to date) of prognostic pathways in breast and ovarian cancer and their molecular subtypes. Our analyses show that only a small subset of the gene sets found statistically significant using standard measures achieve significance by SAPS. We identify new prognostic signatures in breast and ovarian cancer and their corresponding molecular subtypes, and we show that prognostic signatures in ER negative breast cancer are more similar to prognostic signatures in ovarian cancer than to prognostic signatures in ER positive breast cancer. SAPS is a powerful new method for deriving robust prognostic biological signatures from clinically annotated genomic datasets. Andrew H. Beck, Nicholas W. Knoblauch, Marco M. Hefti, Jennifer Kaplan, Stuart J. Schnitt, Aedín C. Culhane, Markus S. Schröder, Thomas Risch, John Quackenbush, Benjamin Haibe-Kains |
PLoS Comput. Biol. | 9 |
| 2012 | nEASE: a method for gene ontology subclassification of high-throughput gene expression dataabstractUNLABELLED: High-throughput technologies can identify genes whose expression profiles correlate with specific phenotypes; however, placing these genes into a biological context remains challenging. To help address this issue, we developed nested Expression Analysis Systematic Explorer (nEASE). nEASE complements traditional gene ontology enrichment approaches by determining statistically enriched gene ontology subterms within a list of genes based on co-annotation. Here, we overview an open-source software version of the nEASE algorithm. nEASE can be used either stand-alone or as part of a pathway discovery pipeline. AVAILABILITY: nEASE is implemented within the Multiple Experiment Viewer software package available at http://www.tm4.org/mev. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas W. Chittenden, Eleanor Howe, Jennifer M. Taylor, Jessica Cara Mar, Martin J. Aryee, Harold Gómez, Razvan Sultana, John C. Braisted, Sarita J. Nair, John Quackenbush |
Bioinform. | 10 |
| 2012 | iBBiG: iterative binary bi-clustering of gene setsabstractMOTIVATION: Meta-analysis of genomics data seeks to identify genes associated with a biological phenotype across multiple datasets; however, merging data from different platforms by their features (genes) is challenging. Meta-analysis using functionally or biologically characterized gene sets simplifies data integration is biologically intuitive and is seen as having great potential, but is an emerging field with few established statistical methods. RESULTS: We transform gene expression profiles into binary gene set profiles by discretizing results of gene set enrichment analyses and apply a new iterative bi-clustering algorithm (iBBiG) to identify groups of gene sets that are coordinately associated with groups of phenotypes across multiple studies. iBBiG is optimized for meta-analysis of large numbers of diverse genomics data that may have unmatched samples. It does not require prior knowledge of the number or size of clusters. When applied to simulated data, it outperforms commonly used clustering methods, discovers overlapping clusters of diverse sizes and is robust in the presence of noise. We apply it to meta-analysis of breast cancer studies, where iBBiG extracted novel gene set-phenotype association that predicted tumor metastases within tumor subtypes. AVAILABILITY: Implemented in the Bioconductor package iBBiG CONTACT: [email protected]. Daniel Gusenleitner, Eleanor Howe, Stefan Bentink, John Quackenbush, Aedín C. Culhane |
Bioinform. | 4 |
| 2012 | Viral Perturbations of Host Networks Reflect Disease EtiologyabstractMany human diseases, arising from mutations of disease susceptibility genes (genetic diseases), are also associated with viral infections (virally implicated diseases), either in a directly causal manner or by indirect associations. Here we examine whether viral perturbations of host interactome may underlie such virally implicated disease relationships. Using as models two different human viruses, Epstein-Barr virus (EBV) and human papillomavirus (HPV), we find that host targets of viral proteins reside in network proximity to products of disease susceptibility genes. Expression changes in virally implicated disease tissues and comorbidity patterns cluster significantly in the network vicinity of viral targets. The topological proximity found between cellular targets of viral proteins and disease genes was exploited to uncover a novel pathway linking HPV to Fanconi anemia. Natali Gulbahce, Amélie Dricot, Megha Padi, Danielle Byrdsong, Rachel Franchi, Deok-Sun Lee, Orit Rozenblatt-Rosen, Jessica Cara Mar, Michael A. Calderwood, Amy Baldwin, Balaji Santhanam, Pascal Falter-Braun, Nicolas Simonis, Kyung-Won Huh, Karin Hellner, Miranda Grace, Alyce Chen, Renee Rubio, Jarrod A. Marto, Nicholas A. Christakis, Elliott Kieff, Frederick P. Roth, Jennifer Roecklein-Canfield, James A. DeCaprio, Michael E. Cusick, John Quackenbush, David E. Hill, Karl Münger, Marc Vidal, Albert-László Barabási |
PLoS Comput. Biol. | 28 |
| 2011 | RNA-Seq analysis in MeVabstractAbstract Summary: RNA-Seq is an exciting methodology that leverages the power of high-throughput sequencing to measure RNA transcript counts at an unprecedented accuracy. However, the data generated from this process are extremely large and biologist-friendly tools with which to analyze it are sorely lacking. MultiExperiment Viewer (MeV) is a Java-based desktop application that allows advanced analysis of gene expression data through an intuitive graphical user interface. Here, we report a significant enhancement to MeV that allows analysis of RNA-Seq data with these familiar, powerful tools. We also report the addition to MeV of several RNA-Seq-specific functions, addressing the differences in analysis requirements between this data type and traditional gene expression data. These tools include automatic conversion functions from raw count data to processed RPKM or FPKM values and differential expression detection and functional annotation enrichment detection based on published methods. Availability: MeV version 4.7 is written in Java and is freely available for download under the terms of the open-source Artistic License version 2.0. The website (http://mev.tm4.org/) hosts a full user manual as well as a short quick-start guide suitable for new users. Contact: [email protected] Eleanor Howe, Raktim Sinha, Daniel Schlauch, John Quackenbush |
Bioinform. | 4 |
| 2011 | Defining an informativeness metric for clustering gene expression dataabstractMOTIVATION: Unsupervised 'cluster' analysis is an invaluable tool for exploratory microarray data analysis, as it organizes the data into groups of genes or samples in which the elements share common patterns. Once the data are clustered, finding the optimal number of informative subgroups within a dataset is a problem that, while important for understanding the underlying phenotypes, is one for which there is no robust, widely accepted solution. RESULTS: To address this problem we developed an 'informativeness metric' based on a simple analysis of variance statistic that identifies the number of clusters which best separate phenotypic groups. The performance of the informativeness metric has been tested on both experimental and simulated datasets, and we contrast these results with those obtained using alternative methods such as the gap statistic. AVAILABILITY: The method has been implemented in the Bioconductor R package attract; it is also freely available from http://compbio.dfci.harvard.edu/pubs/attract_1.0.1.zip. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jessica Cara Mar, Christine A. Wells, John Quackenbush |
Bioinform. | 3 |
| 2011 | Exome sequencing-based copy-number variation and loss of heterozygosity detection: ExomeCNVabstractMOTIVATION: The ability to detect copy-number variation (CNV) and loss of heterozygosity (LOH) from exome sequencing data extends the utility of this powerful approach that has mainly been used for point or small insertion/deletion detection. RESULTS: We present ExomeCNV, a statistical method to detect CNV and LOH using depth-of-coverage and B-allele frequencies, from mapped short sequence reads, and we assess both the method's power and the effects of confounding variables. We apply our method to a cancer exome resequencing dataset. As expected, accuracy and resolution are dependent on depth-of-coverage and capture probe design. AVAILABILITY: CRAN package 'ExomeCNV'. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jarupon Fah Sathirapongsasuti, Hane Lee, Basil A. J. Horst, Georg Brunner, Alistair J. Cochran, Scott Binder, John Quackenbush, Stanley F. Nelson |
Bioinform. | 7 |
| 2011 | survcomp: an R/Bioconductor package for performance assessment and comparison of survival modelsabstractAbstract Summary: The survcomp package provides functions to assess and statistically compare the performance of survival/risk prediction models. It implements state-of-the-art statistics to (i) measure the performance of risk prediction models; (ii) combine these statistical estimates from multiple datasets using a meta-analytical framework; and (iii) statistically compare the performance of competitive models. Availability: The R/Bioconductor package survcomp is provided open source under the Artistic-2.0 License with a user manual containing installation, operating instructions and use case scenarios on real datasets. survcomp requires R version 2.13.0 or higher. http://bioconductor.org/packages/release/bioc/html/survcomp.html Contact: [email protected]; [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Markus S. Schröder, Aedín C. Culhane, John Quackenbush, Benjamin Haibe-Kains |
Bioinform. | 3 |
| 2011 | Multiple-input multiple-output causal strategies for gene selectionabstractBACKGROUND: Traditional strategies for selecting variables in high dimensional classification problems aim to find sets of maximally relevant variables able to explain the target variations. If these techniques may be effective in generalization accuracy they often do not reveal direct causes. The latter is essentially related to the fact that high correlation (or relevance) does not imply causation. In this study, we show how to efficiently incorporate causal information into gene selection by moving from a single-input single-output to a multiple-input multiple-output setting. RESULTS: We show in synthetic case study that a better prioritization of causal variables can be obtained by considering a relevance score which incorporates a causal term. In addition we show, in a meta-analysis study of six publicly available breast cancer microarray datasets, that the improvement occurs also in terms of accuracy. The biological interpretation of the results confirms the potential of a causal approach to gene selection. CONCLUSIONS: Integrating causal information into gene selection algorithms is effective both in terms of prediction accuracy and biological interpretation. Gianluca Bontempi, Benjamin Haibe-Kains, Christine Desmedt, Christos P. Sotiriou, John Quackenbush |
BMC Bioinform. | 5 |
| 2011 | GCOD - GeneChip Oncology DatabaseabstractBACKGROUND: DNA microarrays have become a nearly ubiquitous tool for the study of human disease, and nowhere is this more true than in cancer. With hundreds of studies and thousands of expression profiles representing the majority of human cancers completed and in public databases, the challenge has been effectively accessing and using this wealth of data. DESCRIPTION: To address this issue we have collected published human cancer gene expression datasets generated on the Affymetrix GeneChip platform, and carefully annotated those studies with a focus on providing accurate sample annotation. To facilitate comparison between datasets, we implemented a consistent data normalization and transformation protocol and then applied stringent quality control procedures to flag low-quality assays. CONCLUSION: The resulting resource, the GeneChip Oncology Database, is available through a publicly accessible website that provides several query options and analytical tools through an intuitive interface. Fenglong Liu, Joseph White, Corina Antonescu, John Quackenbush |
BMC Bioinform. | 4 |
| 2009 | EditorialabstractThe transforming aspect of the Human Genome Project was not the completion of the genome sequence itself, but rather the technologies that enabled, and were enabled by, the sequencing of that first reference genome. The evolution of ‘omic science through microarray transcriptomics, metabolomics, proteomics, and whole-genome SNP-omics has in many ways come full circle with a new focus on genomics and genome sequencing. Next-generation sequencing technologies have begun to revolutionise genomics and their effects are becoming increasingly widespread. The 1000 genomes project (http://www.1000genomes.org/) will create a new map of genetic variation for our genome going far beyond the detail captured in the HapMap. Other projects are helping to catalogue genes involved in cancer, alternative splicing in different tissues and transcription factor binding, for example. The growing number of robust applications and the steadily falling cost for generating sequence-based data suggest that these next-generation technologies will continue to rapidly open new applications in the biological sciences and generate new opportunities for software and algorithm development. Given the vast amount of data produced (currently greater than a gigabase per run, with this constantly increasing as well), developing a sound data storage and management solution and creating informatics tools to effectively analyze the data are essential to successful application of the technology. During the past year, a large number of new software applications and algorithms have been developed to deal with this new data. A recent advert in Nature from the Illumina, one of the providers of next-generation sequencing technology, highlighted significant papers in the area of bioinformatics; our journal published 7 of the 16 listed papers. In addition to those cited in the advertisement, there have been many other tools and algorithms published in Bioinformatics that are relevant to next-generation sequencing applications. To celebrate this contribution we have gathered these together in a ‘Bioinformatics for Next Generation Sequencing’ virtual issue (http://www.oxfordjournals.org/our_journals/bioinformatics/nextgenerationsequencing.html). This will be a living resource that we will continually update to include the very latest papers in this area to help researchers keep abreast of the latest developments. To date, the majority of the papers have described methods to take the short sequences produced by the Illumina Genome Analyzer and Applied Biosystems SOLiD machines and align them to a reference genome. This is a crucial and basic requirement for many applications and a variety of techniques have been applied to make the tools sufficiently fast to deal with millions of sequences. We have also included papers that address the issue of assembly of these short reads. Now that there are many of these tools available the Bioinformatics community has begun to make applications that are useful for specific applications such as identifying likely sites of interaction in CHIP-seq. A summary of the inaugural collection is included in the Table 1. We sincerely hope that you find this resource useful and that the collected references lead to additional development in an area that we view as critical to the continued development of genomics and bioinformatics. Tools recently published in the journal Tools recently published in the journal Alex Bateman, John Quackenbush |
Bioinform. | 2 |
| 2009 | Papers on normalization, variable selection, classification or clustering of microarray dataabstractOver the last decade or so, there have been large numbers of methods published on approaches for normalization, variable (gene) selection, classification and clustering of microarray data. As indicated in the scope document for Bioinformatics, this requires papers describing new methods for these problems to meet a very high standard, showing important improvement in results for real biological data, as well as novelty. In this editorial, we describe some standards that need to be met for papers in these areas to be seriously considered. We ask that prospective authors consider these points carefully before submission of their papers to Bioinformatics. The role of simulation: Simulation can be useful in investigating the properties of various methods of data analysis. Yet, there are important barriers to credible use of simulation in microarray studies, largely due to what we do not know about the statistical distribution of measured gene expression levels. First, the distribution across transcripts of true expression values is dependent on the biological state of the tissue or cell, and for a given state this is unknown, even in distributional form, and may further exhibit gene- and platform-specific effects. Second, the correlation within biological replicates of true expression is unknown, and is likely unknowable in detail given that it is expressed by a correlation matrix with on the order of a billion entries. Third, the distribution of changes from one biological state to another is unknown. Fourth, the correlation in observational errors in gene expression across genes is unknown and similarly probably unknowable in detail. On the other hand, the measurement error for a given transcript has been well described by several authors (Ideker et al., 2001; Rocke and Durbin, 2001). Given this gap between knowledge and simulation specification, it is likely that any new method can be shown to be superior to some other method(s) by careful choice of simulation parameters, since simulations often include biases in the distributions selected and in other assumptions of the models. Thus, while simulation may still be worthwhile, and a useful tool for exploring robustness and parameter space of a new method, it is insufficient evidence for superiority of a new method without substantial support from significant improvement in results from analysis of real data. Normalization: Normalization necessarily involves a trade-off between its positive role in reducing variability, and its potentially negative role in increasing bias. There are a number of good image analysis, preprocessing, transformation and normalization methods extant for single- and dual-color DNA microarrays. To show that a new method is better requires comparison demonstrating that results in differential expression analysis, classification or clustering are better with the new normalization method than with previous methods. Not one but several previous methods should be chosen for comparison including the most widely used approaches. Several datasets should be used, including spike-in and dilution studies when feasible, as well as ‘real’ biological datasets. Showing that more genes are differentially expressed using a normalization method is not compelling evidence of superiority without a good estimate of the false-positive rate or a compelling biological analysis of the resulting differentially expressed genes. Variable selection: Typically, new variable selection methods are proposed as part of a classification or clustering strategy, and demonstrating superiority of the variable selection method usually means demonstrating superiority of the combined methodology. It is quite important that metrics for evaluation be used that are robust to intra-array correlations and variable selection artifacts. For example, in cross-validation studies in which variable selection is followed by a classification method, selection of variables using all the data and then cross-validating the classification accuracy introduces substantial bias, making classification methods appear more accurate than they really are (Ambroise and McLachlan, 2002). It is important that any method be compared with several of the most widely used existing methods, including baseline approaches such as filtering by t-score or forward stepwise analysis. Such comparisons should be performed on more than one biological dataset. Further, the method must demonstrate significant improvement over existing methods; incremental improvements will not be considered of sufficient interest to warrant review. Classification and prediction: New classification or prediction methods for microarray data enter a crowded arena. From long-standing techniques such as logistic regression and linear discriminant analysis to the more modern support vector machines and neural networks, most known classification methods have already been applied to microarray data. To show that a newly proposed classification method is a real advance, a substantial improvement in performance needs to be shown over a reasonable selection of existing datasets and methods, including commonly used or simple methods. This is because, consciously or subconsciously, the developer of a new method optimizes its characteristics against the datasets to be used for evaluation. Variable selection and parameter choice for all methods needs to be done strictly in the training set (whether there is one training set or many as in cross validation). Resampling methods like permuting the class labels on the arrays or the bootstrap can be used to provide robust estimates of the significance of differential expression, but do not in themselves give estimates of classification performance except to show that the performance is better than chance. Experience shows that there is considerable noise in classification accuracy experiments, so modest increases in achieved accuracy are usually not convincing. Experience also shows that classification performance in a microarray problem depends strongly on the dataset, and less on the variable selection and classification methods. More than modest differences are required to excite interest in a new method. Authors should keep in mind the ‘No Free Lunch Theorems’ of Wolpert and Macready (1997) which demonstrated that there is no optimization/classification method that outperforms all others in all circumstances (Wolpert, 1996). Clustering: Demonstrating superiority of a clustering method is in many ways more difficult than demonstrating superiority in a classification method. Usually, there is no ground truth against which to compare the clustering results. Defining a criterion (e.g. the Rand index) and showing that a clustering method achieves better scores on this criterion is often not compelling, since such criteria are easily optimized (again, consciously or subconsciously) to ensure superiority. For reasons discussed above, simulation is also not usually sufficient. Ideally, a new clustering method would demonstrate novel biological insights or some attractive statistical properties not available from previous methods, including several commonly used methods. Requiring new biological findings is a difficult standard, but a necessary one to insure that new published methods are useful and likely to be used. To conclude, microarrays remain a useful technology to address a wide array of biological problems and the optimal analysis of these data to extract meaningful results still pose many bioinformatics challenges. However, with a number of successful methods already addressing the well-established microarray data analysis problems, publication of new methods in this area requires either identification of a new challenge and formulation of a new problem or development of a substantially better methodology then those existing that can be benchmarked on a variety of datasets. We hope that suggestions provided above for evaluation and validation of such new methods would increase the likelihood of them supporting biological discoveries in the future. Funding: DMR to NIH grants P42-ES04699 and R01-HG003352. David M. Rocke, Trey Ideker, Olga G. Troyanskaya, John Quackenbush, Joaquín Dopazo |
Bioinform. | 4 |
| 2009 | An improved empirical bayes approach to estimating differential gene expression in microarray time-course data: BETR (Bayesian Estimation of Temporal Regulation)abstractBACKGROUND: Microarray gene expression time-course experiments provide the opportunity to observe the evolution of transcriptional programs that cells use to respond to internal and external stimuli. Most commonly used methods for identifying differentially expressed genes treat each time point as independent and ignore important correlations, including those within samples and between sampling times. Therefore they do not make full use of the information intrinsic to the data, leading to a loss of power. RESULTS: We present a flexible random-effects model that takes such correlations into account, improving our ability to detect genes that have sustained differential expression over more than one time point. By modeling the joint distribution of the samples that have been profiled across all time points, we gain sensitivity compared to a marginal analysis that examines each time point in isolation. We assign each gene a probability of differential expression using an empirical Bayes approach that reduces the effective number of parameters to be estimated. CONCLUSIONS: Based on results from theory, simulated data, and application to the genomic data presented here, we show that BETR has increased power to detect subtle differential expression in time-series data. The open-source R package betr is available through Bioconductor. BETR has also been incorporated in the freely-available, open-source MeV software tool available from http://www.tm4.org/mev.html. Martin J. Aryee, José A. Gutiérrez-Pabello, Igor Kramnik, Tapabrata Maiti, John Quackenbush |
BMC Bioinform. | 5 |
| 2009 | Data-driven normalization strategies for high-throughput quantitative RT-PCRabstractBACKGROUND: High-throughput real-time quantitative reverse transcriptase polymerase chain reaction (qPCR) is a widely used technique in experiments where expression patterns of genes are to be profiled. Current stage technology allows the acquisition of profiles for a moderate number of genes (50 to a few thousand), and this number continues to grow. The use of appropriate normalization algorithms for qPCR-based data is therefore a highly important aspect of the data preprocessing pipeline. RESULTS: We present and evaluate two data-driven normalization methods that directly correct for technical variation and represent robust alternatives to standard housekeeping gene-based approaches. We evaluated the performance of these methods against a single gene housekeeping gene method and our results suggest that quantile normalization performs best. These methods are implemented in freely-available software as an R package qpcrNorm distributed through the Bioconductor project. CONCLUSION: The utility of the approaches that we describe can be demonstrated most clearly in situations where standard housekeeping genes are regulated by some experimental condition. For large qPCR-based data sets, our approaches represent robust, data-driven strategies for normalization. Jessica Cara Mar, Yasumasa Kimura, Kate Schroder, Katharine M. Irvine, Yoshihide Hayashizaki, Harukazu Suzuki, David A. Hume, John Quackenbush |
BMC Bioinform. | 8 |
| 2009 | Decomposition of Gene Expression State Space TrajectoriesabstractRepresenting and analyzing complex networks remains a roadblock to creating dynamic network models of biological processes and pathways. The study of cell fate transitions can reveal much about the transcriptional regulatory programs that underlie these phenotypic changes and give rise to the coordinated patterns in expression changes that we observe. The application of gene expression state space trajectories to capture cell fate transitions at the genome-wide level is one approach currently used in the literature. In this paper, we analyze the gene expression dataset of Huang et al. (2005) which follows the differentiation of promyelocytes into neutrophil-like cells in the presence of inducers dimethyl sulfoxide and all-trans retinoic acid. Huang et al. (2005) build on the work of Kauffman (2004) who raised the attractor hypothesis, stating that cells exist in an expression landscape and their expression trajectories converge towards attractive sites in this landscape. We propose an alternative interpretation that explains this convergent behavior by recognizing that there are two types of processes participating in these cell fate transitions-core processes that include the specific differentiation pathways of promyelocytes to neutrophils, and transient processes that capture those pathways and responses specific to the inducer. Using functional enrichment analyses, specific biological examples and an analysis of the trajectories and their core and transient components we provide a validation of our hypothesis using the Huang et al. (2005) dataset. Jessica Cara Mar, John Quackenbush |
PLoS Comput. Biol. | 2 |
| 2007 | Stochasticity and Networks in Genomic DataabstractTwo trends are driving innovation and discovery in biological sciences: technologies that allow holistic surveys of genes, proteins, and metabolites and a realization that biological processes are driven by complex networks of interacting biological molecules. However, there is a gap between the gene lists emerging from genome sequencing projects and the network diagrams that are essential if we are to understand the link between genotype and phenotype. 'Omic technologies such as DNA microarrays were once heralded as providing a window into those networks, but so far their success has been limited. Although many techniques have been developed to deal with microarray data, to date their ability to extract network relationships has been limited. We believed that by imposing constraints on the networks, based on associations reported through articles indexed in PubMed, we could more effectively extract biologically relevant results from microarray data and develop testable hypotheses that could then be validated in the laboratory. Using literature networks as constraints on a Bayesian network analysis of microarray data, we show that we are able to recover evidence for a wide range of known networks and pathways, even in experiments not explicitly designed to probe them. With a putative gene-interaction network, the problem of producing viable models of the cell remains. While systems biology approaches that attempt to develop quantitative, predictive models of cellular processes have received great attention, it is surprising to note that the starting point for all cellular gene expression, the transcription of RNA, has not been described and measured in a population of living cells. To address this problem, we propose a simple (and obvious) model for transcript levels based on Poisson statistics and provide supporting experimental evidence for genes known to be expressed at high, moderate, and low levels. Although what we describe as a microscopic process, occurring at the level of an individual cell, the data we provide uses a small number of cells where the echoes of the underlying stochastic processes can be seen. Not only do these data confirm our model, but this general strategy opens up a potential new approach, Mesoscopic Biology, that can be used to assess the natural variability of processes occurring at the cellular level in biological systems. Together these two approaches open new avenues of investigation that may help us in our eventual understanding of the function of biological systems, addressing many of the important questions that have arisen in the context of systems biology. Our ultimate goal will be to create predictive models that allow one to examine the current state of a biological system and to estimate the likelihood that, at some later time, the system will have evolved to a new state. Such an approach, if successful, could have a wide range of applications spanning laboratory, clinical, and translational biology. John Quackenbush |
BIBE | 1 |
| 2007 | Microarray blob-defect removal improves array analysisabstractMOTIVATION: New generation Affymetrix oligonucleotide microarrays often have blob-like image defects that will require investigators to either repeat their hybridization assays or analyze their data with the defects left in place. We investigated the effect of analyzing a spike-in experiment on Affymetrix ENCODE tiling arrays in the presence of simulated blobs covering between 1 and 9% of the array area. Using two different ChIP-chip tiling array analysis programs (Affymetrix tiling array software, TAS, and model-based analysis of tiling arrays, MAT), we found that even the smallest blob defects significantly decreased the sensitivity and increased the false discovery rate (FDR) of the spike-in target prediction. RESULTS: We introduced a new software tool, the microarray blob remover (MBR), which allows rapid visualization, detection and removal of various blob defects from the .CEL files of different types of Affymetrix microarrays. It is shown that using MBR significantly improves the sensitivity and FDR of a tiling array analysis compared to leaving the affected probes in the analysis. AVAILABILITY: The MBR software and the sample array .CEL files used in this article are available at: http://liulab.dfci.harvard.edu/Software/MBR/MBR.htm Jun S. Song, Kaveh Maghsoudi, Wei Li 0036, John Quackenbush, Xiaole Shirley Liu |
Bioinform. | 5 |
| 2006 | It is time to end the patenting of softwareabstractOne of the most significant outcomes of genomics has been a rapid increase in the rate that we as a community can generate data on interesting biological systems. Rapid improvements in technologies such as DNA microarrays and proteomics applications have produced a climate where the challenge is no longer collecting high quality data but rather managing and analyzing it. As we in the bioinformatics community have addressed this challenge, we have had to carefully consider the way in which the results of our intellectual efforts—the software tools that we develop—are made available to the wider research community. Increasingly, bioinformatics scientists are coming to call for development in an open source environment in which software is distributed with its underlying source code under a license that generally allows broad reuse and redistribution of the code under certain, usually minimal, restrictions. This model for software distribution has certain distinct advantages that add significant value to tools and techniques that we create. Developing software in a free and open community allows us to build rapidly on the advances of others, rather than having to re-invent or re-engineer software that came before and in doing so greatly accelerates the progress of research both in bioinformatics and in the fields that it touches. Further, making source code available provides an opportunity for methods to be checked and verified in ways that would not be possible if the software or the method it was based upon were proprietary. Admittedly, this model may not work best for all software, but in the research community, which is funded primarily by the public sector, we believe it makes sense for scientists to share their software just as they share their discoveries through publications. Of course, open source software can be copyrighted, which gives the developers of the software proper credit for their creative efforts, much like putting one's name on a research paper gives credit for the results described within it. However, there is another way to protect software: patents. Starting in 1981 (Author Webpage), the United States patent office (USPTO) decided that they would issue patents to software, even though copyright protection was already available. Patents offer much stronger protection than copyright, since they protect not just a software implementation, but also the algorithms and ideas underlying the software. If a program is patented, then others are not permitted to re-implement the algorithms on their own unless they are willing to pay a fee to the patent holder. In practice, software patents have been a boon for attorneys and the USPTO, which is currently granting around 300 000 new patents per year, but a disaster for scientists and engineers who want to build useful artifacts. They have spurred the development of a mini-industry of shell companies whose primary goal is to file patents on software and then to sue others in the hope of collecting monetary damages. These companies often do not even attempt to commercialize their software, but exist solely to extract money from the productive efforts of others. The availability of patents has also spurred large software companies to file literally hundreds of patent applications, many of which have been granted, on their software. Sometimes these patents are filed to defend their turf, warding off competitors, and other times they are filed defensively, to protect against the vulturous shell companies that might attempt to sue. Regardless of the reason, it seems clear that the efforts to file and defend patents primarily benefit lawyers, predatory special interests and others who are not developing software themselves. Although patents are intended to encourage innovation by disclosing the method used to achieve a particular goal while protecting the inventor's intellectual property, in the software world they act primarily to prevent it. A patent is supposed to instruct ‘someone skilled in the art’ how to ‘practice’ a particular invention while claiming rights to the method and use of that invention. In practice, software patents generally disclose very little of practical use and then issue broad, vague claims about the applicability of their method that often goes far beyond what one might reasonably infer they have described. As a result, software patents often create roadblocks to the development of new methods and drain away valuable resources from companies attempting to build new tools. Consider one example. There has been a lengthy patent dispute between Research in Motion (RIM), the makers of the popular Blackberry hand-held PDAs, and a small company called NTP, whose only noteworthy assets are a series of software patents (Austen and Guernsey, 2005). This dispute recently threatened to shut down Blackberry's enormously popular email service, despite the fact that NTP has no competing service to offer. The basic facts are not in dispute: several years ago, NTP was granted a patent on technology for wireless email services. Around the same time, RIM independently invented similar algorithms, and unlike NTP, RIM created a device and started selling it, eventually becoming a highly successful company. NTP sued, and is now seeking some $1 billion USD in damages. Judges initially ruled in NTP's favor, and RIM appealed to a higher court, while continuing to ask the US patent office to rule that the original patents were invalid. Finally, in March 2006, RIM settled the case (under tremendous pressure from the judge) and paid $612.5 million to NTP. Thus the patent ‘blackmail’ strategy worked: RIM invested enormous amounts of time and money, finally paying a huge sum to NTP, and Blackberry users are no better off. All of this to allow RIM to provide a service based on a technology that they clearly developed and implemented. Now consider a hypothetical example relevant to bioinformatics: imagine that the BLAST algorithm (Altschul et al, 1990) had been patented, and further that the patent holders insisted on collecting fees from everyone who wished to use it. The incredibly valuable BLAST servers at NCBI would never have even been built, nor would the hundreds of other sites providing BLAST search services for numerous genome databases. Countless discoveries built on the use of BLAST would have been missed or at best slowed down. BLAST servers all over the world would be shut down or forced to pay fees that would produce no benefit to the scientific community. This scenario is not far-fetched: patent applications have already been filed for numerous sequence alignment algorithms, though fortunately none of them pre-dates the free availability of BLAST. If they did, the patent claims themselves would likely be broad enough that they would prevent BLAST or any other sequence alignment method from being used without a license. We believe that the practice of issuing patents for software should end. While we realize that we may not be able to make patent offices stop issuing software patents, we can at least discourage members of our own research community from patenting software. One way we can take action is to reject any manuscripts submitted to scientific journals if they describe software that is protected by patents or is subject to a pending patent application. We urge all our scientific colleagues to denounce software patents, as we have, and to embrace the practice of open-source development. Otherwise you may find yourself one day asking a lawyer's permission to run software that you wrote yourself. John Quackenbush, Steven Salzberg |
Bioinform. | 1 |
| 2006 | A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TABabstractBACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML. Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma |
BMC Bioinform. | 12 |
| 2005 | MeSHer: identifying biological concepts in microarray assays based on PubMed references and MeSH termsabstractUNLABELLED: MeSHer uses a simple statistical approach to identify biological concepts in the form of Medical Subject Headings (MeSH terms) obtained from the PubMed database that are significantly overrepresented within the identified gene set relative to those associated with the overall collection of genes on the underlying DNA microarray platform. As a demonstration, we apply this approach to gene lists acquired from a published study of the effects of angiotensin II (Ang II) treatment on cardiac gene expression and demonstrate that this approach can aid in the interpretation of the resulting 'significant' gene set. AVAILABILITY: The software is available at http://www.tm4.org. SUPPLEMENTARY INFORMATION: Results from the analysis of significant genes from the published Ang II study. Amira Djebbari, Svetlana Karamycheva, Eleanor Howe, John Quackenbush |
Bioinform. | 4 |
| 2005 | CGHAnalyzer: a stand-alone software package for cancer genome analysis using array-based DNA copy number dataabstractSUMMARY: This synopsis provides an overview of array-based comparative genomic hybridization data display, abstraction and analysis using CGHAnalyzer, a software suite, designed specifically for this purpose. CGHAnalyzer can be used to simultaneously load copy number data from multiple platforms, query and describe large, heterogeneous datasets and export results. Additionally, CGHAnalyzer employs a host of algorithms for microarray analysis that include hierarchical clustering and class differentiation. AVAILABILITY: CGHAnalyzer, the accompanying manual, documentation and sample data are available for download at http://acgh.afcri.upenn.edu. This is a Java-based application built in the framework of the TIGR MeV that can run on Microsoft Windows, Macintosh OSX and a variety of Unix-based platforms. It requires the installation of the free Java Runtime Environment 1.4.1 (or more recent) (http://www.java.sun.com). Adam A. Margolin, Joel Greshock, Tara L. Naylor, Yael Mosse, John M. Maris, Graham Bignell, Alexander I. Saeed, John Quackenbush, Barbara L. Weber |
Bioinform. | 8 |
| 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasetsabstractAbstract TGICL is a pipeline for analysis of large Expressed Sequence Tags (EST) and mRNA databases in which the sequences are first clustered based on pairwise sequence similarity, and then assembled by individual clusters (optionally with quality values) to produce longer, more complete consensus sequences. The system can run on multi-CPU architectures including SMP and PVM. Availability: http://www.tigr.org/tdb/tgi/software/ Contact: [email protected]; [email protected] * To whom correspondence should be addressed. Geo Pertea, Xiaoqiu Huang 0001, Valentin Antonescu, Razvan Sultana, Svetlana Karamycheva, Yuandan Lee, Joseph White, Foo Cheung, Babak Parvizi, Jennifer Tsai, John Quackenbush |
Bioinform. | 12 |
| 2002 | An open letter to the scientific journalsabstractCatherine A. Ball, Gavin Sherlock, Helen Parkinson, Philippe Rocca-Sera, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hin Catherine A. Ball, Gavin Sherlock, Helen E. Parkinson, Philippe Rocca-Serra, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hingamp, Frank C. P. Holstege, Martin Ringwald, Paul T. Spellman, Christian J. Stoeckert Jr., Jason E. Stewart, Ronald C. Taylor, Alvis Brazma, John Quackenbush |
Bioinform. | 17 |
| 2002 | Genesis: cluster analysis of microarray dataabstractA versatile, platform independent and easy to use Java suite for large-scale gene expression analysis was developed. Genesis integrates various tools for microarray data analysis such as filters, normalization and visualization tools, distance measures as well as common clustering algorithms including hierarchical clustering, self-organizing maps, k-means, principal component analysis, and support vector machines. The results of the clustering are transparent across all implemented methods and enable the analysis of the outcome of different algorithms and parameters. Additionally, mapping of gene expression data onto chromosomal sequences was implemented to enhance promoter analysis and investigation of transcriptional control mechanisms. Alexander Sturn, John Quackenbush, Zlatko Trajanoski |
Bioinform. | 2 |