EDBT 2026 Demo / reviewers in the wild / expert
Andrea Califano
dblp:55/1524
· DBLP profile ↗
37ranked-venue papers
15as first author
1since 2021 · last 2026
0000-0003-4742-3679ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 15 · 11 first-authorGraphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
15 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
9 papers |
3D vision · 48% Image recognition and object detection · 39% Segmentation and scene understanding · 8% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.8 | 4 | 2016 | ARACNe-AP: gene network reverse engineering through adaptive partitioning inference of mutual information · Bioinform. 2016 The Cyni framework for network inference in Cytoscape · Bioinform. 2015 DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015 |
Bioinformatics and computational biology
gene expression analysis |
0.4 | 2 | 2018 | iterClust: a statistical framework for iterative clustering analysis · Bioinform. 2018 Analysis of Gene Expression Microarrays for Phenotype Classification · ISMB 2000 |
Bioinformatics and computational biology › gene regulation › gene regulatory network
gene regulatory network analysis |
0.3 | 1 | 2018 | ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's disease · Bioinform. 2018 |
Bioinformatics and computational biology
functional genomics |
0.2 | 1 | 2016 | ScreenBEAM: a novel meta-analysis algorithm for functional genomics screens via Bayesian hierarchical modeling · Bioinform. 2016 |
Bioinformatics and computational biology › drug discovery
high-throughput screening |
0.2 | 1 | 2016 | Detection and removal of spatial bias in multiwell assays · Bioinform. 2016 |
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification |
0.2 | 1 | 2015 | DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015 |
Bioinformatics and computational biology
cancer genomics |
0.2 | 1 | 2015 | DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015 |
Bioinformatics and computational biology
genomics |
0.2 | 1 | 2015 | DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015 |
Bioinformatics and computational biology › biological network › network biology
network inference |
0.2 | 1 | 2015 | The Cyni framework for network inference in Cytoscape · Bioinform. 2015 |
Bioinformatics and computational biology
systems biology |
0.2 | 1 | 2015 | The Cyni framework for network inference in Cytoscape · Bioinform. 2015 |
Bioinformatics and computational biology › genomics › computational genomics
integrative genomics |
0.1 | 1 | 2010 | geWorkbench: an open source platform for integrative genomics · Bioinform. 2010 |
Bioinformatics and computational biology
transcriptomics |
0.1 | 1 | 2018 | ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's disease · Bioinform. 2018 |
Bioinformatics and computational biology
software infrastructure |
0.1 | 1 | 2015 | The Cyni framework for network inference in Cytoscape · Bioinform. 2015 |
Bioinformatics and computational biology › sequence analysis › motif discovery
pattern discovery |
0.1 | 2 | 2000 | SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000 Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000 |
Information retrieval
indexing |
0.0 | 4 | 1994 | Multidimensional Indexing for Recognizing Visual Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 1994 FLASH: a fast look-up algorithm for string homology · CVPR 1993 Systematic design of indexing strategies for object recognition · CVPR 1993 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 2 | 2000 | SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000 FLASH: a fast look-up algorithm for string homology · CVPR 1993 |
Bioinformatics and computational biology › protein sequence analysis
protein family analysis |
0.0 | 2 | 2000 | Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000 SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000 |
Bioinformatics and computational biology › protein function prediction
protein classification |
0.0 | 1 | 2000 | Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.0 | 2 | 1993 | FLASH: A Fast Look-Up Algorithm for String Homology · ISMB 1993 FLASH: a fast look-up algorithm for string homology · CVPR 1993 |
Computer vision › Image recognition and object detection
shape recognition |
0.0 | 2 | 1994 | Multidimensional Indexing for Recognizing Visual Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 1994 Multidimensional indexing for recognizing visual shapes · CVPR 1991 |
Computer vision › 3D vision
3d object recognition |
0.0 | 2 | 1992 | A Complete and Extendable Approach to Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1992 Visual recognition using concurrent and layered parameter networks · CVPR 1989 |
Computer vision › 3D vision › 3d object recognition
feature-based object recognition |
0.0 | 2 | 1992 | A Complete and Extendable Approach to Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1992 Visual recognition using concurrent and layered parameter networks · CVPR 1989 |
Image and video processing
feature extraction |
0.0 | 2 | 1992 | The Multiple Window Parameter Transform · IEEE Trans. Pattern Anal. Mach. Intell. 1992 Generalized neighborhoods: a new approach to complex parameter feature extraction · CVPR 1989 |
Computer vision › Image recognition and object detection › object recognition › model-based object recognition
indexing-based recognition |
0.0 | 1 | 1993 | Systematic design of indexing strategies for object recognition · CVPR 1993 |
Computer vision › Image recognition and object detection
object recognition |
0.0 | 1 | 1993 | Systematic design of indexing strategies for object recognition · CVPR 1993 |
Information retrieval › indexing
probabilistic indexing |
0.0 | 1 | 1993 | FLASH: a fast look-up algorithm for string homology · CVPR 1993 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.0 | 1 | 1992 | The Multiple Window Parameter Transform · IEEE Trans. Pattern Anal. Mach. Intell. 1992 |
Bioinformatics and computational biology › sequence analysis › motif discovery
conserved motif identification |
0.0 | 1 | 2000 | SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000 |
Computer vision › 3D vision
3d shape representation |
0.0 | 1 | 1991 | Multidimensional indexing for recognizing visual shapes · CVPR 1991 |
Indexing and storage engines
multidimensional indexing |
0.0 | 1 | 1991 | Multidimensional indexing for recognizing visual shapes · CVPR 1991 |
Methods — techniques the papers use, named apart from their topics
ARACNe algorithm · 0.4master regulator analysis · 0.3iterative clustering · 0.3differential gene expression analysis · 0.3spatial autocorrelation · 0.2principal component analysis · 0.2mutual information estimation · 0.2bayesian hierarchical modeling · 0.2adaptive partitioning · 0.2genetical-genomic information theory · 0.2invariant computation · 0.0table look-up · 0.0vote threshold estimation · 0.0quantization analysis · 0.0relaxation · 0.0parameter transforms · 0.0evidence integration · 0.0shape autocorrelation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | pyVIPER: a fast and scalable Python package for protein activity estimation and master regulator analysis of single-cell RNA sequencing dataabstractSingle-cell sequencing has revolutionized biomedical research by offering insights into cellular heterogeneity at unprecedented resolution. Yet, the low signal-to-noise ratio characteristic of single-cell RNA sequencing (scRNA-seq) challenges quantitative analyses. Gene regulatory network (GRN) analysis can help overcome this obstacle, enabling the mechanistic elucidation of cellular state determinants. For instance, the VIPER algorithm can identify Master Regulator proteins from gene expression data. However, as the size and complexity of scRNA-seq datasets grow, the demand for scalable tools supporting the analysis of datasets with up to hundreds of thousands of cells becomes increasingly critical in its original implementation in R. RESULTS: To address this challenge, we introduce pyVIPER, a Python-based tool for protein activity inference from transcriptional data. pyVIPER supports flexible data transformation/postprocessing modules, enrichment analysis algorithms, and features a novel data structure for GRNs manipulation. It integrates seamlessly with scverse, scanpy and widely adopted machine learning libraries. By leveraging PyTorch-based GPU acceleration and optimized core operations, benchmarking demonstrates orders-of-magnitude improvements in runtime efficiency compared to R-based VIPER, reducing analysis time for large datasets from hours to minutes. CONCLUSIONS: pyVIPER is a fast, memory-efficient, and highly scalable Python toolkit for protein activity inference in large-scale scRNA-seq datasets. Its scalability and hardware acceleration enables high-throughput VIPER-based analysis of virtually any single-cell dataset while facilitating integration with other Python-based, including state-of-the-art machine learning workflows. Taken together, these features make pyVIPER a valuable resource to expand the applicability of mechanistic regulatory network-based analysis in single-cell research. Alexander L. E. Wang, Luca Zanella, Zizhao Lin, Filippo Riva, Heeju Noh, Gabriel M. Aizenman, Lukas Vlahos, Miquel Anglada-Girotto, Rowan Cassius, Léo Dupire, Aziz Zafar, Andrea Califano, Alessandro Vasciaveo |
BMC Bioinform. | 12 |
| 2019 | SJARACNe: a scalable software tool for gene network reverse engineering from big dataabstractSUMMARY: Over the last two decades, we have observed an exponential increase in the number of generated array or sequencing-based transcriptomic profiles. Reverse engineering of biological networks from high-throughput gene expression profiles has been one of the grand challenges in systems biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective and widely-used tools to address this challenge. However, existing ARACNe implementations do not efficiently process big input data with thousands of samples. Here we present an improved implementation of the algorithm, SJARACNe, to solve this big data problem, based on sophisticated software engineering. The new scalable SJARACNe package achieves a dramatic improvement in computational performance in both time and memory usage and implements new features while preserving the network inference accuracy of the original algorithm. Given that large-sampled transcriptomic data is increasingly available and ARACNe is extremely demanding for network reconstruction, the scalable SJARACNe will allow even researchers with modest computational resources to efficiently construct complex regulatory and signaling networks from thousands of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: SJARACNe is implemented in C++ (computational core) and Python (pipelining scripting wrapper, ≥3.6.1). It is freely available at https://github.com/jyyulab/SJARACNe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alireza Khatamian, Evan O. Paull, Andrea Califano, Jiyang Yu |
Bioinform. | 3 |
| 2018 | iterClust: a statistical framework for iterative clustering analysisabstractMotivation: In a scenario where populations A, B1 and B2 (subpopulations of B) exist, pronounced differences between A and B may mask subtle differences between B1 and B2. Results: Here we present iterClust, an iterative clustering framework, which can separate more pronounced differences (e.g. A and B) in starting iterations, followed by relatively subtle differences (e.g. B1 and B2), providing a comprehensive clustering trajectory. Availability and implementation: iterClust is implemented as a Bioconductor R package. Supplementary information: Supplementary data are available at Bioinformatics online. Hongxu Ding, Wanxin Wang, Andrea Califano |
Bioinform. | 3 |
| 2018 | ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's diseaseabstractMotivation: In an effort to better understand the molecular drivers of synaptic and neurophysiologic dysfunction in Alzheimer's disease (AD), we analyzed neuronal gene expression data from human AD brain tissue to identify master regulators of synaptic gene expression. Results: Master regulator analysis identifies ZCCHC17 as normally supporting the expression of a network of synaptic genes, and predicts that ZCCHC17 dysfunction in AD leads to lower expression of these genes. We demonstrate that ZCCHC17 is normally expressed in neurons and is reduced early in the course of AD pathology. We show that ZCCHC17 loss in rat neurons leads to lower expression of the majority of the predicted synaptic targets and that ZCCHC17 drives the expression of a similar gene network in humans and rats. These findings support a conserved function for ZCCHC17 between species and identify ZCCHC17 loss as an important early driver of lower synaptic gene expression in AD. Availability and implementation: Matlab and R scripts used in this paper are available at https://github.com/afteich/AD_ZCC. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zeljko Tomljanovic, Mitesh Patel 0002, William Shin, Andrea Califano, Andrew F. Teich |
Bioinform. | 4 |
| 2017 | Systematic, network-based characterization of therapeutic target inhibitorsabstractA large fraction of the proteins that are being identified as key tumor dependencies represent poor pharmacological targets or lack clinically-relevant small-molecule inhibitors. Availability of fully generalizable approaches for the systematic and efficient prioritization of tumor-context specific protein activity inhibitors would thus have significant translational value. Unfortunately, inhibitor effects on protein activity cannot be directly measured in systematic and proteome-wide fashion by conventional biochemical assays. We introduce OncoLead, a novel network based approach for the systematic prioritization of candidate inhibitors for arbitrary targets of therapeutic interest. In vitro and in vivo validation confirmed that OncoLead analysis can recapitulate known inhibitors as well as prioritize novel, context-specific inhibitors of difficult targets, such as MYC and STAT3. We used OncoLead to generate the first unbiased drug/regulator interaction map, representing compounds modulating the activity of cancer-relevant transcription factors, with potential in precision medicine. Mariano J. Alvarez, Brygida C. Bisikirska, Alexander Lachmann, Ronald Realubit, Sergey Pampou, Jorida Coku, Charles Karan, Andrea Califano |
PLoS Comput. Biol. | 9 |
| 2016 | Detection and removal of spatial bias in multiwell assaysabstractMOTIVATION: Multiplex readout assays are now increasingly being performed using microfluidic automation in multiwell format. For instance, the Library of Integrated Network-based Cellular Signatures (LINCS) has produced gene expression measurements for tens of thousands of distinct cell perturbations using a 384-well plate format. This dataset is by far the largest 384-well gene expression measurement assay ever performed. We investigated the gene expression profiles of a million samples from the LINCS dataset and found that the vast majority (96%) of the tested plates were affected by a significant 2D spatial bias. RESULTS: Using a novel algorithm combining spatial autocorrelation detection and principal component analysis, we could remove most of the spatial bias from the LINCS dataset and show in parallel a dramatic improvement of similarity between biological replicates assayed in different plates. The proposed methodology is fully general and can be applied to any highly multiplexed assay performed in multiwell format. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Lachmann, Federico Manuel Giorgi, Mariano J. Alvarez, Andrea Califano |
Bioinform. | 4 |
| 2016 | ARACNe-AP: gene network reverse engineering through adaptive partitioning inference of mutual informationabstractUNLABELLED: The accurate reconstruction of gene regulatory networks from large scale molecular profile datasets represents one of the grand challenges of Systems Biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective tools to accomplish this goal. However, the initial Fixed Bandwidth (FB) implementation is both inefficient and unable to deal with sample sets providing largely uneven coverage of the probability density space. Here, we present a completely new implementation of the algorithm, based on an Adaptive Partitioning strategy (AP) for estimating the Mutual Information. The new AP implementation (ARACNe-AP) achieves a dramatic improvement in computational performance (200× on average) over the previous methodology, while preserving the Mutual Information estimator and the Network inference accuracy of the original algorithm. Given that the previous version of ARACNe is extremely demanding, the new version of the algorithm will allow even researchers with modest computational resources to build complex regulatory networks from hundreds of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: A JAVA cross-platform command line executable of ARACNe, together with all source code and a detailed usage guide are freely available on Sourceforge (http://sourceforge.net/projects/aracne-ap). JAVA version 8 or higher is required. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander Lachmann, Federico Manuel Giorgi, Gonzalo López, Andrea Califano |
Bioinform. | 4 |
| 2016 | ScreenBEAM: a novel meta-analysis algorithm for functional genomics screens via Bayesian hierarchical modelingabstractMOTIVATION: Functional genomics (FG) screens, using RNAi or CRISPR technology, have become a standard tool for systematic, genome-wide loss-of-function studies for therapeutic target discovery. As in many large-scale assays, however, off-target effects, variable reagents' potency and experimental noise must be accounted for appropriately control for false positives. Indeed, rigorous statistical analysis of high-throughput FG screening data remains challenging, particularly when integrative analyses are used to combine multiple sh/sgRNAs targeting the same gene in the library. METHOD: We use large RNAi and CRISPR repositories that are publicly available to evaluate a novel meta-analysis approach for FG screens via Bayesian hierarchical modeling, Screening Bayesian Evaluation and Analysis Method (ScreenBEAM). RESULTS: Results from our analysis show that the proposed strategy, which seamlessly combines all available data, robustly outperforms classical algorithms developed for microarray data sets as well as recent approaches designed for next generation sequencing technologies. Remarkably, the ScreenBEAM algorithm works well even when the quality of FG screens is relatively low, which accounts for about 80-95% of the public datasets. AVAILABILITY AND IMPLEMENTATION: R package and source code are available at: https://github.com/jyyu/ScreenBEAM. CONTACT: [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiyang Yu, Andrea Califano |
Bioinform. | 3 |
| 2015 | DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypesabstractUNLABELLED: Identification of driver mutations in human diseases is often limited by cohort size and availability of appropriate statistical models. We propose a method for the systematic discovery of genetic alterations that are causal determinants of disease, by prioritizing genes upstream of functional disease drivers, within regulatory networks inferred de novo from experimental data. Here we present the implementation of Driver-gene Inference by Genetical-Genomic Information Theory as an R-system package. AVAILABILITY AND IMPLEMENTATION: The diggit package is freely available under the GPL-2 license from Bioconductor (http://www.bioconductor.org). Mariano J. Alvarez, James C. Chen, Andrea Califano |
Bioinform. | 3 |
| 2015 | The Cyni framework for network inference in CytoscapeabstractMOTIVATION: Research on methods for the inference of networks from biological data is making significant advances, but the adoption of network inference in biomedical research practice is lagging behind. Here, we present Cyni, an open-source 'fill-in-the-algorithm' framework that provides common network inference functionality and user interface elements. Cyni allows the rapid transformation of Java-based network inference prototypes into apps of the popular open-source Cytoscape network analysis and visualization ecosystem. Merely placing the resulting app in the Cytoscape App Store makes the method accessible to a worldwide community of biomedical researchers by mouse click. In a case study, we illustrate the transformation of an ARACNE implementation into a Cytoscape app. AVAILABILITY AND IMPLEMENTATION: Cyni, its apps, user guides, documentation and sample code are available from the Cytoscape App Store http://apps.cytoscape.org/apps/cynitoolbox CONTACT: [email protected]. Oriol Guitart, Manjunath Kustagi, Frank Rügheimer, Andrea Califano, Benno Schwikowski |
Bioinform. | 4 |
| 2013 | Improving Breast Cancer Survival Analysis through Competition-Based Multidimensional ModelingabstractBreast cancer is the most common malignancy in women and is responsible for hundreds of thousands of deaths annually. As with most cancers, it is a heterogeneous disease and different breast cancer subtypes are treated differently. Understanding the difference in prognosis for breast cancer based on its molecular and phenotypic features is one avenue for improving treatment by matching the proper treatment with molecular subtypes of the disease. In this work, we employed a competition-based approach to modeling breast cancer prognosis using large datasets containing genomic and clinical information and an online real-time leaderboard program used to speed feedback to the modeling team and to encourage each modeler to work towards achieving a higher ranked submission. We find that machine learning methods combined with molecular features selected based on expert prior knowledge can improve survival predictions compared to current best-in-class methodologies and that ensemble models trained across multiple user submissions systematically outperform individual models within the ensemble. We also find that model scores are highly consistent across multiple independent evaluations. This study serves as the pilot phase of a much larger competition open to the whole research community, with the goal of understanding general strategies for model optimization using clinical and molecular profiling data and providing an objective, transparent system for assessing prognostic models. Erhan Bilal, Janusz Dutkowski, Justin Guinney, In Sock Jang, Benjamin A. Logsdon, Gaurav Pandey 0002, Benjamin A. Sauerwine, Yishai Shimoni, Hans Kristian Moen Vollan, Brigham H. Mecham, Oscar M. Rueda, Jorg Tost, Christina Curtis, Mariano J. Alvarez, Vessela N. Kristensen, Samuel Aparicio, Anne-Lise Børresen-Dale, Carlos Caldas, Andrea Califano, Stephen H. Friend, Trey Ideker, Eric E. Schadt, Gustavo Stolovitzky, Adam A. Margolin |
PLoS Comput. Biol. | 19 |
| 2012 | Using systems and structure biology tools to dissect cellular phenotypesabstractThe Center for the Multiscale Analysis of Genetic Networks (MAGNet, http://magnet.c2b2.columbia.edu) was established in 2005, with the mission of providing the biomedical research community with Structural and Systems Biology algorithms and software tools for the dissection of molecular interactions and for the interaction-based elucidation of cellular phenotypes. Over the last 7 years, MAGNet investigators have developed many novel analysis methodologies, which have led to important biological discoveries, including understanding the role of the DNA shape in protein-DNA binding specificity and the discovery of genes causally related to the presentation of malignant phenotypes, including lymphoma, glioma, and melanoma. Software tools implementing these methodologies have been broadly adopted by the research community and are made freely available through geWorkbench, the Center's integrated analysis platform. Additionally, MAGNet has been instrumental in organizing and developing key conferences and meetings focused on the emerging field of systems biology and regulatory genomics, with special focus on cancer-related research. Aris Floratos, Barry Honig, Dana Pe'er, Andrea Califano |
J. Am. Medical Informatics Assoc. | 4 |
| 2010 | geWorkbench: an open source platform for integrative genomicsabstractSUMMARY: geWorkbench (genomics Workbench) is an open source Java desktop application that provides access to an integrated suite of tools for the analysis and visualization of data from a wide range of genomics domains (gene expression, sequence, protein structure and systems biology). More than 70 distinct plug-in modules are currently available implementing both classical analyses (several variants of clustering, classification, homology detection, etc.) as well as state of the art algorithms for the reverse engineering of regulatory networks and for protein structure prediction, among many others. geWorkbench leverages standards-based middleware technologies to provide seamless access to remote data, annotation and computational servers, thus, enabling researchers with limited local resources to benefit from available public infrastructure. AVAILABILITY: The project site (http://www.geworkbench.org) includes links to self-extracting installers for most operating system (OS) platforms as well as instructions for building the application from scratch using the source code [which is freely available from the project's SVN (subversion) repository]. geWorkbench support is available through the end-user and developer forums of the caBIG Molecular Analysis Tools Knowledge Center, https://cabig-kc.nci.nih.gov/Molecular/forums/ Aris Floratos, Kenneth Smith 0001, Zhou Ji, John Watkinson, Andrea Califano |
Bioinform. | 5 |
| 2009 | New JBI emphasis on translational bioinformatics
Edward H. Shortliffe, Andrea Califano, Lawrence Hunter |
J. Biomed. Informatics | 2 |
| 2007 | ChIP-on-chip significance analysis reveals ubiquitous transcription factor bindingabstractChIP-on-chip technology provides a genome-scale view of transcription factor (TF)/target interactions and a systems-level window into transcriptional regulatory networks. However, while many studies have used ChIP-on-chip data to effectively discover new TF targets, statistical methods have fallen short of developing an accurate model to disassociate signals caused by experimental noise from those caused by true biological variation, thus leveraging the technology to provide high confidence predictions of the full range of interactions. This paper presents a novel method to accurately model the significance of binding events measured by ChIP-on-chip data. For each arrayed probe representing a genomic segment, a ChIP-on-chip microarray measures intensity levels for the IP channel, which is enriched in genomic fragments bound by an immunoprecipitated TF, and the WCE channel, which represents random genomic fragments. Statistical significance is inferred by computing the conditional probability, p ( M | A ), where M = log 2 ( I P W C E ) MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGnbqtcqGH9aqpcyGGSbaBcqGGVbWBcqGGNbWzcqaIYaGmdaqadaqaamaalaaabaGaemysaKKaemiuaafabaGaem4vaCLaem4qamKaemyraueaaaGaayjkaiaawMcaaaaa@3B1B@ and A = log 2 ( I P ) + log 2 ( W C E ) 2 MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGbbqqcqGH9aqpdaWcaaqaaiGbcYgaSjabc+gaVjabcEgaNjabikdaYiabcIcaOiabdMeajjabdcfaqjabcMcaPiabgUcaRiGbcYgaSjabc+gaVjabcEgaNjabikdaYiabcIcaOiabdEfaxjabdoeadjabdweafjabcMcaPaqaaiabikdaYaaaaaa@43C2@ (Fig. 1 ). A kernel density estimation procedure is used to calculate the joint probability, p ( M , A ), and for each average intensity value, the mean of the null distribution (i.e. distribution for unbound probes) is inferred as M ^ A = arg max M p ( M | A ) MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaOGaeyypa0ZaaCbeaeaacyGGHbqycqGGYbGCcqGGNbWzcyGGTbqBcqGGHbqycqGG4baEaSqaaiabd2eanbqabaGccqWGWbaCcqGGOaakcqWGnbqtcqGG8baFcqWGbbqqcqGGPaqkaaa@4089@ . The distribution of p ( M | A ), for M < M ^ A MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaaaa@2F16@ , is then projected across M ^ A MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaaaa@2F16@ to yield the inferred null distribution, which is used to assign statistical significance scores. Probes for replicate experiments and probes with genomic locations within the fragmentation length (~500 bp) are integrated to produce a single significance score for each genomic region. (Left) Magnitude versus amplitude (MA) plot of a ChIP-on-chip hybridization. The x-axis represents the average log2 intensity of the IP and WCE channels, and the y-axis represents the log2 ratio of IP/WCE. The black line represents the mean of the inferred null distribution, and the colored lines represent confidence intervals of .1, .01, and .001 probability. The model reveals an intensity dependent mean and variance of the null distribution, and a large number of probes are significantly enriched in the IP channel. (Right) The axes are the same as in the left panel, and colors represent the -log10 p-value of the null distribution. The method is tested on six different ChIP-on-chip arrays representing replicate experiments for three different TFs (NOTCH1, MYC and HES1). For each experiment, this analysis reveals an order of magnitude more genomic binding events than detected by traditional methods, predicting several thousand interactions for each TF and suggesting previously unappreciated complexity of transcriptional regulatory networks. Several independent experiments are used to provide evidence about the validity of these predictions. First, biochemical validation of more than 20 predicted targets by gene specific ChIP and qPCR confirm the accuracy of false discovery rate statistics computed by the method. Second, binding site enrichment analysis indicates that the strength of binding site signals are maintained over several thousand promoters. Finally, gene expression analysis reveals a coordinated downregulation of gene expression for the entire range of predicted NOTCH1 bound genes upon NOTCH1 inhibition experiments in cell lines, indicating that a large percentage of bound genes are also functionally regulated by NOTCH1. Adam A. Margolin, Teresa Palomero, Adolfo A. Ferrando, Andrea Califano, Gustavo Stolovitzky |
BMC Bioinform. | 4 |
| 2007 | Hui-Huang Hsu, Advanced Data Mining Technologies in Bioinformatics, Editors: Idea Group Publishing, pp 239
Andrea Califano |
J. Biomed. Informatics | 1 |
| 2006 | Genome-Wide Discovery of Modulators of Transcriptional Interactions in Human B Lymphocytes
Kai Wang 0045, Ilya Nemenman, Nilanjana Banerjee, Adam A. Margolin, Andrea Califano |
RECOMB | 5 |
| 2006 | ARACNE: An Algorithm for the Reconstruction of Gene Regulatory Networks in a Mammalian Cellular ContextabstractBACKGROUND: Elucidating gene regulatory networks is crucial for understanding normal cell physiology and complex pathologic phenotypes. Existing computational methods for the genome-wide "reverse engineering" of such networks have been successful only for lower eukaryotes with simple genomes. Here we present ARACNE, a novel algorithm, using microarray expression profiles, specifically designed to scale up to the complexity of regulatory networks in mammalian cells, yet general enough to address a wider range of network deconvolution problems. This method uses an information theoretic approach to eliminate the majority of indirect interactions inferred by co-expression methods. RESULTS: We prove that ARACNE reconstructs the network exactly (asymptotically) if the effect of loops in the network topology is negligible, and we show that the algorithm works well in practice, even in the presence of numerous loops and complex topologies. We assess ARACNE's ability to reconstruct transcriptional regulatory networks using both a realistic synthetic dataset and a microarray dataset from human B cells. On synthetic datasets ARACNE achieves very low error rates and outperforms established methods, such as Relevance Networks and Bayesian Networks. Application to the deconvolution of genetic networks in human B cells demonstrates ARACNE's ability to infer validated transcriptional targets of the cMYC proto-oncogene. We also study the effects of misestimation of mutual information on network reconstruction, and show that algorithms based on mutual information ranking are more resilient to estimation errors. CONCLUSION: ARACNE shows promise in identifying direct transcriptional interactions in mammalian cellular networks, a problem that has challenged existing reverse engineering algorithms. This approach should enhance our ability to use microarray data to elucidate functional mechanisms that underlie cellular processes and to identify molecular targets of pharmacological compounds in mammalian cellular networks. Adam A. Margolin, Ilya Nemenman, Katia Basso, Chris Wiggins 0001, Gustavo Stolovitzky, Riccardo Dalla Favera, Andrea Califano |
BMC Bioinform. | 7 |
| 2000 | Analysis of Gene Expression Microarrays for Phenotype Classification
Andrea Califano, Gustavo Stolovitzky, Yuhai Tu |
ISMB | 1 |
| 2000 | Systematic and automated discovery of patterns in PROSITE familiesabstractPROSITE is a method for protein classification which relies on a database of biologically significant sites and patterns in protein sequences. Most patterns in PROSITE have been gathered by a a labor intensive combination of experimental characterization of functional residues and sequence alignment. In this paper we present a new and efficient supervised learning procedure, based on the Splash deterministic pattern discovery algorithm and on a framework to assess the statistical significance of patterns. We demonstrate its application to the fully automatic discovery of patterns in 974 PROSITE families. For these families, Splash generates patterns with better specificity and/or sensitivity in 28%, identical statistics in 48%, and worse statistics in 15% of the cases; for the remaining families, patterns exhibited mixed behavior. Second, we have characterized the amount of overlap, on the sequences, between newly discovered patterns and those in PROSITE. In about 75% of the cases, Splash patterns identify sequence sites that overlap more than 50% with those reported in PROSITE. Of the 272 patterns which perform strictly better than the corresponding PROSITE pattern, 178 show more than 70% overlap with the PROSITE pattern. Third, our results suggest that the statistical significance of discovered patterns correlates well with their biological significance. Finally, we use the trypsin subfamily of serine proteases to illustrate the use of this method to exhaustively discover all motifs in a family that are statistically and biologically significant. The complete analysis is sufficiently rapid, taking less than a day for all PROSITE families, to enable the use this methodology for routine curation of existing motif and profile databases. Reece K. Hart, Ajay K. Royyuru, Gustavo Stolovitzky, Andrea Califano |
RECOMB | 4 |
| 2000 | SPLASH: structural pattern localization analysis by sequential histogramsabstractMOTIVATION: The discovery of sparse amino acid patterns that match repeatedly in a set of protein sequences is an important problem in computational biology. Statistically significant patterns, that is patterns that occur more frequently than expected, may identify regions that have been preserved by evolution and which may therefore play a key functional or structural role. Sparseness can be important because a handful of non-contiguous residues may play a key role, while others, in between, may be changed without significant loss of function or structure. Similar arguments may be applied to conserved DNA patterns. Available sparse pattern discovery algorithms are either inefficient or impose limitations on the type of patterns that can be discovered. RESULTS: This paper introduces a deterministic pattern discovery algorithm, called Splash, which can find sparse amino or nucleic acid patterns matching identically or similarly in a set of protein or DNA sequences. Sparse patterns of any length, up to the size of the input sequence, can be discovered without significant loss in performances. Splash is extremely efficient and embarrassingly parallel by nature. Large databases, such as a complete genome or the non-redundant SWISS-PROT database can be processed in a few hours on a typical workstation. Alternatively, a protein family or superfamily, with low overall homology, can be analyzed to discover common functional or structural signatures. Some examples of biologically interesting motifs discovered by Splash are reported for the histone I and for the G-Protein Coupled Receptor families. Due to its efficiency, Splash can be used to systematically and exhaustively identify conserved regions in protein family sets. These can then be used to build accurate and sensitive PSSM or HMM models for sequence analysis. AVAILABILITY: Splash is available to non-commercial research centers upon request, conditional on the signing of a test field agreement. CONTACT: [email protected], Splash main page http://www.research.ibm.com/splash Andrea Califano |
Bioinform. | 1 |
| 1998 | A High-Dimensional Indexing Scheme for Scalable Fingerprint-Based Identification
Andrea Califano, Bob Germain, Scott Colville |
ACCV (1) | 1 |
| 1996 | Data- and Model-Driven Multiresolution Processing
Andrea Califano, Rick Kjeldsen, Ruud M. Bolle |
Comput. Vis. Image Underst. | 1 |
| 1994 | Multidimensional Indexing for Recognizing Visual ShapesabstractThis paper introduces an analytical framework for studying some properties of model acquisition and recognition techniques based on indexing. The goal is to demonstrate that several problems previously associated with the approach can be attributed to the low dimensionality of invariants used. These include limited index selectivity. Excessive accumulation of votes in the look-up table buckets, and excessive sensitivity to quantization parameters. Theoretical results demonstrate that using high-dimensional, highly descriptive global invariants produces better results in terms of accuracy, false positive suppression, and computation time. A practical example of high-dimensional global invariants is introduced and used to implement a 2-D shape acquisition/recognition system. The acquisition/recognition system is based on a two-step table look-up mechanism. First, local curve descriptors are obtained by correlating image contour information at short range. Then, seven-dimensional global invariants are computed by correlating triplets of local curve descriptors at longer range. This experimental system is meant to illustrate the behavior of a high-dimensional indexing scheme. Indeed, its performance shows good agreement with the analytical model with respect to database size, fault tolerance, and recognition speed. Model acquisition time is linear to cubic in the number of object features. Object recognition time is constant to linear in the number of models in the database and linear to cubic in the number of features in the image. The system has been tested extensively. With more than 250 arbitrary shapes in the database. Unsupervised shape and subpart acquisition is demonstrated.> Andrea Califano, Rakesh Mohan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | Systematic design of indexing strategies for object recognitionabstractThe authors analyze how parameters of indexing based recognition systems affects their performance. The main result is that increasing the dimensionality of the indices leads to significantly improved discrimination, false positive suppression and reduced recognition times. With increase in index dimensionality, coarser quantization is required to allow index match in the presence of noise. The authors' analysis also allows estimation of votes thresholds for recognition and estimation of the amount of occlusion that can be tolerated by indexing schemes for given levels of recognition confidence.> Andrea Califano, Rakesh Mohan |
CVPR | 1 |
| 1993 | FLASH: a fast look-up algorithm for string homologyabstractA key issue in managing large amounts of data is the availability of efficient, accurate, ad selective techniques to detect homology (similarity) between newly recovered and previously acquired sequences. The algorithm presented is based on a probabilistic indexing framework which requires minimal access to the database for each match. A highly redundant number of descriptive tuples from the sequences of interest are generated and used as indices in a table look-up paradigm. Theoretical and experimental results on the sensitivity and accuracy of the approach are provided. These include the probability of correct and random matches and the storage and computational requirements. An experimental system is implemented for a database containing the complete genome of the bacteria E. Coli (approximately 2 million nucleotides). Search time is a few seconds on a workstation class machine. The algorithm is shown to scale well to databases containing billions of nucleotides with performances that are orders of magnitude better than the fastest of the current techniques.> Andrea Califano, Isidore Rigoutsos |
CVPR | 1 |
| 1993 | FLASH: A Fast Look-Up Algorithm for String Homology
Andrea Califano, Isidore Rigoutsos |
ISMB | 1 |
| 1992 | A Complete and Extendable Approach to Visual RecognitionabstractA framework for 3D object recognition is presented. Its flexibility and extensibility are accomplished through a uniform, parallel, and modular recognition architecture. Concurrent and stacked parameter transforms reconstruct a variety of features from the input scene. At each stage, constraint satisfaction networks collect and fuse the evidence obtained through the parameter transforms, ensuring a globally consistent interpretation of the input scene and allowing for the integration of diverse types of information. The final interpretation of the scene is a small consistent subset of the many initial hypotheses about partial features, primitive features, feature assemblies, and 3D objects computed by the various parameter transforms. A complete, integrated, and implemented system that extracts planar surfaces, patches of quadrics of revolution, and planar intersection curves of these surfaces from a depth map viewing 3D objects is described. Experimental results on the recognition behavior of the system are presented.> Ruud M. Bolle, Andrea Califano, Rick Kjeldsen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | The Multiple Window Parameter TransformabstractThe multiwindow transform, an extension of parameter transform techniques that increase performance and scope by exploiting the long-range correlated information contained in multiple portions of an image, is presented. Multiple-window transforms allow the extraction of high-dimensional features with improvement in accuracy over conventional techniques while keeping linear to low-order-polynomial computational and space requirements with respect to image size and dimensionality of the features. Using correlated information provides a direct link between extracted features and supporting regions in the image. This, coupled with evidence integration techniques, is used to suppress noisy or nonexistent feature hypotheses. Parameter spaces are implemented as constraint satisfaction networks, where feature hypotheses with overlapping support in the image compete. After an iterative relaxation phase, surviving hypotheses have disjoint support, forming a segmentation of the image. Examples show the performance and provide insight about the behavior.> Andrea Califano, Ruud M. Bolle |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1991 | Multidimensional indexing for recognizing visual shapesabstractA homogeneous approach for acquisition, storage, and recognition of nonparametric shapes from images, using a novel shape representation based on shape autocorrelation operators is presented. A theoretical and experimental analysis of the computational complexity, recognition performance with increasing database size, and fault tolerance of the approach is presented. The system has been tested extensively with more than 300 arbitrary shapes in the database. Using a set of complex shapes, the recognition behavior with respect to occlusion, geometric transformation, and cluttered environments is studied. Unsupervised shape and subpart acquisition is demonstrated.> Andrea Califano, Rakesh Mohan |
CVPR | 1 |
| 1990 | Generalized Shape Autocorrelation
Andrea Califano, Rakesh Mohan |
AAAI | 1 |
| 1990 | Active 3D object modelsabstractA novel approach is presented for pruning the amount of search needed to match image features to object models. The technique relies on active networks which capture various visibility and geometric constraints between features of a model to prune these features from search space during matching. The networks, which can be efficiently implemented in Boolean logic, integrate harmoniously with the previous work in feature recognition and object matching. A method is proposed for clustering model features (vsets) and four types of constraints which assist in building the networks. The authors show, both analytically and empirically, the dramatic reduction in search provided by activation nets.> Ruud M. Bolle, Andrea Califano, Rick Kjeldsen, Rakesh Mohan |
ICCV | 2 |
| 1990 | Data and model driven foveationabstractA general framework for multiresolution visual recognition is introduced. The input is processed simultaneously at a coarse resolution throughout the image and at finer resolution within a small window. An approach for controlling the movement of the high-resolution window is described which allows for the unification of a variety of data and model-driven behavioral paradigms. Three modes have been implemented, one based on large unexplained areas in the data, one on conflicts in the object-model database, and one on a 2D space-filling algorithm. It is argued that this kind of multiresolution processing not only is useful in limiting the computational time, but also can be a deciding factor in making the entire vision problem a tractable and stable one. To demonstrate the approach, a class of 3D surface textures is introduced as a feature for recognition in the system considered. Surface texture recognition typically requires higher-resolution processing than required for the extraction of the underlying surface. As an example, surface texture is used to discriminate between a ping-pong ball and a golf ball.> Andrea Califano, Rick Kjeldsen, Ruud M. Bolle |
ICPR (1) | 1 |
| 1989 | Visual recognition using concurrent and layered parameter networksabstractA vision system to recognize 3-D objects is presented. A novel notion of generalized feature makes it possible to develop a homogeneous architecture to support recognition from simple partial features to complex feature assemblies and 3-D objects. Layered concurrent parameter transforms vote for feature hypotheses on the basis of image data and previously reconstructed features. Recognition networks, motivated by connectionist systems, collect votes, fuse evidence from various sources and ensure global consistency. In addition, they provide an integrated solution to the segmentation problem. The highly modular design allows fundamentally different types of features to interact in a coherent way, utilizing redundancies for more robust recognition. Within this paradigm, a system extracts planar patches, patches of quadrics of revolution, and the intersection curves of these surfaces (lines and conic sections in three-space) from a depth map. Reconstructed features index into a model database to form consistent object hypotheses. Experimental results detailing the recognition behavior or real depth maps are included.> Ruud M. Bolle, Andrea Califano, Rick Kjeldsen, Russell W. Taylor |
CVPR | 2 |
| 1989 | Generalized neighborhoods: a new approach to complex parameter feature extractionabstractA generalized neighborhood concept is presented which extends the usual techniques for feature extraction using parameter transforms. Generalized neighborhoods allow operators to use the joint information contained in distant portions of the same feature; i.e. to utilize the long-distance correlation present in the image. The generalized neighborhood techniques, by correlating local information over different portions of the image, produce up to two orders of magnitude improvement in accuracy over conventional techniques. The response also becomes more complicated; false features may be detected due to a peculiar form of correlated noise. A general framework, motivated by connectionist networks, is presented which eliminates this behaviour by introducing competitive processes in the parameter spaces. A novel approach to the generation of lateral inhibition links in the networks is proposed which is consistent with generalized neighborhoods. Experiments are provided that show results on range data. Complex surfaces and 3-D surface-intersection curves are reconstructed from the data.> Andrea Califano, Ruud M. Bolle, Russell W. Taylor |
CVPR | 1 |
| 1989 | A Homogeneous Framework for Visual Recognition
Rick Kjeldsen, Ruud M. Bolle, Andrea Califano, Russell W. Taylor |
IJCAI | 3 |
| 1988 | Feature Recognition Using Correlated Information Contained in Multiple Neighborboods
Andrea Califano |
AAAI | 1 |