Andrea Califano

dblp:55/1524 · DBLP profile ↗
← Back
37ranked-venue papers
15as first author
1since 2021 · last 2026
0000-0003-4742-3679ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 15 · 11 first-authorGraphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
15 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
9 papers
3D vision · 48% Image recognition and object detection · 39% Segmentation and scene understanding · 8%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.842016
ARACNe-AP: gene network reverse engineering through adaptive partitioning inference of mutual information · Bioinform. 2016
The Cyni framework for network inference in Cytoscape · Bioinform. 2015
DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015
Bioinformatics and computational biology
gene expression analysis
0.422018
iterClust: a statistical framework for iterative clustering analysis · Bioinform. 2018
Analysis of Gene Expression Microarrays for Phenotype Classification · ISMB 2000
Bioinformatics and computational biology › gene regulation › gene regulatory network
gene regulatory network analysis
0.312018
ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's disease · Bioinform. 2018
Bioinformatics and computational biology
functional genomics
0.212016
ScreenBEAM: a novel meta-analysis algorithm for functional genomics screens via Bayesian hierarchical modeling · Bioinform. 2016
Bioinformatics and computational biology › drug discovery
high-throughput screening
0.212016
Detection and removal of spatial bias in multiwell assays · Bioinform. 2016
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification
0.212015
DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015
Bioinformatics and computational biology
cancer genomics
0.212015
DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015
Bioinformatics and computational biology
genomics
0.212015
DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes · Bioinform. 2015
Bioinformatics and computational biology › biological network › network biology
network inference
0.212015
The Cyni framework for network inference in Cytoscape · Bioinform. 2015
Bioinformatics and computational biology
systems biology
0.212015
The Cyni framework for network inference in Cytoscape · Bioinform. 2015
Bioinformatics and computational biology › genomics › computational genomics
integrative genomics
0.112010
geWorkbench: an open source platform for integrative genomics · Bioinform. 2010
Bioinformatics and computational biology
transcriptomics
0.112018
ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's disease · Bioinform. 2018
Bioinformatics and computational biology
software infrastructure
0.112015
The Cyni framework for network inference in Cytoscape · Bioinform. 2015
Bioinformatics and computational biology › sequence analysis › motif discovery
pattern discovery
0.122000
SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000
Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000
Information retrieval
indexing
0.041994
Multidimensional Indexing for Recognizing Visual Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 1994
FLASH: a fast look-up algorithm for string homology · CVPR 1993
Systematic design of indexing strategies for object recognition · CVPR 1993
Bioinformatics and computational biology
sequence analysis
0.022000
SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000
FLASH: a fast look-up algorithm for string homology · CVPR 1993
Bioinformatics and computational biology › protein sequence analysis
protein family analysis
0.022000
Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000
SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000
Bioinformatics and computational biology › protein function prediction
protein classification
0.012000
Systematic and automated discovery of patterns in PROSITE families · RECOMB 2000
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.021993
FLASH: A Fast Look-Up Algorithm for String Homology · ISMB 1993
FLASH: a fast look-up algorithm for string homology · CVPR 1993
Computer vision › Image recognition and object detection
shape recognition
0.021994
Multidimensional Indexing for Recognizing Visual Shapes · IEEE Trans. Pattern Anal. Mach. Intell. 1994
Multidimensional indexing for recognizing visual shapes · CVPR 1991
Computer vision › 3D vision
3d object recognition
0.021992
A Complete and Extendable Approach to Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1992
Visual recognition using concurrent and layered parameter networks · CVPR 1989
Computer vision › 3D vision › 3d object recognition
feature-based object recognition
0.021992
A Complete and Extendable Approach to Visual Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1992
Visual recognition using concurrent and layered parameter networks · CVPR 1989
Image and video processing
feature extraction
0.021992
The Multiple Window Parameter Transform · IEEE Trans. Pattern Anal. Mach. Intell. 1992
Generalized neighborhoods: a new approach to complex parameter feature extraction · CVPR 1989
Computer vision › Image recognition and object detection › object recognition › model-based object recognition
indexing-based recognition
0.011993
Systematic design of indexing strategies for object recognition · CVPR 1993
Computer vision › Image recognition and object detection
object recognition
0.011993
Systematic design of indexing strategies for object recognition · CVPR 1993
Information retrieval › indexing
probabilistic indexing
0.011993
FLASH: a fast look-up algorithm for string homology · CVPR 1993
Computer vision › Segmentation and scene understanding
image segmentation
0.011992
The Multiple Window Parameter Transform · IEEE Trans. Pattern Anal. Mach. Intell. 1992
Bioinformatics and computational biology › sequence analysis › motif discovery
conserved motif identification
0.012000
SPLASH: structural pattern localization analysis by sequential histograms · Bioinform. 2000
Computer vision › 3D vision
3d shape representation
0.011991
Multidimensional indexing for recognizing visual shapes · CVPR 1991
Indexing and storage engines
multidimensional indexing
0.011991
Multidimensional indexing for recognizing visual shapes · CVPR 1991

Methods — techniques the papers use, named apart from their topics

ARACNe algorithm · 0.4master regulator analysis · 0.3iterative clustering · 0.3differential gene expression analysis · 0.3spatial autocorrelation · 0.2principal component analysis · 0.2mutual information estimation · 0.2bayesian hierarchical modeling · 0.2adaptive partitioning · 0.2genetical-genomic information theory · 0.2invariant computation · 0.0table look-up · 0.0vote threshold estimation · 0.0quantization analysis · 0.0relaxation · 0.0parameter transforms · 0.0evidence integration · 0.0shape autocorrelation · 0.0
YearPublicationVenuePosition
2026 pyVIPER: a fast and scalable Python package for protein activity estimation and master regulator analysis of single-cell RNA sequencing data
abstract
Single-cell sequencing has revolutionized biomedical research by offering insights into cellular heterogeneity at unprecedented resolution. Yet, the low signal-to-noise ratio characteristic of single-cell RNA sequencing (scRNA-seq) challenges quantitative analyses. Gene regulatory network (GRN) analysis can help overcome this obstacle, enabling the mechanistic elucidation of cellular state determinants. For instance, the VIPER algorithm can identify Master Regulator proteins from gene expression data. However, as the size and complexity of scRNA-seq datasets grow, the demand for scalable tools supporting the analysis of datasets with up to hundreds of thousands of cells becomes increasingly critical in its original implementation in R. RESULTS: To address this challenge, we introduce pyVIPER, a Python-based tool for protein activity inference from transcriptional data. pyVIPER supports flexible data transformation/postprocessing modules, enrichment analysis algorithms, and features a novel data structure for GRNs manipulation. It integrates seamlessly with scverse, scanpy and widely adopted machine learning libraries. By leveraging PyTorch-based GPU acceleration and optimized core operations, benchmarking demonstrates orders-of-magnitude improvements in runtime efficiency compared to R-based VIPER, reducing analysis time for large datasets from hours to minutes. CONCLUSIONS: pyVIPER is a fast, memory-efficient, and highly scalable Python toolkit for protein activity inference in large-scale scRNA-seq datasets. Its scalability and hardware acceleration enables high-throughput VIPER-based analysis of virtually any single-cell dataset while facilitating integration with other Python-based, including state-of-the-art machine learning workflows. Taken together, these features make pyVIPER a valuable resource to expand the applicability of mechanistic regulatory network-based analysis in single-cell research.
Alexander L. E. Wang, Luca Zanella, Zizhao Lin, Filippo Riva, Heeju Noh, Gabriel M. Aizenman, Lukas Vlahos, Miquel Anglada-Girotto, Rowan Cassius, Léo Dupire, Aziz Zafar, Andrea Califano, Alessandro Vasciaveo
BMC Bioinform.12
2019 SJARACNe: a scalable software tool for gene network reverse engineering from big data
abstract
SUMMARY: Over the last two decades, we have observed an exponential increase in the number of generated array or sequencing-based transcriptomic profiles. Reverse engineering of biological networks from high-throughput gene expression profiles has been one of the grand challenges in systems biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective and widely-used tools to address this challenge. However, existing ARACNe implementations do not efficiently process big input data with thousands of samples. Here we present an improved implementation of the algorithm, SJARACNe, to solve this big data problem, based on sophisticated software engineering. The new scalable SJARACNe package achieves a dramatic improvement in computational performance in both time and memory usage and implements new features while preserving the network inference accuracy of the original algorithm. Given that large-sampled transcriptomic data is increasingly available and ARACNe is extremely demanding for network reconstruction, the scalable SJARACNe will allow even researchers with modest computational resources to efficiently construct complex regulatory and signaling networks from thousands of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: SJARACNe is implemented in C++ (computational core) and Python (pipelining scripting wrapper, ≥3.6.1). It is freely available at https://github.com/jyyulab/SJARACNe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alireza Khatamian, Evan O. Paull, Andrea Califano, Jiyang Yu
Bioinform.3
2018 iterClust: a statistical framework for iterative clustering analysis
abstract
Motivation: In a scenario where populations A, B1 and B2 (subpopulations of B) exist, pronounced differences between A and B may mask subtle differences between B1 and B2. Results: Here we present iterClust, an iterative clustering framework, which can separate more pronounced differences (e.g. A and B) in starting iterations, followed by relatively subtle differences (e.g. B1 and B2), providing a comprehensive clustering trajectory. Availability and implementation: iterClust is implemented as a Bioconductor R package. Supplementary information: Supplementary data are available at Bioinformatics online.
Hongxu Ding, Wanxin Wang, Andrea Califano
Bioinform.3
2018 ZCCHC17 is a master regulator of synaptic gene expression in Alzheimer's disease
abstract
Motivation: In an effort to better understand the molecular drivers of synaptic and neurophysiologic dysfunction in Alzheimer's disease (AD), we analyzed neuronal gene expression data from human AD brain tissue to identify master regulators of synaptic gene expression. Results: Master regulator analysis identifies ZCCHC17 as normally supporting the expression of a network of synaptic genes, and predicts that ZCCHC17 dysfunction in AD leads to lower expression of these genes. We demonstrate that ZCCHC17 is normally expressed in neurons and is reduced early in the course of AD pathology. We show that ZCCHC17 loss in rat neurons leads to lower expression of the majority of the predicted synaptic targets and that ZCCHC17 drives the expression of a similar gene network in humans and rats. These findings support a conserved function for ZCCHC17 between species and identify ZCCHC17 loss as an important early driver of lower synaptic gene expression in AD. Availability and implementation: Matlab and R scripts used in this paper are available at https://github.com/afteich/AD_ZCC. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Zeljko Tomljanovic, Mitesh Patel 0002, William Shin, Andrea Califano, Andrew F. Teich
Bioinform.4
2017 Systematic, network-based characterization of therapeutic target inhibitors
abstract
A large fraction of the proteins that are being identified as key tumor dependencies represent poor pharmacological targets or lack clinically-relevant small-molecule inhibitors. Availability of fully generalizable approaches for the systematic and efficient prioritization of tumor-context specific protein activity inhibitors would thus have significant translational value. Unfortunately, inhibitor effects on protein activity cannot be directly measured in systematic and proteome-wide fashion by conventional biochemical assays. We introduce OncoLead, a novel network based approach for the systematic prioritization of candidate inhibitors for arbitrary targets of therapeutic interest. In vitro and in vivo validation confirmed that OncoLead analysis can recapitulate known inhibitors as well as prioritize novel, context-specific inhibitors of difficult targets, such as MYC and STAT3. We used OncoLead to generate the first unbiased drug/regulator interaction map, representing compounds modulating the activity of cancer-relevant transcription factors, with potential in precision medicine.
Mariano J. Alvarez, Brygida C. Bisikirska, Alexander Lachmann, Ronald Realubit, Sergey Pampou, Jorida Coku, Charles Karan, Andrea Califano
PLoS Comput. Biol.9
2016 Detection and removal of spatial bias in multiwell assays
abstract
MOTIVATION: Multiplex readout assays are now increasingly being performed using microfluidic automation in multiwell format. For instance, the Library of Integrated Network-based Cellular Signatures (LINCS) has produced gene expression measurements for tens of thousands of distinct cell perturbations using a 384-well plate format. This dataset is by far the largest 384-well gene expression measurement assay ever performed. We investigated the gene expression profiles of a million samples from the LINCS dataset and found that the vast majority (96%) of the tested plates were affected by a significant 2D spatial bias. RESULTS: Using a novel algorithm combining spatial autocorrelation detection and principal component analysis, we could remove most of the spatial bias from the LINCS dataset and show in parallel a dramatic improvement of similarity between biological replicates assayed in different plates. The proposed methodology is fully general and can be applied to any highly multiplexed assay performed in multiwell format. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Lachmann, Federico Manuel Giorgi, Mariano J. Alvarez, Andrea Califano
Bioinform.4
2016 ARACNe-AP: gene network reverse engineering through adaptive partitioning inference of mutual information
abstract
UNLABELLED: The accurate reconstruction of gene regulatory networks from large scale molecular profile datasets represents one of the grand challenges of Systems Biology. The Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNe) represents one of the most effective tools to accomplish this goal. However, the initial Fixed Bandwidth (FB) implementation is both inefficient and unable to deal with sample sets providing largely uneven coverage of the probability density space. Here, we present a completely new implementation of the algorithm, based on an Adaptive Partitioning strategy (AP) for estimating the Mutual Information. The new AP implementation (ARACNe-AP) achieves a dramatic improvement in computational performance (200× on average) over the previous methodology, while preserving the Mutual Information estimator and the Network inference accuracy of the original algorithm. Given that the previous version of ARACNe is extremely demanding, the new version of the algorithm will allow even researchers with modest computational resources to build complex regulatory networks from hundreds of gene expression profiles. AVAILABILITY AND IMPLEMENTATION: A JAVA cross-platform command line executable of ARACNe, together with all source code and a detailed usage guide are freely available on Sourceforge (http://sourceforge.net/projects/aracne-ap). JAVA version 8 or higher is required. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alexander Lachmann, Federico Manuel Giorgi, Gonzalo López, Andrea Califano
Bioinform.4
2016 ScreenBEAM: a novel meta-analysis algorithm for functional genomics screens via Bayesian hierarchical modeling
abstract
MOTIVATION: Functional genomics (FG) screens, using RNAi or CRISPR technology, have become a standard tool for systematic, genome-wide loss-of-function studies for therapeutic target discovery. As in many large-scale assays, however, off-target effects, variable reagents' potency and experimental noise must be accounted for appropriately control for false positives. Indeed, rigorous statistical analysis of high-throughput FG screening data remains challenging, particularly when integrative analyses are used to combine multiple sh/sgRNAs targeting the same gene in the library. METHOD: We use large RNAi and CRISPR repositories that are publicly available to evaluate a novel meta-analysis approach for FG screens via Bayesian hierarchical modeling, Screening Bayesian Evaluation and Analysis Method (ScreenBEAM). RESULTS: Results from our analysis show that the proposed strategy, which seamlessly combines all available data, robustly outperforms classical algorithms developed for microarray data sets as well as recent approaches designed for next generation sequencing technologies. Remarkably, the ScreenBEAM algorithm works well even when the quality of FG screens is relatively low, which accounts for about 80-95% of the public datasets. AVAILABILITY AND IMPLEMENTATION: R package and source code are available at: https://github.com/jyyu/ScreenBEAM. CONTACT: [email protected], [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiyang Yu, Andrea Califano
Bioinform.3
2015 DIGGIT: a Bioconductor package to infer genetic variants driving cellular phenotypes
abstract
UNLABELLED: Identification of driver mutations in human diseases is often limited by cohort size and availability of appropriate statistical models. We propose a method for the systematic discovery of genetic alterations that are causal determinants of disease, by prioritizing genes upstream of functional disease drivers, within regulatory networks inferred de novo from experimental data. Here we present the implementation of Driver-gene Inference by Genetical-Genomic Information Theory as an R-system package. AVAILABILITY AND IMPLEMENTATION: The diggit package is freely available under the GPL-2 license from Bioconductor (http://www.bioconductor.org).
Mariano J. Alvarez, James C. Chen, Andrea Califano
Bioinform.3
2015 The Cyni framework for network inference in Cytoscape
abstract
MOTIVATION: Research on methods for the inference of networks from biological data is making significant advances, but the adoption of network inference in biomedical research practice is lagging behind. Here, we present Cyni, an open-source 'fill-in-the-algorithm' framework that provides common network inference functionality and user interface elements. Cyni allows the rapid transformation of Java-based network inference prototypes into apps of the popular open-source Cytoscape network analysis and visualization ecosystem. Merely placing the resulting app in the Cytoscape App Store makes the method accessible to a worldwide community of biomedical researchers by mouse click. In a case study, we illustrate the transformation of an ARACNE implementation into a Cytoscape app. AVAILABILITY AND IMPLEMENTATION: Cyni, its apps, user guides, documentation and sample code are available from the Cytoscape App Store http://apps.cytoscape.org/apps/cynitoolbox CONTACT: [email protected].
Oriol Guitart, Manjunath Kustagi, Frank Rügheimer, Andrea Califano, Benno Schwikowski
Bioinform.4
2013 Improving Breast Cancer Survival Analysis through Competition-Based Multidimensional Modeling
abstract
Breast cancer is the most common malignancy in women and is responsible for hundreds of thousands of deaths annually. As with most cancers, it is a heterogeneous disease and different breast cancer subtypes are treated differently. Understanding the difference in prognosis for breast cancer based on its molecular and phenotypic features is one avenue for improving treatment by matching the proper treatment with molecular subtypes of the disease. In this work, we employed a competition-based approach to modeling breast cancer prognosis using large datasets containing genomic and clinical information and an online real-time leaderboard program used to speed feedback to the modeling team and to encourage each modeler to work towards achieving a higher ranked submission. We find that machine learning methods combined with molecular features selected based on expert prior knowledge can improve survival predictions compared to current best-in-class methodologies and that ensemble models trained across multiple user submissions systematically outperform individual models within the ensemble. We also find that model scores are highly consistent across multiple independent evaluations. This study serves as the pilot phase of a much larger competition open to the whole research community, with the goal of understanding general strategies for model optimization using clinical and molecular profiling data and providing an objective, transparent system for assessing prognostic models.
Erhan Bilal, Janusz Dutkowski, Justin Guinney, In Sock Jang, Benjamin A. Logsdon, Gaurav Pandey 0002, Benjamin A. Sauerwine, Yishai Shimoni, Hans Kristian Moen Vollan, Brigham H. Mecham, Oscar M. Rueda, Jorg Tost, Christina Curtis, Mariano J. Alvarez, Vessela N. Kristensen, Samuel Aparicio, Anne-Lise Børresen-Dale, Carlos Caldas, Andrea Califano, Stephen H. Friend, Trey Ideker, Eric E. Schadt, Gustavo Stolovitzky, Adam A. Margolin
PLoS Comput. Biol.19
2012 Using systems and structure biology tools to dissect cellular phenotypes
abstract
The Center for the Multiscale Analysis of Genetic Networks (MAGNet, http://magnet.c2b2.columbia.edu) was established in 2005, with the mission of providing the biomedical research community with Structural and Systems Biology algorithms and software tools for the dissection of molecular interactions and for the interaction-based elucidation of cellular phenotypes. Over the last 7 years, MAGNet investigators have developed many novel analysis methodologies, which have led to important biological discoveries, including understanding the role of the DNA shape in protein-DNA binding specificity and the discovery of genes causally related to the presentation of malignant phenotypes, including lymphoma, glioma, and melanoma. Software tools implementing these methodologies have been broadly adopted by the research community and are made freely available through geWorkbench, the Center's integrated analysis platform. Additionally, MAGNet has been instrumental in organizing and developing key conferences and meetings focused on the emerging field of systems biology and regulatory genomics, with special focus on cancer-related research.
Aris Floratos, Barry Honig, Dana Pe'er, Andrea Califano
J. Am. Medical Informatics Assoc.4
2010 geWorkbench: an open source platform for integrative genomics
abstract
SUMMARY: geWorkbench (genomics Workbench) is an open source Java desktop application that provides access to an integrated suite of tools for the analysis and visualization of data from a wide range of genomics domains (gene expression, sequence, protein structure and systems biology). More than 70 distinct plug-in modules are currently available implementing both classical analyses (several variants of clustering, classification, homology detection, etc.) as well as state of the art algorithms for the reverse engineering of regulatory networks and for protein structure prediction, among many others. geWorkbench leverages standards-based middleware technologies to provide seamless access to remote data, annotation and computational servers, thus, enabling researchers with limited local resources to benefit from available public infrastructure. AVAILABILITY: The project site (http://www.geworkbench.org) includes links to self-extracting installers for most operating system (OS) platforms as well as instructions for building the application from scratch using the source code [which is freely available from the project's SVN (subversion) repository]. geWorkbench support is available through the end-user and developer forums of the caBIG Molecular Analysis Tools Knowledge Center, https://cabig-kc.nci.nih.gov/Molecular/forums/
Aris Floratos, Kenneth Smith 0001, Zhou Ji, John Watkinson, Andrea Califano
Bioinform.5
2009 New JBI emphasis on translational bioinformatics
Edward H. Shortliffe, Andrea Califano, Lawrence Hunter
J. Biomed. Informatics2
2007 ChIP-on-chip significance analysis reveals ubiquitous transcription factor binding
abstract
ChIP-on-chip technology provides a genome-scale view of transcription factor (TF)/target interactions and a systems-level window into transcriptional regulatory networks. However, while many studies have used ChIP-on-chip data to effectively discover new TF targets, statistical methods have fallen short of developing an accurate model to disassociate signals caused by experimental noise from those caused by true biological variation, thus leveraging the technology to provide high confidence predictions of the full range of interactions. This paper presents a novel method to accurately model the significance of binding events measured by ChIP-on-chip data. For each arrayed probe representing a genomic segment, a ChIP-on-chip microarray measures intensity levels for the IP channel, which is enriched in genomic fragments bound by an immunoprecipitated TF, and the WCE channel, which represents random genomic fragments. Statistical significance is inferred by computing the conditional probability, p ( M | A ), where M = log 2 ( I P W C E ) MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGnbqtcqGH9aqpcyGGSbaBcqGGVbWBcqGGNbWzcqaIYaGmdaqadaqaamaalaaabaGaemysaKKaemiuaafabaGaem4vaCLaem4qamKaemyraueaaaGaayjkaiaawMcaaaaa@3B1B@ and A = log 2 ( I P ) + log 2 ( W C E ) 2 MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacqWGbbqqcqGH9aqpdaWcaaqaaiGbcYgaSjabc+gaVjabcEgaNjabikdaYiabcIcaOiabdMeajjabdcfaqjabcMcaPiabgUcaRiGbcYgaSjabc+gaVjabcEgaNjabikdaYiabcIcaOiabdEfaxjabdoeadjabdweafjabcMcaPaqaaiabikdaYaaaaaa@43C2@ (Fig. 1 ). A kernel density estimation procedure is used to calculate the joint probability, p ( M , A ), and for each average intensity value, the mean of the null distribution (i.e. distribution for unbound probes) is inferred as M ^ A = arg max M p ( M | A ) MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaOGaeyypa0ZaaCbeaeaacyGGHbqycqGGYbGCcqGGNbWzcyGGTbqBcqGGHbqycqGG4baEaSqaaiabd2eanbqabaGccqWGWbaCcqGGOaakcqWGnbqtcqGG8baFcqWGbbqqcqGGPaqkaaa@4089@ . The distribution of p ( M | A ), for M < M ^ A MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaaaa@2F16@ , is then projected across M ^ A MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaacuWGnbqtgaqcamaaBaaaleaacqWGbbqqaeqaaaaa@2F16@ to yield the inferred null distribution, which is used to assign statistical significance scores. Probes for replicate experiments and probes with genomic locations within the fragmentation length (~500 bp) are integrated to produce a single significance score for each genomic region. (Left) Magnitude versus amplitude (MA) plot of a ChIP-on-chip hybridization. The x-axis represents the average log2 intensity of the IP and WCE channels, and the y-axis represents the log2 ratio of IP/WCE. The black line represents the mean of the inferred null distribution, and the colored lines represent confidence intervals of .1, .01, and .001 probability. The model reveals an intensity dependent mean and variance of the null distribution, and a large number of probes are significantly enriched in the IP channel. (Right) The axes are the same as in the left panel, and colors represent the -log10 p-value of the null distribution. The method is tested on six different ChIP-on-chip arrays representing replicate experiments for three different TFs (NOTCH1, MYC and HES1). For each experiment, this analysis reveals an order of magnitude more genomic binding events than detected by traditional methods, predicting several thousand interactions for each TF and suggesting previously unappreciated complexity of transcriptional regulatory networks. Several independent experiments are used to provide evidence about the validity of these predictions. First, biochemical validation of more than 20 predicted targets by gene specific ChIP and qPCR confirm the accuracy of false discovery rate statistics computed by the method. Second, binding site enrichment analysis indicates that the strength of binding site signals are maintained over several thousand promoters. Finally, gene expression analysis reveals a coordinated downregulation of gene expression for the entire range of predicted NOTCH1 bound genes upon NOTCH1 inhibition experiments in cell lines, indicating that a large percentage of bound genes are also functionally regulated by NOTCH1.
Adam A. Margolin, Teresa Palomero, Adolfo A. Ferrando, Andrea Califano, Gustavo Stolovitzky
BMC Bioinform.4
2007 Hui-Huang Hsu, Advanced Data Mining Technologies in Bioinformatics, Editors: Idea Group Publishing, pp 239
Andrea Califano
J. Biomed. Informatics1
2006 Genome-Wide Discovery of Modulators of Transcriptional Interactions in Human B Lymphocytes
Kai Wang 0045, Ilya Nemenman, Nilanjana Banerjee, Adam A. Margolin, Andrea Califano
RECOMB5
2006 ARACNE: An Algorithm for the Reconstruction of Gene Regulatory Networks in a Mammalian Cellular Context
abstract
BACKGROUND: Elucidating gene regulatory networks is crucial for understanding normal cell physiology and complex pathologic phenotypes. Existing computational methods for the genome-wide "reverse engineering" of such networks have been successful only for lower eukaryotes with simple genomes. Here we present ARACNE, a novel algorithm, using microarray expression profiles, specifically designed to scale up to the complexity of regulatory networks in mammalian cells, yet general enough to address a wider range of network deconvolution problems. This method uses an information theoretic approach to eliminate the majority of indirect interactions inferred by co-expression methods. RESULTS: We prove that ARACNE reconstructs the network exactly (asymptotically) if the effect of loops in the network topology is negligible, and we show that the algorithm works well in practice, even in the presence of numerous loops and complex topologies. We assess ARACNE's ability to reconstruct transcriptional regulatory networks using both a realistic synthetic dataset and a microarray dataset from human B cells. On synthetic datasets ARACNE achieves very low error rates and outperforms established methods, such as Relevance Networks and Bayesian Networks. Application to the deconvolution of genetic networks in human B cells demonstrates ARACNE's ability to infer validated transcriptional targets of the cMYC proto-oncogene. We also study the effects of misestimation of mutual information on network reconstruction, and show that algorithms based on mutual information ranking are more resilient to estimation errors. CONCLUSION: ARACNE shows promise in identifying direct transcriptional interactions in mammalian cellular networks, a problem that has challenged existing reverse engineering algorithms. This approach should enhance our ability to use microarray data to elucidate functional mechanisms that underlie cellular processes and to identify molecular targets of pharmacological compounds in mammalian cellular networks.
Adam A. Margolin, Ilya Nemenman, Katia Basso, Chris Wiggins 0001, Gustavo Stolovitzky, Riccardo Dalla Favera, Andrea Califano
BMC Bioinform.7
2000 Analysis of Gene Expression Microarrays for Phenotype Classification
Andrea Califano, Gustavo Stolovitzky, Yuhai Tu
ISMB1
2000 Systematic and automated discovery of patterns in PROSITE families
abstract
PROSITE is a method for protein classification which relies on a database of biologically significant sites and patterns in protein sequences. Most patterns in PROSITE have been gathered by a a labor intensive combination of experimental characterization of functional residues and sequence alignment. In this paper we present a new and efficient supervised learning procedure, based on the Splash deterministic pattern discovery algorithm and on a framework to assess the statistical significance of patterns. We demonstrate its application to the fully automatic discovery of patterns in 974 PROSITE families. For these families, Splash generates patterns with better specificity and/or sensitivity in 28%, identical statistics in 48%, and worse statistics in 15% of the cases; for the remaining families, patterns exhibited mixed behavior. Second, we have characterized the amount of overlap, on the sequences, between newly discovered patterns and those in PROSITE. In about 75% of the cases, Splash patterns identify sequence sites that overlap more than 50% with those reported in PROSITE. Of the 272 patterns which perform strictly better than the corresponding PROSITE pattern, 178 show more than 70% overlap with the PROSITE pattern. Third, our results suggest that the statistical significance of discovered patterns correlates well with their biological significance. Finally, we use the trypsin subfamily of serine proteases to illustrate the use of this method to exhaustively discover all motifs in a family that are statistically and biologically significant. The complete analysis is sufficiently rapid, taking less than a day for all PROSITE families, to enable the use this methodology for routine curation of existing motif and profile databases.
Reece K. Hart, Ajay K. Royyuru, Gustavo Stolovitzky, Andrea Califano
RECOMB4
2000 SPLASH: structural pattern localization analysis by sequential histograms
abstract
MOTIVATION: The discovery of sparse amino acid patterns that match repeatedly in a set of protein sequences is an important problem in computational biology. Statistically significant patterns, that is patterns that occur more frequently than expected, may identify regions that have been preserved by evolution and which may therefore play a key functional or structural role. Sparseness can be important because a handful of non-contiguous residues may play a key role, while others, in between, may be changed without significant loss of function or structure. Similar arguments may be applied to conserved DNA patterns. Available sparse pattern discovery algorithms are either inefficient or impose limitations on the type of patterns that can be discovered. RESULTS: This paper introduces a deterministic pattern discovery algorithm, called Splash, which can find sparse amino or nucleic acid patterns matching identically or similarly in a set of protein or DNA sequences. Sparse patterns of any length, up to the size of the input sequence, can be discovered without significant loss in performances. Splash is extremely efficient and embarrassingly parallel by nature. Large databases, such as a complete genome or the non-redundant SWISS-PROT database can be processed in a few hours on a typical workstation. Alternatively, a protein family or superfamily, with low overall homology, can be analyzed to discover common functional or structural signatures. Some examples of biologically interesting motifs discovered by Splash are reported for the histone I and for the G-Protein Coupled Receptor families. Due to its efficiency, Splash can be used to systematically and exhaustively identify conserved regions in protein family sets. These can then be used to build accurate and sensitive PSSM or HMM models for sequence analysis. AVAILABILITY: Splash is available to non-commercial research centers upon request, conditional on the signing of a test field agreement. CONTACT: [email protected], Splash main page http://www.research.ibm.com/splash
Andrea Califano
Bioinform.1
1998 A High-Dimensional Indexing Scheme for Scalable Fingerprint-Based Identification
Andrea Califano, Bob Germain, Scott Colville
ACCV (1)1
1996 Data- and Model-Driven Multiresolution Processing
Andrea Califano, Rick Kjeldsen, Ruud M. Bolle
Comput. Vis. Image Underst.1
1994 Multidimensional Indexing for Recognizing Visual Shapes
abstract
This paper introduces an analytical framework for studying some properties of model acquisition and recognition techniques based on indexing. The goal is to demonstrate that several problems previously associated with the approach can be attributed to the low dimensionality of invariants used. These include limited index selectivity. Excessive accumulation of votes in the look-up table buckets, and excessive sensitivity to quantization parameters. Theoretical results demonstrate that using high-dimensional, highly descriptive global invariants produces better results in terms of accuracy, false positive suppression, and computation time. A practical example of high-dimensional global invariants is introduced and used to implement a 2-D shape acquisition/recognition system. The acquisition/recognition system is based on a two-step table look-up mechanism. First, local curve descriptors are obtained by correlating image contour information at short range. Then, seven-dimensional global invariants are computed by correlating triplets of local curve descriptors at longer range. This experimental system is meant to illustrate the behavior of a high-dimensional indexing scheme. Indeed, its performance shows good agreement with the analytical model with respect to database size, fault tolerance, and recognition speed. Model acquisition time is linear to cubic in the number of object features. Object recognition time is constant to linear in the number of models in the database and linear to cubic in the number of features in the image. The system has been tested extensively. With more than 250 arbitrary shapes in the database. Unsupervised shape and subpart acquisition is demonstrated.>
Andrea Califano, Rakesh Mohan
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Systematic design of indexing strategies for object recognition
abstract
The authors analyze how parameters of indexing based recognition systems affects their performance. The main result is that increasing the dimensionality of the indices leads to significantly improved discrimination, false positive suppression and reduced recognition times. With increase in index dimensionality, coarser quantization is required to allow index match in the presence of noise. The authors' analysis also allows estimation of votes thresholds for recognition and estimation of the amount of occlusion that can be tolerated by indexing schemes for given levels of recognition confidence.>
Andrea Califano, Rakesh Mohan
CVPR1
1993 FLASH: a fast look-up algorithm for string homology
abstract
A key issue in managing large amounts of data is the availability of efficient, accurate, ad selective techniques to detect homology (similarity) between newly recovered and previously acquired sequences. The algorithm presented is based on a probabilistic indexing framework which requires minimal access to the database for each match. A highly redundant number of descriptive tuples from the sequences of interest are generated and used as indices in a table look-up paradigm. Theoretical and experimental results on the sensitivity and accuracy of the approach are provided. These include the probability of correct and random matches and the storage and computational requirements. An experimental system is implemented for a database containing the complete genome of the bacteria E. Coli (approximately 2 million nucleotides). Search time is a few seconds on a workstation class machine. The algorithm is shown to scale well to databases containing billions of nucleotides with performances that are orders of magnitude better than the fastest of the current techniques.>
Andrea Califano, Isidore Rigoutsos
CVPR1
1993 FLASH: A Fast Look-Up Algorithm for String Homology
Andrea Califano, Isidore Rigoutsos
ISMB1
1992 A Complete and Extendable Approach to Visual Recognition
abstract
A framework for 3D object recognition is presented. Its flexibility and extensibility are accomplished through a uniform, parallel, and modular recognition architecture. Concurrent and stacked parameter transforms reconstruct a variety of features from the input scene. At each stage, constraint satisfaction networks collect and fuse the evidence obtained through the parameter transforms, ensuring a globally consistent interpretation of the input scene and allowing for the integration of diverse types of information. The final interpretation of the scene is a small consistent subset of the many initial hypotheses about partial features, primitive features, feature assemblies, and 3D objects computed by the various parameter transforms. A complete, integrated, and implemented system that extracts planar surfaces, patches of quadrics of revolution, and planar intersection curves of these surfaces from a depth map viewing 3D objects is described. Experimental results on the recognition behavior of the system are presented.>
Ruud M. Bolle, Andrea Califano, Rick Kjeldsen
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 The Multiple Window Parameter Transform
abstract
The multiwindow transform, an extension of parameter transform techniques that increase performance and scope by exploiting the long-range correlated information contained in multiple portions of an image, is presented. Multiple-window transforms allow the extraction of high-dimensional features with improvement in accuracy over conventional techniques while keeping linear to low-order-polynomial computational and space requirements with respect to image size and dimensionality of the features. Using correlated information provides a direct link between extracted features and supporting regions in the image. This, coupled with evidence integration techniques, is used to suppress noisy or nonexistent feature hypotheses. Parameter spaces are implemented as constraint satisfaction networks, where feature hypotheses with overlapping support in the image compete. After an iterative relaxation phase, surviving hypotheses have disjoint support, forming a segmentation of the image. Examples show the performance and provide insight about the behavior.>
Andrea Califano, Ruud M. Bolle
IEEE Trans. Pattern Anal. Mach. Intell.1
1991 Multidimensional indexing for recognizing visual shapes
abstract
A homogeneous approach for acquisition, storage, and recognition of nonparametric shapes from images, using a novel shape representation based on shape autocorrelation operators is presented. A theoretical and experimental analysis of the computational complexity, recognition performance with increasing database size, and fault tolerance of the approach is presented. The system has been tested extensively with more than 300 arbitrary shapes in the database. Using a set of complex shapes, the recognition behavior with respect to occlusion, geometric transformation, and cluttered environments is studied. Unsupervised shape and subpart acquisition is demonstrated.>
Andrea Califano, Rakesh Mohan
CVPR1
1990 Generalized Shape Autocorrelation
Andrea Califano, Rakesh Mohan
AAAI1
1990 Active 3D object models
abstract
A novel approach is presented for pruning the amount of search needed to match image features to object models. The technique relies on active networks which capture various visibility and geometric constraints between features of a model to prune these features from search space during matching. The networks, which can be efficiently implemented in Boolean logic, integrate harmoniously with the previous work in feature recognition and object matching. A method is proposed for clustering model features (vsets) and four types of constraints which assist in building the networks. The authors show, both analytically and empirically, the dramatic reduction in search provided by activation nets.>
Ruud M. Bolle, Andrea Califano, Rick Kjeldsen, Rakesh Mohan
ICCV2
1990 Data and model driven foveation
abstract
A general framework for multiresolution visual recognition is introduced. The input is processed simultaneously at a coarse resolution throughout the image and at finer resolution within a small window. An approach for controlling the movement of the high-resolution window is described which allows for the unification of a variety of data and model-driven behavioral paradigms. Three modes have been implemented, one based on large unexplained areas in the data, one on conflicts in the object-model database, and one on a 2D space-filling algorithm. It is argued that this kind of multiresolution processing not only is useful in limiting the computational time, but also can be a deciding factor in making the entire vision problem a tractable and stable one. To demonstrate the approach, a class of 3D surface textures is introduced as a feature for recognition in the system considered. Surface texture recognition typically requires higher-resolution processing than required for the extraction of the underlying surface. As an example, surface texture is used to discriminate between a ping-pong ball and a golf ball.>
Andrea Califano, Rick Kjeldsen, Ruud M. Bolle
ICPR (1)1
1989 Visual recognition using concurrent and layered parameter networks
abstract
A vision system to recognize 3-D objects is presented. A novel notion of generalized feature makes it possible to develop a homogeneous architecture to support recognition from simple partial features to complex feature assemblies and 3-D objects. Layered concurrent parameter transforms vote for feature hypotheses on the basis of image data and previously reconstructed features. Recognition networks, motivated by connectionist systems, collect votes, fuse evidence from various sources and ensure global consistency. In addition, they provide an integrated solution to the segmentation problem. The highly modular design allows fundamentally different types of features to interact in a coherent way, utilizing redundancies for more robust recognition. Within this paradigm, a system extracts planar patches, patches of quadrics of revolution, and the intersection curves of these surfaces (lines and conic sections in three-space) from a depth map. Reconstructed features index into a model database to form consistent object hypotheses. Experimental results detailing the recognition behavior or real depth maps are included.>
Ruud M. Bolle, Andrea Califano, Rick Kjeldsen, Russell W. Taylor
CVPR2
1989 Generalized neighborhoods: a new approach to complex parameter feature extraction
abstract
A generalized neighborhood concept is presented which extends the usual techniques for feature extraction using parameter transforms. Generalized neighborhoods allow operators to use the joint information contained in distant portions of the same feature; i.e. to utilize the long-distance correlation present in the image. The generalized neighborhood techniques, by correlating local information over different portions of the image, produce up to two orders of magnitude improvement in accuracy over conventional techniques. The response also becomes more complicated; false features may be detected due to a peculiar form of correlated noise. A general framework, motivated by connectionist networks, is presented which eliminates this behaviour by introducing competitive processes in the parameter spaces. A novel approach to the generation of lateral inhibition links in the networks is proposed which is consistent with generalized neighborhoods. Experiments are provided that show results on range data. Complex surfaces and 3-D surface-intersection curves are reconstructed from the data.>
Andrea Califano, Ruud M. Bolle, Russell W. Taylor
CVPR1
1989 A Homogeneous Framework for Visual Recognition
Rick Kjeldsen, Ruud M. Bolle, Andrea Califano, Russell W. Taylor
IJCAI3
1988 Feature Recognition Using Correlated Information Contained in Multiple Neighborboods
Andrea Califano
AAAI1