Emmanuel Barillot

dblp:22/5886 · DBLP profile ↗
← Back
55ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-2724-2002ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 51 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 AstroLogics: a simulation-based framework for the analysis of boolean model ensembles
abstract
MOTIVATION: Boolean networks (BNs) have emerged as versatile tools for modeling cellular regulatory mechanisms due to their ability to capture key biological features despite their simplicity. Multiple BN synthesis methods have emerged in recent decades, aiming to infer BNs with dynamics that correspond to experimental data. Often, these methods generate multiple BN candidates or "model ensembles". While these ensembles are valuable for representing cell populations and their heterogeneity, they are typically treated as a single component without examining their constituent features. RESULTS: We present AstroLogics, a novel framework designed to analyze and identify differences in both dynamical behavior and logical regulation within a BN model ensemble. The framework calculates dynamical distances between BNs through exploration of their state transition graphs (STGs), enabling clustering of similarly functioning models that may represent different cellular fates or signaling mechanisms. AstroLogics also identifies key logical properties that govern each cluster, highlighting the core regulatory structures that differentiate model behaviors. Our approach leverages MaBoSS, a stochastic simulation tool that implements the Boolean Kinetic Monte-Carlo algorithm to address time interpretation in BNs. This probabilistic estimation method allows efficient probing of BN dynamics through stochastic simulations, overcoming the computational limitations of exhaustive STG analysis. Our framework also provides powerful visualization and classification of the BN ensemble. Through multiple use cases, we demonstrate how AstroLogics facilitates comprehensive analyses of model diversity and discovery of key regulatory structures within a BN ensemble. AVAILABILITY AND IMPLEMENTATION: The AstroLogics package, along with tutorials and datasets, are available at https://github.com/sysbio-curie/AstroLogics.
Saran Pankaew, Vincent Noel, Loïc Paulevé, Denis Thieffry, Emmanuel Barillot, Laurence Calzone
Bioinform.5
2025 NeKo: A tool for automatic network construction from prior knowledge
abstract
Biological networks provide a structured framework for analyzing the dynamic interplay and interactions between molecular entities, facilitating deeper insights into cellular functions and biological processes. Network construction often requires extensive manual curation based on scientific literature and public databases, a time-consuming and laborious task. To address this challenge, we introduce NeKo, a Python package to automate the construction of biological networks by integrating and prioritizing molecular interactions from various databases. NeKo allows users to provide their molecules of interest (e.g., genes, proteins or phosphosites), select interaction resources and apply flexible strategies to build networks based on prior knowledge. Users can filter interactions by various criteria, such as direct or indirect links and signed or unsigned interactions, to tailor the network to their needs and downstream analysis. We demonstrate some of NeKo's capabilities in two use cases: first we construct a network based on transcriptomics from medulloblastoma; in the second, we model drug synergies. NeKo streamlines the network-building process, making it more accessible and efficient for researchers.
Marco Ruscone, Eirini Tsirvouli, Andrea Checcoli, Dénes Türei, Emmanuel Barillot, Julio Saez-Rodriguez, Loredana Martignetti, Åsmund Flobak, Laurence Calzone
PLoS Comput. Biol.5
2024 Building multiscale models with PhysiBoSS, an agent-based modeling tool
abstract
Multiscale models provide a unique tool for analyzing complex processes that study events occurring at different scales across space and time. In the context of biological systems, such models can simulate mechanisms happening at the intracellular level such as signaling, and at the extracellular level where cells communicate and coordinate with other cells. These models aim to understand the impact of genetic or environmental deregulation observed in complex diseases, describe the interplay between a pathological tissue and the immune system, and suggest strategies to revert the diseased phenotypes. The construction of these multiscale models remains a very complex task, including the choice of the components to consider, the level of details of the processes to simulate, or the fitting of the parameters to the data. One additional difficulty is the expert knowledge needed to program these models in languages such as C++ or Python, which may discourage the participation of non-experts. Simplifying this process through structured description formalisms-coupled with a graphical interface-is crucial in making modeling more accessible to the broader scientific community, as well as streamlining the process for advanced users. This article introduces three examples of multiscale models which rely on the framework PhysiBoSS, an add-on of PhysiCell that includes intracellular descriptions as continuous time Boolean models to the agent-based approach. The article demonstrates how to construct these models more easily, relying on PhysiCell Studio, the PhysiCell Graphical User Interface. A step-by-step tutorial is provided as Supplementary Material and all models are provided at https://physiboss.github.io/tutorial/.
Marco Ruscone, Andrea Checcoli, Randy W. Heiland, Emmanuel Barillot, Paul Macklin, Laurence Calzone, Vincent Noel
Briefings Bioinform.4
2024 Maboss for HPC environments: implementations of the continuous time Boolean model simulator for large CPU clusters and GPU accelerators
abstract
BACKGROUND: Computational models in systems biology are becoming more important with the advancement of experimental techniques to query the mechanistic details responsible for leading to phenotypes of interest. In particular, Boolean models are well fit to describe the complexity of signaling networks while being simple enough to scale to a very large number of components. With the advance of Boolean model inference techniques, the field is transforming from an artisanal way of building models of moderate size to a more automatized one, leading to very large models. In this context, adapting the simulation software for such increases in complexity is crucial. RESULTS: We present two new developments in the continuous time Boolean simulators: MaBoSS.MPI, a parallel implementation of MaBoSS which can exploit the computational power of very large CPU clusters, and MaBoSS.GPU, which can use GPU accelerators to perform these simulations. CONCLUSION: These implementations enable simulation and exploration of the behavior of very large models, thus becoming a valuable analysis tool for the systems biology community.
Adam Smelko, Miroslav Kratochvíl, Emmanuel Barillot, Vincent Noel
BMC Bioinform.3
2023 Multiscale model of the different modes of cancer cell invasion
abstract
MOTIVATION: Mathematical models of biological processes altered in cancer are built using the knowledge of complex networks of signaling pathways, detailing the molecular regulations inside different cell types, such as tumor cells, immune and other stromal cells. If these models mainly focus on intracellular information, they often omit a description of the spatial organization among cells and their interactions, and with the tumoral microenvironment. RESULTS: We present here a model of tumor cell invasion simulated with PhysiBoSS, a multiscale framework, which combines agent-based modeling and continuous time Markov processes applied on Boolean network models. With this model, we aim to study the different modes of cell migration and to predict means to block it by considering not only spatial information obtained from the agent-based simulation but also intracellular regulation obtained from the Boolean model. Our multiscale model integrates the impact of gene mutations with the perturbation of the environmental conditions and allows the visualization of the results with 2D and 3D representations. The model successfully reproduces single and collective migration processes and is validated on published experiments on cell invasion. In silico experiments are suggested to search for possible targets that can block the more invasive tumoral phenotypes. AVAILABILITY AND IMPLEMENTATION: https://github.com/sysbio-curie/Invasion_model_PhysiBoSS.
Marco Ruscone, Arnau Montagud, Philippe Chavrier, Olivier Destaing, Isabelle Bonnet, Andrei Yu. Zinovyev, Emmanuel Barillot, Vincent Noel, Laurence Calzone
Bioinform.7
2022 BIODICA: a computational environment for Independent Component Analysis of omics data
abstract
SUMMARY: We developed BIODICA, an integrated computational environment for application of independent component analysis (ICA) to bulk and single-cell molecular profiles, interpretation of the results in terms of biological functions and correlation with metadata. The computational core is the novel Python package stabilized-ica which provides interface to several ICA algorithms, a stabilization procedure, meta-analysis and component interpretation tools. BIODICA is equipped with a user-friendly graphical user interface, allowing non-experienced users to perform the ICA-based omics data analysis. The results are provided in interactive ways, thus facilitating communication with biology experts. AVAILABILITY AND IMPLEMENTATION: BIODICA is implemented in Java, Python and JavaScript. The source code is freely available on GitHub under the MIT and the GNU LGPL licenses. BIODICA is supported on all major operating systems. URL: https://sysbio-curie.github.io/biodica-environment/.
Nicolas Captier, Jane Merlevede, Askhat Molkenov, Ainur Seisenova, Altynbek Zhubanchaliyev, Petr V. Nazarov, Emmanuel Barillot, Ulykbek Kairov, Andrei Yu. Zinovyev
Bioinform.7
2021 Personalized logical models to investigate cancer response to BRAF treatments in melanomas and colorectal cancers
abstract
The study of response to cancer treatments has benefited greatly from the contribution of different omics data but their interpretation is sometimes difficult. Some mathematical models based on prior biological knowledge of signaling pathways facilitate this interpretation but often require fitting of their parameters using perturbation data. We propose a more qualitative mechanistic approach, based on logical formalism and on the sole mapping and interpretation of omics data, and able to recover differences in sensitivity to gene inhibition without model training. This approach is showcased by the study of BRAF inhibition in patients with melanomas and colorectal cancers who experience significant differences in sensitivity despite similar omics profiles. We first gather information from literature and build a logical model summarizing the regulatory network of the mitogen-activated protein kinase (MAPK) pathway surrounding BRAF, with factors involved in the BRAF inhibition resistance mechanisms. The relevance of this model is verified by automatically assessing that it qualitatively reproduces response or resistance behaviors identified in the literature. Data from over 100 melanoma and colorectal cancer cell lines are then used to validate the model's ability to explain differences in sensitivity. This generic model is transformed into personalized cell line-specific logical models by integrating the omics information of the cell lines as constraints of the model. The use of mutations alone allows personalized models to correlate significantly with experimental sensitivities to BRAF inhibition, both from drug and CRISPR targeting, and even better with the joint use of mutations and RNA, supporting multi-omics mechanistic models. A comparison of these untrained models with learning approaches highlights similarities in interpretation and complementarity depending on the size of the datasets. This parsimonious pipeline, which can easily be extended to other biological questions, makes it possible to explore the mechanistic causes of the response to treatment, on an individualized basis.
Jonas Béal, Lorenzo Pantolini, Vincent Noel, Emmanuel Barillot, Laurence Calzone
PLoS Comput. Biol.4
2020 cd2sbgnml: bidirectional conversion between CellDesigner and SBGN formats
abstract
MOTIVATION: CellDesigner is a well-established biological map editor used in many large-scale scientific efforts. However, the interoperability between the Systems Biology Graphical Notation (SBGN) Markup Language (SBGN-ML) and the CellDesigner's proprietary Systems Biology Markup Language (SBML) extension formats remains a challenge due to the proprietary extensions used in CellDesigner files. RESULTS: We introduce a library named cd2sbgnml and an associated web service for bidirectional conversion between CellDesigner's proprietary SBML extension and SBGN-ML formats. We discuss the functionality of the cd2sbgnml converter, which was successfully used for the translation of comprehensive large-scale diagrams such as the RECON Human Metabolic network and the complete Atlas of Cancer Signalling Network, from the CellDesigner file format into SBGN-ML. AVAILABILITY AND IMPLEMENTATION: The cd2sbgnml conversion library and the web service were developed in Java, and distributed under the GNU Lesser General Public License v3.0. The sources along with a set of examples are available on GitHub (https://github.com/sbgn/cd2sbgnml and https://github.com/sbgn/cd2sbgnml-webservice, respectively). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Irina Balaur, Ludovic Roy, Alexander Mazein, S. Gökberk Karaca, Ugur Dogrusoz, Emmanuel Barillot, Andrei Yu. Zinovyev
Bioinform.6
2020 cd2sbgnml: bidirectional conversion between CellDesigner and SBGN formats
abstract
Bioinformatics (2019) doi: 10.1093/bioinformatics/btz969. The following funding acknowledgement was omitted from the above article: This work was also supported by the Innovative Medicines Initiative Joint Undertaking under grant agreement no. IMI 115446 (eTRIKS) to Charles Auffray and Rudi Balling, resources of which are composed of financial contributions from the European Union’s Seventh Framework Programme (2007-2013) and EFPIA companies. This has now been corrected.
Irina Balaur, Ludovic Roy, Alexander Mazein, S. Gökberk Karaca, Ugur Dogrusoz, Emmanuel Barillot, Andrei Yu. Zinovyev
Bioinform.6
2020 Exact solving and sensitivity analysis of stochastic continuous time Boolean models
abstract
BACKGROUND: Solutions to stochastic Boolean models are usually estimated by Monte Carlo simulations, but as the state space of these models can be enormous, there is an inherent uncertainty about the accuracy of Monte Carlo estimates and whether simulations have reached all attractors. Moreover, these models have timescale parameters (transition rates) that the probability values of stationary solutions depend on in complex ways, raising the necessity of parameter sensitivity analysis. We address these two issues by an exact calculation method for this class of models. RESULTS: We show that the stationary probability values of the attractors of stochastic (asynchronous) continuous time Boolean models can be exactly calculated. The calculation does not require Monte Carlo simulations, instead it uses graph theoretical and matrix calculation methods previously applied in the context of chemical kinetics. In this version of the asynchronous updating framework the states of a logical model define a continuous time Markov chain and for a given initial condition the stationary solution is fully defined by the right and left nullspace of the master equation's kinetic matrix. We use topological sorting of the state transition graph and the dependencies between the nullspaces and the kinetic matrix to derive the stationary solution without simulations. We apply this calculation to several published Boolean models to analyze the under-explored question of the effect of transition rates on the stationary solutions and show they can be sensitive to parameter changes. The analysis distinguishes processes robust or, alternatively, sensitive to parameter values, providing both methodological and biological insights. CONCLUSION: Up to an intermediate size (the biggest model analyzed is 23 nodes) stochastic Boolean models can be efficiently solved by an exact matrix method, without using Monte Carlo simulations. Sensitivity analysis with respect to the model's timescale parameters often reveals a small subset of all parameters that primarily determine the stationary probability of attractor states.
Mihály Koltai, Vincent Noel, Andrei Yu. Zinovyev, Laurence Calzone, Emmanuel Barillot
BMC Bioinform.5
2020 Collective intelligence defines biological functions in Wikipedia as communities in the hidden protein connection network
abstract
English Wikipedia, containing more than five millions articles, has approximately eleven thousands web pages devoted to proteins or genes most of which were generated by the Gene Wiki project. These pages contain information about interactions between proteins and their functional relationships. At the same time, they are interconnected with other Wikipedia pages describing biological functions, diseases, drugs and other topics curated by independent, not coordinated collective efforts. Therefore, Wikipedia contains a directed network of protein functional relations or physical interactions embedded into the global network of the encyclopedia terms, which defines hidden (indirect) functional proximity between proteins. We applied the recently developed reduced Google Matrix (REGOMAX) algorithm in order to extract the network of hidden functional connections between proteins in Wikipedia. In this network we discovered tight communities which reflect areas of interest in molecular biology or medicine and can be considered as definitions of biological functions shaped by collective intelligence. Moreover, by comparing two snapshots of Wikipedia graph (from years 2013 and 2017), we studied the evolution of the network of direct and hidden protein connections. We concluded that the hidden connections are more dynamic compared to the direct ones and that the size of the hidden interaction communities grows with time. We recapitulate the results of Wikipedia protein community analysis and annotation in the form of an interactive online map, which can serve as a portal to the Gene Wiki project.
Andrei Yu. Zinovyev, Urszula Czerwinska, Laura Cantini, Emmanuel Barillot, Klaus M. Frahm, Dima Shepelyansky
PLoS Comput. Biol.4
2019 Application of Atlas of Cancer Signalling Network in preclinical studies
abstract
Cancer initiation and progression are associated with multiple molecular mechanisms. The knowledge of these mechanisms is expanding and should be converted into guidelines for tackling the disease. Here, we discuss the formalization of biological knowledge into a comprehensive resource: the Atlas of Cancer Signalling Network (ACSN) and the Google Maps-based tool NaviCell, which supports map navigation. The application of ACSN for omics data visualization, in the context of signalling maps, is possible via the NaviCell Web Service module and through the NaviCom tool. It allows generation of network-based molecular portraits of cancer using multilevel omics data. We review how these resources and tools are applied for cancer preclinical studies. Structural analysis of the maps together with omics data helps to rationalize the synergistic effects of drugs and allows design of complex disease stage-specific druggable interventions. The use of ACSN modules and maps as signatures of biological functions can help in cancer data analysis and interpretation. In addition, they empowered finding of associations between perturbations in particular molecular mechanisms and the risk to develop a specific type of cancer. These approaches are helpful, among others, to study the interplay between molecular mechanisms of cancer. It opens an opportunity to decipher how gene interactions govern the hallmarks of cancer in specific contexts. We discuss a perspective to develop a flexible methodology and a pipeline to enable systematic omics data analysis in the context of signalling network maps, for stratifying patients and suggesting interventions points and drug repositioning in cancer and other diseases.
L. Cristobal Monraz Gomez, Maria Kondratova, Jean-Marie Ravel, Emmanuel Barillot, Andrei Yu. Zinovyev, Inna Kuperstein
Briefings Bioinform.4
2019 Conceptual and computational framework for logical modelling of biological networks deregulated in diseases
abstract
Mathematical models can serve as a tool to formalize biological knowledge from diverse sources, to investigate biological questions in a formal way, to test experimental hypotheses, to predict the effect of perturbations and to identify underlying mechanisms. We present a pipeline of computational tools that performs a series of analyses to explore a logical model's properties. A logical model of initiation of the metastatic process in cancer is used as a transversal example. We start by analysing the structure of the interaction network constructed from the literature or existing databases. Next, we show how to translate this network into a mathematical object, specifically a logical model, and how robustness analyses can be applied to it. We explore the visualization of the stable states, defined as specific attractors of the model, and match them to cellular fates or biological read-outs. With the different tools we present here, we explain how to assign to each solution of the model a probability and how to identify genetic interactions using mutant phenotype probabilities. Finally, we connect the model to relevant experimental data: we present how some data analyses can direct the construction of the network, and how the solutions of a mathematical model can also be compared with experimental data, with a particular focus on high-throughput data in cancer biology. A step-by-step tutorial is provided as a Supplementary Material and all models, tools and scripts are provided on an accompanying website: https://github.com/sysbio-curie/Logical_modelling_pipeline.
Arnau Montagud, Pauline Traynard, Loredana Martignetti, Eric Bonnet, Emmanuel Barillot, Andrei Yu. Zinovyev, Laurence Calzone
Briefings Bioinform.5
2019 Community-driven roadmap for integrated disease maps
abstract
The Disease Maps Project builds on a network of scientific and clinical groups that exchange best practices, share information and develop systems biomedicine tools. The project aims for an integrated, highly curated and user-friendly platform for disease-related knowledge. The primary focus of disease maps is on interconnected signaling, metabolic and gene regulatory network pathways represented in standard formats. The involvement of domain experts ensures that the key disease hallmarks are covered and relevant, up-to-date knowledge is adequately represented. Expert-curated and computer readable, disease maps may serve as a compendium of knowledge, allow for data-supported hypothesis generation or serve as a scaffold for the generation of predictive mathematical models. This article summarizes the 2nd Disease Maps Community meeting, highlighting its important topics and outcomes. We outline milestones on the roadmap for the future development of disease maps, including creating and maintaining standardized disease maps; sharing parts of maps that encode common human disease mechanisms; providing technical solutions for complexity management of maps; and Web tools for in-depth exploration of such maps. A dedicated discussion was focused on mathematical modeling approaches, as one of the main goals of disease map development is the generation of mathematically interpretable representations to predict disease comorbidity or drug response and to suggest drug repositioning, altogether supporting clinical decisions.
Marek Ostaszewski, Stephan Gebel, Inna Kuperstein, Alexander Mazein, Andrei Yu. Zinovyev, Ugur Dogrusoz, Jan Hasenauer, Ronan M. T. Fleming, Nicolas Le Novère, Piotr Gawron, Thomas S. Ligon, Anna Niarakis, David P. Nickerson, Daniel Weindl, Rudi Balling, Emmanuel Barillot, Charles Auffray, Reinhard Schneider 0002
Briefings Bioinform.16
2019 Assessing reproducibility of matrix factorization methods in independent transcriptomes
abstract
MOTIVATION: Matrix factorization (MF) methods are widely used in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). MF algorithms have never been compared based on the between-datasets reproducibility of their outputs in similar independent datasets. Lack of this knowledge might have a crucial impact when generalizing the predictions made in a study to others. RESULTS: We systematically test widely used MF methods on several transcriptomic datasets collected from the same cancer type (14 colorectal, 8 breast and 4 ovarian cancer transcriptomic datasets). Inspired by concepts of evolutionary bioinformatics, we design a novel framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the MF methods for their ability to produce generalizable components. We show that a particular protocol of application of independent component analysis (ICA), accompanied by a stabilization procedure, leads to a significant increase in the between-datasets reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other standard methods. We developed a user-friendly tool for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors associated to biological processes or to technological artifacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping. AVAILABILITY AND IMPLEMENTATION: The RBH construction tool is available from http://goo.gl/DzpwYp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Laura Cantini, Ulykbek Kairov, Aurélien de Reyniès, Emmanuel Barillot, François Radvanyi, Andrei Yu. Zinovyev
Bioinform.4
2019 PhysiBoSS: a multi-scale agent-based modelling framework integrating physical dimension and cell signalling
abstract
MOTIVATION: Due to the complexity and heterogeneity of multicellular biological systems, mathematical models that take into account cell signalling, cell population behaviour and the extracellular environment are particularly helpful. We present PhysiBoSS, an open source software which combines intracellular signalling using Boolean modelling (MaBoSS) and multicellular behaviour using agent-based modelling (PhysiCell). RESULTS: PhysiBoSS provides a flexible and computationally efficient framework to explore the effect of environmental and genetic alterations of individual cells at the population level, bridging the critical gap from single-cell genotype to single-cell phenotype and emergent multicellular behaviour. PhysiBoSS thus becomes very useful when studying heterogeneous population response to treatment, mutation effects, different modes of invasion or isomorphic morphogenesis events. To concretely illustrate a potential use of PhysiBoSS, we studied heterogeneous cell fate decisions in response to TNF treatment. We explored the effect of different treatments and the behaviour of several resistant mutants. We highlighted the importance of spatial information on the population dynamics by considering the effect of competition for resources like oxygen. AVAILABILITY AND IMPLEMENTATION: PhysiBoSS is freely available on GitHub (https://github.com/sysbio-curie/PhysiBoSS), with a Docker image (https://hub.docker.com/r/gletort/physiboss/). It is distributed as open source under the BSD 3-clause license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gaëlle Letort, Arnau Montagud, Gautier Stoll, Randy W. Heiland, Emmanuel Barillot, Paul Macklin, Andrei Yu. Zinovyev, Laurence Calzone
Bioinform.5
2019 Metabolic and signalling network maps integration: application to cross-talk studies and omics data analysis in cancer
abstract
BACKGROUND: The interplay between metabolic processes and signalling pathways remains poorly understood. Global, detailed and comprehensive reconstructions of human metabolism and signalling pathways exist in the form of molecular maps, but they have never been integrated together. We aim at filling in this gap by integrating of both signalling and metabolic pathways allowing a visual exploration of multi-level omics data and study of cross-regulatory circuits between these processes in health and in disease. RESULTS: We combined two comprehensive manually curated network maps. Atlas of Cancer Signalling Network (ACSN), containing mechanisms frequently implicated in cancer; and ReconMap 2.0, a comprehensive reconstruction of human metabolic network. We linked ACSN and ReconMap 2.0 maps via common players and represented the two maps as interconnected layers using the NaviCell platform for maps exploration ( https://navicell.curie.fr/pages/maps_ReconMap%202.html ). In addition, proteins catalysing metabolic reactions in ReconMap 2.0 were not previously visually represented on the map canvas. This precluded visualisation of omics data in the context of ReconMap 2.0. We suggested a solution for displaying protein nodes on the ReconMap 2.0 map in the vicinity of the corresponding reaction or process nodes. This permits multi-omics data visualisation in the context of both map layers. Exploration and shuttling between the two map layers is possible using Google Maps-like features of NaviCell. The integrated networks ACSN-ReconMap 2.0 are accessible online and allows data visualisation through various modes such as markers, heat maps, bar-plots, glyphs and map staining. The integrated networks were applied for comparison of immunoreactive and proliferative ovarian cancer subtypes using transcriptomic, copy number and mutation multi-omics data. A certain number of metabolic and signalling processes specifically deregulated in each of the ovarian cancer sub-types were identified. CONCLUSIONS: As knowledge evolves and new omics data becomes more heterogeneous, gathering together existing domains of biology under common platforms is essential. We believe that an integrated ACSN-ReconMap 2.0 networks will help in understanding various disease mechanisms and discovery of new interactions at the intersection of cell signalling and metabolism. In addition, the successful integration of metabolic and signalling networks allows broader systems biology approach application for data interpretation and retrieval of intervention points to tackle simultaneously the key players coordinating signalling and metabolism in human diseases.
Nicolas Sompairac, Jennifer Modamio, Emmanuel Barillot, Ronan M. T. Fleming, Andrei Yu. Zinovyev, Inna Kuperstein
BMC Bioinform.3
2018 QuantumClone: clonal assessment of functional mutations in cancer based on a genotype-aware method for clonal reconstruction
abstract
Motivation: In cancer, clonal evolution is assessed based on information coming from single nucleotide variants and copy number alterations. Nonetheless, existing methods often fail to accurately combine information from both sources to truthfully reconstruct clonal populations in a given tumor sample or in a set of tumor samples coming from the same patient. Moreover, previously published methods detect clones from a single set of variants. As a result, compromises have to be done between stringent variant filtering [reducing dispersion in variant allele frequency estimates (VAFs)] and using all biologically relevant variants. Results: We present a framework for defining cancer clones using most reliable variants of high depth of coverage and assigning functional mutations to the detected clones. The key element of our framework is QuantumClone, a method for variant clustering into clones based on VAFs, genotypes of corresponding regions and information about tumor purity. We validated QuantumClone and our framework on simulated data. We then applied our framework to whole genome sequencing data for 19 neuroblastoma trios each including constitutional, diagnosis and relapse samples. We confirmed an enrichment of damaging variants within such pathways as MAPK (mitogen-activated protein kinases), neuritogenesis, epithelial-mesenchymal transition, cell survival and DNA repair. Most pathways had more damaging variants in the expanding clones compared to shrinking ones, which can be explained by the increased total number of variants between these two populations. Functional mutational rate varied for ancestral clones and clones shrinking or expanding upon treatment, suggesting changes in clone selection mechanisms at different time points of tumor evolution. Availability and implementation: Source code and binaries of the QuantumClone R package are freely available for download at https://CRAN.R-project.org/package=QuantumClone. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Paul Deveau, Leo Colmet Daage, Derek Oldridge, Virginie Bernard, Angela Bellini, Mathieu Chicard, Nathalie Clement, Eve Lapouble, Valerie Combaret, Anne Boland, Vincent Meyer, Jean-François Deleuze, Isabelle Janoueix-Lerosey, Emmanuel Barillot, Olivier Delattre, John M. Maris, Gudrun Schleiermacher, Valentina Boeva
Bioinform.14
2018 Effective normalization for copy number variation in Hi-C data
abstract
BACKGROUND: Normalization is essential to ensure accurate analysis and proper interpretation of sequencing data, and chromosome conformation capture data such as Hi-C have particular challenges. Although several methods have been proposed, the most widely used type of normalization of Hi-C data usually casts estimation of unwanted effects as a matrix balancing problem, relying on the assumption that all genomic regions interact equally with each other. RESULTS: In order to explore the effect of copy-number variations on Hi-C data normalization, we first propose a simulation model that predict the effects of large copy-number changes on a diploid Hi-C contact map. We then show that the standard approaches relying on equal visibility fail to correct for unwanted effects in the presence of copy-number variations. We thus propose a simple extension to matrix balancing methods that model these effects. Our approach can either retain the copy-number variation effects (LOIC) or remove them (CAIC). We show that this leads to better downstream analysis of the three-dimensional organization of rearranged genomes. CONCLUSIONS: Taken together, our results highlight the importance of using dedicated methods for the analysis of Hi-C cancer data. Both CAIC and LOIC methods perform well on simulated and real Hi-C data sets, each fulfilling different needs.
Nicolas Servant, Nelle Varoquaux, Edith Heard, Emmanuel Barillot, Jean-Philippe Vert
BMC Bioinform.4
2017 MaBoSS 2.0: an environment for stochastic Boolean modeling
abstract
MOTIVATION: Modeling of signaling pathways is an important step towards the understanding and the treatment of diseases such as cancers, HIV or auto-immune diseases. MaBoSS is a software that allows to simulate populations of cells and to model stochastically the intracellular mechanisms that are deregulated in diseases. MaBoSS provides an output of a Boolean model in the form of time-dependent probabilities, for all biological entities (genes, proteins, phenotypes, etc.) of the model. RESULTS: We present a new version of MaBoSS (2.0), including an updated version of the core software and an environment. With this environment, the needs for modeling signaling pathways are facilitated, including model construction, visualization, simulations of mutations, drug treatments and sensitivity analyses. It offers a framework for automated production of theoretical predictions. AVAILABILITY AND IMPLEMENTATION: MaBoSS software can be found at https://maboss.curie.fr , including tutorials on existing models and examples of models. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gautier Stoll, Barthélémy Caron, Eric Viara, Aurélien Dugourd, Andrei Yu. Zinovyev, Aurélien Naldi, Guido Kroemer, Emmanuel Barillot, Laurence Calzone
Bioinform.8
2016 SV-Bay: structural variant detection in cancer genomes using a Bayesian approach with correction for GC-content and read mappability
abstract
MOTIVATION: Whole genome sequencing of paired-end reads can be applied to characterize the landscape of large somatic rearrangements of cancer genomes. Several methods for detecting structural variants with whole genome sequencing data have been developed. So far, none of these methods has combined information about abnormally mapped read pairs connecting rearranged regions and associated global copy number changes automatically inferred from the same sequencing data file. Our aim was to create a computational method that could use both types of information, i.e. normal and abnormal reads, and demonstrate that by doing so we can highly improve both sensitivity and specificity rates of structural variant prediction. RESULTS: We developed a computational method, SV-Bay, to detect structural variants from whole genome sequencing mate-pair or paired-end data using a probabilistic Bayesian approach. This approach takes into account depth of coverage by normal reads and abnormalities in read pair mappings. To estimate the model likelihood, SV-Bay considers GC-content and read mappability of the genome, thus making important corrections to the expected read count. For the detection of somatic variants, SV-Bay makes use of a matched normal sample when it is available. We validated SV-Bay on simulated datasets and an experimental mate-pair dataset for the CLB-GA neuroblastoma cell line. The comparison of SV-Bay with several other methods for structural variant detection demonstrated that SV-Bay has better prediction accuracy both in terms of sensitivity and false-positive detection rate. AVAILABILITY AND IMPLEMENTATION: https://github.com/InstitutCurie/SV-Bay CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daria Iakovishina, Isabelle Janoueix-Lerosey, Emmanuel Barillot, Mireille Régnier, Valentina Boeva
Bioinform.3
2015 Mathematical Modelling of Molecular Pathways Enabling Tumour Cell Invasion and Migration
abstract
Understanding the etiology of metastasis is very important in clinical perspective, since it is estimated that metastasis accounts for 90% of cancer patient mortality. Metastasis results from a sequence of multiple steps including invasion and migration. The early stages of metastasis are tightly controlled in normal cells and can be drastically affected by malignant mutations; therefore, they might constitute the principal determinants of the overall metastatic rate even if the later stages take long to occur. To elucidate the role of individual mutations or their combinations affecting the metastatic development, a logical model has been constructed that recapitulates published experimental results of known gene perturbations on local invasion and migration processes, and predict the effect of not yet experimentally assessed mutations. The model has been validated using experimental data on transcriptome dynamics following TGF-β-dependent induction of Epithelial to Mesenchymal Transition in lung cancer cell lines. A method to associate gene expression profiles with different stable state solutions of the logical model has been developed for that purpose. In addition, we have systematically predicted alleviating (masking) and synergistic pairwise genetic interactions between the genes composing the model with respect to the probability of acquiring the metastatic phenotype. We focused on several unexpected synergistic genetic interactions leading to theoretically very high metastasis probability. Among them, the synergistic combination of Notch overexpression and p53 deletion shows one of the strongest effects, which is in agreement with a recent published experiment in a mouse model of gut cancer. The mathematical model can recapitulate experimental mutations in both cell line and mouse models. Furthermore, the model predicts new gene perturbations that affect the early steps of metastasis underlying potential intervention points for innovative therapeutic strategies in oncology.
David P. A. Cohen, Loredana Martignetti, Sylvie Robine, Emmanuel Barillot, Andrei Yu. Zinovyev, Laurence Calzone
PLoS Comput. Biol.4
2014 Multi-factor data normalization enables the detection of copy number aberrations in amplicon sequencing data
abstract
MOTIVATION: Because of its low cost, amplicon sequencing, also known as ultra-deep targeted sequencing, is now becoming widely used in oncology for detection of actionable mutations, i.e. mutations influencing cell sensitivity to targeted therapies. Amplicon sequencing is based on the polymerase chain reaction amplification of the regions of interest, a process that considerably distorts the information on copy numbers initially present in the tumor DNA. Therefore, additional experiments such as single nucleotide polymorphism (SNP) or comparative genomic hybridization (CGH) arrays often complement amplicon sequencing in clinics to identify copy number status of genes whose amplification or deletion has direct consequences on the efficacy of a particular cancer treatment. So far, there has been no proven method to extract the information on gene copy number aberrations based solely on amplicon sequencing. RESULTS: Here we present ONCOCNV, a method that includes a multifactor normalization and annotation technique enabling the detection of large copy number changes from amplicon sequencing data. We validated our approach on high and low amplicon density datasets and demonstrated that ONCOCNV can achieve a precision comparable with that of array CGH techniques in detecting copy number aberrations. Thus, ONCOCNV applied on amplicon sequencing data would make the use of additional array CGH or SNP array experiments unnecessary.
Valentina Boeva, Tatiana G. Popova, Maxime Lienard, Sebastien Toffoli, Maud Kamal, Christophe Le Tourneau, David Gentien, Nicolas Servant, Pierre Gestraud, Thomas Rio Frio, Philippe Hupé, Emmanuel Barillot, Jean-François Laes
Bioinform.12
2013 HMCan: a method for detecting chromatin modifications in cancer samples using ChIP-seq data
abstract
MOTIVATION: Cancer cells are often characterized by epigenetic changes, which include aberrant histone modifications. In particular, local or regional epigenetic silencing is a common mechanism in cancer for silencing expression of tumor suppressor genes. Though several tools have been created to enable detection of histone marks in ChIP-seq data from normal samples, it is unclear whether these tools can be efficiently applied to ChIP-seq data generated from cancer samples. Indeed, cancer genomes are often characterized by frequent copy number alterations: gains and losses of large regions of chromosomal material. Copy number alterations may create a substantial statistical bias in the evaluation of histone mark signal enrichment and result in underdetection of the signal in the regions of loss and overdetection of the signal in the regions of gain. RESULTS: We present HMCan (Histone modifications in cancer), a tool specially designed to analyze histone modification ChIP-seq data produced from cancer genomes. HMCan corrects for the GC-content and copy number bias and then applies Hidden Markov Models to detect the signal from the corrected data. On simulated data, HMCan outperformed several commonly used tools developed to analyze histone modification data produced from genomes without copy number alterations. HMCan also showed superior results on a ChIP-seq dataset generated for the repressive histone mark H3K27me3 in a bladder cancer cell line. HMCan predictions matched well with experimental data (qPCR validated regions) and included, for example, the previously detected H3K27me3 mark in the promoter of the DLEC1 gene, missed by other tools we tested.
Haitham Ashoor, Aurélie Hérault, Aurélie Kamoun, François Radvanyi, Vladimir B. Bajic, Emmanuel Barillot, Valentina Boeva
Bioinform.6
2013 OCSANA: optimal combinations of interventions from network analysis
abstract
UNLABELLED: Targeted therapies interfering with specifically one protein activity are promising strategies in the treatment of diseases like cancer. However, accumulated empirical experience has shown that targeting multiple proteins in signaling networks involved in the disease is often necessary. Thus, one important problem in biomedical research is the design and prioritization of optimal combinations of interventions to repress a pathological behavior, while minimizing side-effects. OCSANA (optimal combinations of interventions from network analysis) is a new software designed to identify and prioritize optimal and minimal combinations of interventions to disrupt the paths between source nodes and target nodes. When specified by the user, OCSANA seeks to additionally minimize the side effects that a combination of interventions can cause on specified off-target nodes. With the crucial ability to cope with very large networks, OCSANA includes an exact solution and a novel selective enumeration approach for the combinatorial interventions' problem. AVAILABILITY: The latest version of OCSANA, implemented as a plugin for Cytoscape and distributed under LGPL license, is available together with source code at http://bioinfo.curie.fr/projects/ocsana.
Paola Vera-Licona, Eric Bonnet, Emmanuel Barillot, Andrei Yu. Zinovyev
Bioinform.3
2013 Synthetic Lethality between Gene Defects Affecting a Single Non-essential Molecular Pathway with Reversible Steps
abstract
Systematic analysis of synthetic lethality (SL) constitutes a critical tool for systems biology to decipher molecular pathways. The most accepted mechanistic explanation of SL is that the two genes function in parallel, mutually compensatory pathways, known as between-pathway SL. However, recent genome-wide analyses in yeast identified a significant number of within-pathway negative genetic interactions. The molecular mechanisms leading to within-pathway SL are not fully understood. Here, we propose a novel mechanism leading to within-pathway SL involving two genes functioning in a single non-essential pathway. This type of SL termed within-reversible-pathway SL involves reversible pathway steps, catalyzed by different enzymes in the forward and backward directions, and kinetic trapping of a potentially toxic intermediate. Experimental data with recombinational DNA repair genes validate the concept. Mathematical modeling recapitulates the possibility of kinetic trapping and revealed the potential contributions of synthetic, dosage-lethal interactions in such a genetic system as well as the possibility of within-pathway positive masking interactions. Analysis of yeast gene interaction and pathway data suggests broad applicability of this novel concept. These observations extend the canonical interpretation of synthetic-lethal or synthetic-sick interactions with direct implications to reconstruct molecular pathways and improve therapeutic approaches to diseases such as cancer.
Andrei Yu. Zinovyev, Inna Kuperstein, Emmanuel Barillot, Wolf-Dietrich Heyer
PLoS Comput. Biol.3
2012 Nebula - a web-server for advanced ChIP-seq data analysis
abstract
MOTIVATION: ChIP-seq consists of chromatin immunoprecipitation and deep sequencing of the extracted DNA fragments. It is the technique of choice for accurate characterization of the binding sites of transcription factors and other DNA-associated proteins. We present a web service, Nebula, which allows inexperienced users to perform a complete bioinformatics analysis of ChIP-seq data. RESULTS: Nebula was designed for both bioinformaticians and biologists. It is based on the Galaxy open source framework. Galaxy already includes a large number of functionalities for mapping reads and peak calling. We added the following to Galaxy: (i) peak calling with FindPeaks and a module for immunoprecipitation quality control, (ii) de novo motif discovery with ChIPMunk, (iii) calculation of the density and the cumulative distribution of peak locations relative to gene transcription start sites, (iv) annotation of peaks with genomic features and (v) annotation of genes with peak information. Nebula generates the graphs and the enrichment statistics at each step of the process. During Steps 3-5, Nebula optionally repeats the analysis on a control dataset and compares these results with those from the main dataset. Nebula can also incorporate gene expression (or gene modulation) data during these steps. In summary, Nebula is an innovative web service that provides an advanced ChIP-seq analysis pipeline providing ready-to-publish results. AVAILABILITY: Nebula is available at http://nebula.curie.fr/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valentina Boeva, Alban Lermine, Camille Barette, Christel Guillouf, Emmanuel Barillot
Bioinform.5
2012 Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data
abstract
SUMMARY: More and more cancer studies use next-generation sequencing (NGS) data to detect various types of genomic variation. However, even when researchers have such data at hand, single-nucleotide polymorphism arrays have been considered necessary to assess copy number alterations and especially loss of heterozygosity (LOH). Here, we present the tool Control-FREEC that enables automatic calculation of copy number and allelic content profiles from NGS data, and consequently predicts regions of genomic alteration such as gains, losses and LOH. Taking as input aligned reads, Control-FREEC constructs copy number and B-allele frequency profiles. The profiles are then normalized, segmented and analyzed in order to assign genotype status (copy number and allelic content) to each genomic region. When a matched normal sample is provided, Control-FREEC discriminates somatic from germline events. Control-FREEC is able to analyze overdiploid tumor samples and samples contaminated by normal cells. Low mappability regions can be excluded from the analysis using provided mappability tracks. AVAILABILITY: C++ source code is available at: http://bioinfo.curie.fr/projects/freec/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Valentina Boeva, Tatiana G. Popova, Kevin Bleakley, Pierre Chiche, Julie Cappo, Gudrun Schleiermacher, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot
Bioinform.9
2012 ncPRO-seq: a tool for annotation and profiling of ncRNAs in sRNA-seq data
abstract
SUMMARY: Non-coding RNA (ncRNA) PROfiling in small RNA (sRNA)-seq (ncPRO-seq) is a stand-alone, comprehensive and flexible ncRNA analysis pipeline. It can interrogate and perform detailed profiling analysis on sRNAs derived from annotated non-coding regions in miRBase, Rfam and RepeatMasker, as well as specific regions defined by users. The ncPRO-seq pipeline performs both gene-based and family-based analyses of sRNAs. It also has a module to identify regions significantly enriched with short reads, which cannot be classified under known ncRNA families, thus enabling the discovery of previously unknown ncRNA- or small interfering RNA (siRNA)-producing regions. The ncPRO-seq pipeline supports input read sequences in fastq, fasta and color space format, as well as alignment results in BAM format, meaning that sRNA raw data from the three current major platforms (Roche-454, Illumina-Solexa and Life technologies-SOLiD) can be analyzed with this pipeline. The ncPRO-seq pipeline can be used to analyze read and alignment data, based on any sequenced genome, including mammals and plants. AVAILABILITY: Source code, annotation files, manual and online version are available at http://ncpro.curie.fr/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chong-Jian Chen, Nicolas Servant, Joern Toedling, Alexis Sarazin, Antonin Marchais, Evelyne Duvernois-Berthet, Valérie Cognat, Vincent Colot, Olivier Voinnet, Edith Heard, Constance Ciaudo, Emmanuel Barillot
Bioinform.12
2012 HiTC: exploration of high-throughput 'C' experiments
abstract
Abstract Summary: The R/Bioconductor package HiTC facilitates the exploration of high-throughput 3C-based data. It allows users to import and export ‘C’ data, to transform, normalize, annotate and visualize interaction maps. The package operates within the Bioconductor framework and thus offers new opportunities for future development in this field. Availability and implementation: The R package HiTC is available from the Bioconductor website. A detailed vignette provides additional documentation and help for using the package. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Nicolas Servant, Bryan R. Lajoie, Elphège P. Nora, Luca Giorgetti, Chong-Jian Chen, Edith Heard, Job Dekker, Emmanuel Barillot
Bioinform.8
2011 Control-free calling of copy number alterations in deep-sequencing data using GC-content normalization
abstract
SUMMARY: We present a tool for control-free copy number alteration (CNA) detection using deep-sequencing data, particularly useful for cancer studies. The tool deals with two frequent problems in the analysis of cancer deep-sequencing data: absence of control sample and possible polyploidy of cancer cells. FREEC (control-FREE Copy number caller) automatically normalizes and segments copy number profiles (CNPs) and calls CNAs. If ploidy is known, FREEC assigns absolute copy number to each predicted CNA. To normalize raw CNPs, the user can provide a control dataset if available; otherwise GC content is used. We demonstrate that for Illumina single-end, mate-pair or paired-end sequencing, GC-contentr normalization provides smooth profiles that can be further segmented and analyzed in order to predict CNAs. AVAILABILITY: Source code and sample data are available at http://bioinfo-out.curie.fr/projects/freec/.
Valentina Boeva, Andrei Yu. Zinovyev, Kevin Bleakley, Jean-Philippe Vert, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot
Bioinform.7
2010 girafe - an R/Bioconductor package for functional exploration of aligned next-generation sequencing reads
abstract
UNLABELLED: The R/Bioconductor package girafe facilitates the functional exploration of alignments of sequence reads from next-generation sequencing data to a genome. It allows users to investigate the genomic intervals together with the aligned reads and to work with, visualise and export these intervals. Moreover, the package operates within and extends the ever-growing Bioconductor framework and thus enables users to leverage a multitude of methods for their data in order to answer specific research questions. AVAILABILITY AND IMPLEMENTATION: The R package girafe is available from the Bioconductor web site: http://www.bioconductor.org/packages/release/bioc/html/girafe.html. An extensive vignette and the Bioconductor mailing lists provide additional documentation and help for using the package.
Joern Toedling, Constance Ciaudo, Olivier Voinnet, Edith Heard, Emmanuel Barillot
Bioinform.5
2010 SVDetect: a tool to identify genomic structural variations from paired-end and mate-pair sequencing data
abstract
SUMMARY: We present SVDetect, a program designed to identify genomic structural variations from paired-end and mate-pair next-generation sequencing data produced by the Illumina GA and ABI SOLiD platforms. Applying both sliding-window and clustering strategies, we use anomalously mapped read pairs provided by current short read aligners to localize genomic rearrangements and classify them according to their type, e.g. large insertions-deletions, inversions, duplications and balanced or unbalanced inter-chromosomal translocations. SVDetect outputs predicted structural variants in various file formats for appropriate graphical visualization. AVAILABILITY: Source code and sample data are available at http://svdetect.sourceforge.net/
Bruno Zeitouni, Valentina Boeva, Isabelle Janoueix-Lerosey, Sophie Loeillet, Patricia Legoix-né, Alain Nicolas, Olivier Delattre, Emmanuel Barillot
Bioinform.8
2010 Mathematical Modelling of Cell-Fate Decision in Response to Death Receptor Engagement
abstract
Cytokines such as TNF and FASL can trigger death or survival depending on cell lines and cellular conditions. The mechanistic details of how a cell chooses among these cell fates are still unclear. The understanding of these processes is important since they are altered in many diseases, including cancer and AIDS. Using a discrete modelling formalism, we present a mathematical model of cell fate decision recapitulating and integrating the most consistent facts extracted from the literature. This model provides a generic high-level view of the interplays between NFkappaB pro-survival pathway, RIP1-dependent necrosis, and the apoptosis pathway in response to death receptor-mediated signals. Wild type simulations demonstrate robust segregation of cellular responses to receptor engagement. Model simulations recapitulate documented phenotypes of protein knockdowns and enable the prediction of the effects of novel knockdowns. In silico experiments simulate the outcomes following ligand removal at different stages, and suggest experimental approaches to further validate and specialise the model for particular cell types. We also propose a reduced conceptual model implementing the logic of the decision process. This analysis gives specific predictions regarding cross-talks between the three pathways, as well as the transient role of RIP1 protein in necrosis, and confirms the phenotypes of novel perturbations. Our wild type and mutant simulations provide novel insights to restore apoptosis in defective cells. The model analysis expands our understanding of how cell fate decision is made. Moreover, our current model can be used to assess contradictory or controversial data from the literature. Ultimately, it constitutes a valuable reasoning tool to delineate novel experiments.
Laurence Calzone, Laurent Tournier, Simon Fourquet, Denis Thieffry, Boris Zhivotovsky, Emmanuel Barillot, Andrei Yu. Zinovyev
PLoS Comput. Biol.6
2008 Classification of arrayCGH data using fused SVM
abstract
MOTIVATION: Array-based comparative genomic hybridization (arrayCGH) has recently become a popular tool to identify DNA copy number variations along the genome. These profiles are starting to be used as markers to improve prognosis or diagnosis of cancer, which implies that methods for automated supervised classification of arrayCGH data are needed. Like gene expression profiles, arrayCGH profiles are characterized by a large number of variables usually measured on a limited number of samples. However, arrayCGH profiles have a particular structure of correlations between variables, due to the spatial organization of bacterial artificial chromosomes along the genome. This suggests that classical classification methods, often based on the selection of a small number of discriminative features, may not be the most accurate methods and may not produce easily interpretable prediction rules. RESULTS: We propose a new method for supervised classification of arrayCGH data. The method is a variant of support vector machine that incorporates the biological specificities of DNA copy number variations along the genome as prior knowledge. The resulting classifier is a sparse linear classifier based on a limited number of regions automatically selected on the chromosomes, leading to easy interpretation and identification of discriminative regions of the genome. We test this method on three classification problems for bladder and uveal cancer, involving both diagnosis and prognosis. We demonstrate that the introduction of the new prior on the classifier leads not only to more accurate predictions, but also to the identification of known and new regions of interest in the genome. AVAILABILITY: All data and algorithms are publicly available.
Franck Rapaport, Emmanuel Barillot, Jean-Philippe Vert
ISMB2
2008 ITALICS: an algorithm for normalization and DNA copy number calling for Affymetrix SNP arrays
abstract
MOTIVATION: Affymetrix SNP arrays can be used to determine the DNA copy number measurement of 11 000-500 000 SNPs along the genome. Their high density facilitates the precise localization of genomic alterations and makes them a powerful tool for studies of cancers and copy number polymorphism. Like other microarray technologies it is influenced by non-relevant sources of variation, requiring correction. Moreover, the amplitude of variation induced by non-relevant effects is similar or greater than the biologically relevant effect (i.e. true copy number), making it difficult to estimate non-relevant effects accurately without including the biologically relevant effect. RESULTS: We addressed this problem by developing ITALICS, a normalization method that estimates both biological and non-relevant effects in an alternate, iterative manner, accurately eliminating irrelevant effects. We compared our normalization method with other existing and available methods, and found that ITALICS outperformed these methods for several in-house datasets and one public dataset. These results were validated biologically by quantitative PCR. AVAILABILITY: The R package ITALICS (ITerative and Alternative normaLIzation and Copy number calling for affymetrix Snp arrays) has been submitted to Bioconductor.
Guillem Rigaill, Philippe Hupé, Anna Almeida, Philippe La Rosa, Jean-Philippe Meyniel, Charles Decraene, Emmanuel Barillot
Bioinform.7
2008 BiNoM: a Cytoscape plugin for manipulating and analyzing biological networks
abstract
UNLABELLED: BiNoM (Biological Network Manager) is a new bioinformatics software that significantly facilitates the usage and the analysis of biological networks in standard systems biology formats (SBML, SBGN, BioPAX). BiNoM implements a full-featured BioPAX editor and a method of 'interfaces' for accessing BioPAX content. BiNoM is able to work with huge BioPAX files such as whole pathway databases. In addition, BiNoM allows the analysis of networks created with CellDesigner software and their conversion into BioPAX format. BiNoM comes as a library and as a Cytoscape plugin which adds a rich set of operations to Cytoscape such as path and cycle analysis, clustering sub-networks, decomposition of network into modules, clipboard operations and others. AVAILABILITY: Last version of BiNoM distributed under the LGPL licence together with documentation, source code and API are available at http://bioinfo.curie.fr/projects/binom
Andrei Yu. Zinovyev, Eric Viara, Laurence Calzone, Emmanuel Barillot
Bioinform.4
2007 LICORN: learning cooperative regulation networks from gene expression data
abstract
MOTIVATION: One of the most challenging tasks in the post-genomic era is the reconstruction of transcriptional regulation networks. The goal is to identify, for each gene expressed in a particular cellular context, the regulators affecting its transcription, and the co-ordination of several regulators in specific types of regulation. DNA microarrays can be used to investigate relationships between regulators and their target genes, through simultaneous observations of their RNA levels. RESULTS: We propose a data mining system for inferring transcriptional regulation relationships from RNA expression values. This system is particularly suitable for the detection of cooperative transcriptional regulation. We model regulatory relationships as labelled two-layer gene regulatory networks, and describe a method for the efficient learning of these bipartite networks from discretized expression data sets. We also evaluate the statistical significance of such inferred networks and validate our methods on two public yeast expression data sets. AVAILABILITY: http://www.lri.fr/~elati/licorn.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mohamed Elati, Pierre Neuvial, Monique Bolotin-Fukuhara, Emmanuel Barillot, François Radvanyi, Céline Rouveirol
Bioinform.4
2007 Software package for automatic microarray image analysis (MAIA)
abstract
UNLABELLED: Although various software solutions are currently available for microarray image analysis, one would still expect to develop algorithms ensuring higher level of intelligence and robustness. We present a fully functional software package for automatic processing of the two-color microarray images including spot localization, quantification and quality control. The developed algorithms aim at making ratio estimates more resistant to array contamination and offer automatic tools to evaluate spot quality. AVAILABILITY: A demo version of the software can be downloaded from http://bioinfo.curie.fr/projects/maia. A full version is freely available to non-commercial users upon request from the authors.
Eugene Novikov, Emmanuel Barillot
Bioinform.2
2007 Classification of microarray data using gene networks
abstract
BACKGROUND: Microarrays have become extremely useful for analysing genetic phenomena, but establishing a relation between microarray analysis results (typically a list of genes) and their biological significance is often difficult. Currently, the standard approach is to map a posteriori the results onto gene networks in order to elucidate the functions perturbed at the level of pathways. However, integrating a priori knowledge of the gene networks could help in the statistical analysis of gene expression data and in their biological interpretation. RESULTS: We propose a method to integrate a priori the knowledge of a gene network in the analysis of gene expression data. The approach is based on the spectral decomposition of gene expression profiles with respect to the eigenfunctions of the graph, resulting in an attenuation of the high-frequency components of the expression profiles with respect to the topology of the graph. We show how to derive unsupervised and supervised classification algorithms of expression profiles, resulting in classifiers with biological relevance. We illustrate the method with the analysis of a set of expression profiles from irradiated and non-irradiated yeast strains. CONCLUSION: Including a priori knowledge of a gene network for the analysis of gene expression data leads to good classification performance and improved interpretability of the results.
Franck Rapaport, Andrei Yu. Zinovyev, Marie Dutreix, Emmanuel Barillot, Jean-Philippe Vert
BMC Bioinform.4
2006 VAMP: Visualization and analysis of array-CGH, transcriptome and other molecular profiles
abstract
MOTIVATION: Microarray-based CGH (Comparative Genomic Hybridization), transcriptome arrays and other large-scale genomic technologies are now routinely used to generate a vast amount of genomic profiles. Exploratory analysis of this data is crucial in helping to understand the data and to help form biological hypotheses. This step requires visualization of the data in a meaningful way to visualize the results and to perform first level analyses. RESULTS: We have developed a graphical user interface for visualization and first level analysis of molecular profiles. It is currently in use at the Institut Curie for cancer research projects involving CGH arrays, transcriptome arrays, SNP (single nucleotide polymorphism) arrays, loss of heterozygosity results (LOH), and Chromatin ImmunoPrecipitation arrays (ChIP chips). The interface offers the possibility of studying these different types of information in a consistent way. Several views are proposed, such as the classical CGH karyotype view or genome-wide multi-tumor comparison. Many functionalities for analyzing CGH data are provided by the interface, including looking for recurrent regions of alterations, confrontation to transcriptome data or clinical information, and clustering. Our tool consists of PHP scripts and of an applet written in Java. It can be run on public datasets at http://bioinfo.curie.fr/vamp AVAILABILITY: The VAMP software (Visualization and Analysis of array-CGH,transcriptome and other Molecular Profiles) is available upon request. It can be tested on public datasets at http://bioinfo.curie.fr/vamp. The documentation is available at http://bioinfo.curie.fr/vamp/doc.
Philippe La Rosa, Eric Viara, Philippe Hupé, Gaëlle Pierron, Stéphane Liva, Pierre Neuvial, Isabel Brito 0002, Séverine Lair, Nicolas Servant, Nicolas Robine, Elodie Manié, Caroline Brennetot, Isabelle Janoueix-Lerosey, Virginie Raynal, Nadège Gruel, Céline Rouveirol, Nicolas Stransky, Marc-Henri Stern, Olivier Delattre, Alain Aurias, François Radvanyi, Emmanuel Barillot
Bioinform.22
2006 Computation of recurrent minimal genomic alterations from array-CGH data
abstract
MOTIVATION: The identification of recurrent genomic alterations can provide insight into the initiation and progression of genetic diseases, such as cancer. Array-CGH can identify chromosomal regions that have been gained or lost, with a resolution of approximately 1 mb, for the cutting-edge techniques. The extraction of discrete profiles from raw array-CGH data has been studied extensively, but subsequent steps in the analysis require flexible, efficient algorithms, particularly if the number of available profiles exceeds a few tens or the number of array probes exceeds a few thousands. RESULTS: We propose two algorithms for computing minimal and minimal constrained regions of gain and loss from discretized CGH profiles. The second of these algorithms can handle additional constraints describing relevant regions of copy number change. We have validated these algorithms on two public array-CGH datasets. AVAILABILITY: From the authors, upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Céline Rouveirol, Nicolas Stransky, Philippe Hupé, Philippe La Rosa, Eric Viara, Emmanuel Barillot, François Radvanyi
Bioinform.6
2006 Spatial normalization of array-CGH data
abstract
BACKGROUND: Array-based comparative genomic hybridization (array-CGH) is a recently developed technique for analyzing changes in DNA copy number. As in all microarray analyses, normalization is required to correct for experimental artifacts while preserving the true biological signal. We investigated various sources of systematic variation in array-CGH data and identified two distinct types of spatial effect of no biological relevance as the predominant experimental artifacts: continuous spatial gradients and local spatial bias. Local spatial bias affects a large proportion of arrays, and has not previously been considered in array-CGH experiments. RESULTS: We show that existing normalization techniques do not correct these spatial effects properly. We therefore developed an automatic method for the spatial normalization of array-CGH data. This method makes it possible to delineate and to eliminate and/or correct areas affected by spatial bias. It is based on the combination of a spatial segmentation algorithm called NEM (Neighborhood Expectation Maximization) and spatial trend estimation. We defined quality criteria for array-CGH data, demonstrating significant improvements in data quality with our method for three data sets coming from two different platforms (198, 175 and 26 BAC-arrays). CONCLUSION: We have designed an automatic algorithm for the spatial normalization of BAC CGH-array data, preventing the misinterpretation of experimental artifacts as biologically relevant outliers in the genomic profile. This algorithm is implemented in the R package MANOR (Micro-Array NORmalization), which is described at http://bioinfo.curie.fr/projects/manor and available from the Bioconductor site http://www.bioconductor.org. It can also be tested on the CAPweb bioinformatics platform at http://bioinfo.curie.fr/CAPweb.
Pierre Neuvial, Philippe Hupé, Isabel Brito 0002, Stéphane Liva, Elodie Manié, Caroline Brennetot, François Radvanyi, Alain Aurias, Emmanuel Barillot
BMC Bioinform.9
2006 A noise-resistant algorithm for grid finding in microarray image analysis
Eugene Novikov, Emmanuel Barillot
Mach. Vis. Appl.2
2005 An algorithm for automatic evaluation of the spot quality in two-color DNA microarray experiments
abstract
BACKGROUND: Although DNA microarray technologies are very powerful for the simultaneous quantitative characterization of thousands of genes, the quality of the obtained experimental data is often far from ideal. The measured microarrays images represent a regular collection of spots, and the intensity of light at each spot is proportional to the DNA copy number or to the expression level of the gene whose DNA clone is spotted. Spot quality control is an essential part of microarray image analysis, which must be carried out at the level of individual spot identification. The problem is difficult to formalize due to the diversity of instrumental and biological factors that can influence the result. RESULTS: For each spot we estimate the ratio of measured fluorescence intensities revealing differential gene expression or change in DNA copy numbers between the test and control samples. We also define a set of quality characteristics and a model for combining these characteristics into an overall spot quality value. We have developed a training procedure to evaluate the contribution of each individual characteristic in the overall quality. This procedure uses information available from replicated spots, located in the same array or over a set of replicated arrays. It is assumed that unspoiled replicated spots must have very close ratios, whereas poor spots yield greater diversity in the obtained ratio estimates. CONCLUSION: The developed procedure provides an automatic tool to quantify spot quality and to identify different types of spot deficiency occurring in DNA microarray technology. Quality values assigned to each spot can be used either to eliminate spots or to weight contribution of each ratio estimate in follow-up analysis procedures.
Eugene Novikov, Emmanuel Barillot
BMC Bioinform.2
2004 Analysis of array CGH data: from signal ratio to gain and loss of DNA regions
abstract
MOTIVATION: Genomic DNA regions are frequently lost or gained during tumor progression. Array Comparative Genomic Hybridization (array CGH) technology makes it possible to assess these changes in DNA in cancers, by comparison with a normal reference. The identification of systematically deleted or amplified genomic regions in a set of tumors enables biologists to identify genes involved in cancer progression because tumor suppressor genes are thought to be located in lost genomic regions and oncogenes, in gained regions. Array CGH profiles should also improve the classification of tumors. The achievement of these goals requires a methodology for detecting the breakpoints delimiting altered regions in genomic patterns and assigning a status (normal, gained or lost) to each chromosomal region. RESULTS: We have developed a methodology for the automatic detection of breakpoints from array CGH profile, and the assignment of a status to each chromosomal region. The breakpoint detection step is based on the Adaptive Weights Smoothing (AWS) procedure and provides highly convincing results: our algorithm detects 97, 100 and 94% of breakpoints in simulated data, karyotyping results and manually analyzed profiles, respectively. The percentage of correctly assigned statuses ranges from 98.9 to 99.8% for simulated data and is 100% for karyotyping results. Our algorithm also outperforms other solutions on a public reference dataset. AVAILABILITY: The R package GLAD (Gain and Loss Analysis of DNA) is available upon request.
Philippe Hupé, Nicolas Stransky, Jean-Paul Thiery, François Radvanyi, Emmanuel Barillot
Bioinform.5
2002 Distributing CORBA views from an OODBMS
abstract
The need to distribute objects on the Internet and to offer views from databases has found a solution with the advent of CORBA. Most database management systems now offer CORBA interfaces which are generally simple mapping of the database schema to the CORBA world. This approach does not address all the problems of database interoperation because (i) such a view is static (ii) its semantic is completely bound to the semantic of the schema and it is not possible to re-model it (iii) only one view per database can be offered (iv) access may be limited to reading and no mechanism is given to write in the database through the view. To solve these problems, we have designed a language, the Interface Mapping Definition Language (IMDL) and some tools, grouped in the Interface Mapping Service (IMS). IMDL is used to define CORBA views from OODBMS, while IMS generates an IDL construct and a full CORBA implementation from an IMDL construct and a database schema.
Eric Viara, Guy Vaysseix, Emmanuel Barillot
IDEAS3
2001 XML, bioinformatics and data integration
abstract
Abstract Motivation: The eXtensible Markup Language (XML) is an emerging standard for structuring documents, notably for the World Wide Web. In this paper, the authors present XML and examine its use as a data language for bioinformatics. In particular, XML is compared to other languages, and some of the potential uses of XML in bioinformatics applications are presented. The authors propose to adopt XML for data interchange between databases and other sources of data. Finally the discussion is illustrated by a test case of a pedigree data model in XML. Contact: [email protected] * To whom correspondence should be addressed. 3 Present address: Cereon Genomics, 45 Sidney St, Cambridge, MA 02139, USA.
Frédéric Achard, Guy Vaysseix, Emmanuel Barillot
Bioinform.3
2001 The art of pedigree drawing: algorithmic aspects
abstract
Abstract Motivation: Giving a meaningful representation of a pedigree is not obvious when it includes consanguinity loops, individuals with multiple mates or several related families. Results: We show that finding a perfectly meaningful representation of a pedigree is equivalent to the interval graph sandwich problem and we propose an algorithm for drawing pedigrees. Contact: [email protected] * To whom correspondence should be addressed.
Frédéric Tores, Emmanuel Barillot
Bioinform.2
2000 Context and Interaction in Zooable User Interfaces
abstract
Zoomable User Interfaces (ZUIs) are difficult to use on large information spaces in part because they provide insufficient context. Even after a short period of navigation users no longer know where they are in the information space nor where to find the information they are looking for. We propose a temporary in-place context aid that helps users position themselves in ZUIs. This context layer is a transparent view of the context that is drawn over the users' focus of attention. A second temporary in-place aid is proposed that can be used to view already visited regions of the information space. This history layer is an overlapping transparent layer that adds a history mechanism to ZUIs. We complete these orientation aids with an additional window, a hierarchy tree, that shows users the structure of the information space and their current position within it. Context layers show users their position, history layers show them how they got there, and hierarchy trees show what information is available and where it is.
Stuart Pook, Eric Lecolinet, Guy Vaysseix, Emmanuel Barillot
Advanced Visual Interfaces4
1999 The EyeDB OODBMS
abstract
This paper introduces the EYEDB object oriented database management system (OODBMS). EYEDB implements all the standard features of OODBMS, is language oriented, provides a generic object model and a support for data distribution using CORBA. It can deal with very large databases, and is both efficient and scalable. It is used in the genome project where a huge amount of data have to be managed and intricate data structures needs to be modeled. Online information and a trial version of EYEDB can be obtained from http://www.sysra.com/eyedb.
Eric Viara, Emmanuel Barillot, Guy Vaysseix
IDEAS2
1999 A proposal for a standard CORBA interface for genome maps
abstract
MOTIVATION: The scientific community urgently needs to standardize the exchange of biological data. This is helped by the use of a common protocol and the definition of shared data structures. We have based our standardization work on CORBA, a technology that has become a standard in the past years and allows interoperability between distributed objects. RESULTS: We have defined an IDL specification for genome maps and present it to the scientific community. We have implemented CORBA servers based on this IDL to distribute RHdb and HuGeMap maps. The IDL will co-evolve with the needs of the mapping community. AVAILABILITY: The standard IDL for genome maps is available at http:// corba.ebi.ac.uk/RHdb/EUCORBA/MapIDL.htm l. The IORs to browse maps from Infobiogen and EBI are at http://www.infobiogen.fr/services/Hugemap/IOR and http://corba.ebi.ac.uk/RHdb/EUCORBA/IOR CONTACT: [email protected], [email protected]
Emmanuel Barillot, Ulf Leser, Philip Lijnzaad, Christophe Cussat-Blanc, Kim Jungfer, Frédéric Guyon, Guy Vaysseix, Carsten Helgesen, Patricia Rodriguez-Tomé
Bioinform.1
1999 CoPE: a collaborative pedigree drawing environment
abstract
SUMMARY: We developed a collaborative pedigree environment called CoPE. This environment includes a Java program for drawing pedigrees and a standardized system for pedigree storage. Unlike other existing pedigree programs, this software is particularly intended for epidemiologists in the sense that it allows customized automatic drawing of large numbers of pedigrees and remote and distributed consultation of pedigrees. AVAILABILITY: At http://www.infobiogen.fr/services/CoPE
L. Brun-Samarcq, Sophie Gallina, A. Philippi, Florence Demenais, Guy Vaysseix, Emmanuel Barillot
Bioinform.6
1998 The new Virgil database: a service of rich links
abstract
MOTIVATION: Links between biological objects are frequently used by researchers in biology. However, many of the links found in public databases are insufficiently documented and difficult to retrieve. Virgil introduces the idea of a rich link, i.e. the link itself and the related pieces of information. Virgil was developed to collect, manage and distribute such links. RESULTS: At the moment, Virgil is a prototype database that contains rich links between GDB genes and Genbank sequences. The Virgil data model is rich enough to describe comprehensively a link between two biological objects. Two different means to access the information were developed: a schema-driven Web interface and a CORBA server. AVAILABILITY: http://www.infobiogen. fr/services/virgil/home.html CONTACT: [email protected]
Frédéric Achard, Christophe Cussat-Blanc, Eric Viara, Emmanuel Barillot
Bioinform.4
1998 Zomit: biological data visualization and browsing
abstract
MOTIVATION: The problems caused by the difficulty in visualizing and browsing biological databases have become crucial. Scientists can no longer interact directly with the huge amount of available data. However, future breakthroughs in biology depend on this interaction. We propose a new metaphor for biological data visualization and browsing that allows navigation in very large databases in an intuitive way. The concepts underlying our approach are based on navigation and visualization with zooming, semantic zooming and portals; and on data transformation via magic lenses. We think that these new visualization and navigation techniques should be applied globally to a federation of biological databases. RESULTS: We have implemented a generic tool, called Zomit, that provides an application programming interface for developing servers for such navigation and visualization, and a generic architecture-independent client (Javatrade mark applet) that queries such servers. As an illustration of the capabilities of our approach, we have developed ZoomMap, a prototype browser for the HuGeMap human genome map database. AVAILABILITY: Zomit and ZoomMap are available at the URL http://www.infobiogen.fr/services/zomit.
Stuart Pook, Guy Vaysseix, Emmanuel Barillot
Bioinform.3