EDBT 2026 Demo / reviewers in the wild / expert
Andrei Yu. Zinovyev
dblp:79/6040
· DBLP profile ↗
36ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-9517-7284ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | scBoolSeq: Linking scRNA-seq statistics and Boolean dynamicsabstractBoolean networks are largely employed to model the qualitative dynamics of cell fate processes by describing the change of binary activation states of genes and transcription factors with time. Being able to bridge such qualitative states with quantitative measurements of gene expression in cells, as scRNA-seq, is a cornerstone for data-driven model construction and validation. On one hand, scRNA-seq binarisation is a key step for inferring and validating Boolean models. On the other hand, the generation of synthetic scRNA-seq data from baseline Boolean models provides an important asset to benchmark inference methods. However, linking characteristics of scRNA-seq datasets, including dropout events, with Boolean states is a challenging task. We present scBoolSeq, a method for the bidirectional linking of scRNA-seq data and Boolean activation state of genes. Given a reference scRNA-seq dataset, scBoolSeq computes statistical criteria to classify the empirical gene pseudocount distributions as either unimodal, bimodal, or zero-inflated, and fit a probabilistic model of dropouts, with gene-dependent parameters. From these learnt distributions, scBoolSeq can perform both binarisation of scRNA-seq datasets, and generate synthetic scRNA-seq datasets from Boolean traces, as issued from Boolean networks, using biased sampling and dropout simulation. We present a case study demonstrating the application of scBoolSeq's binarisation scheme in data-driven model inference. Furthermore, we compare synthetic scRNA-seq data generated by scBoolSeq with BoolODE's, data for the same Boolean Network model. The comparison shows that our method better reproduces the statistics of real scRNA-seq datasets, such as the mean-variance and mean-dropout relationships while exhibiting clearly defined trajectories in two-dimensional projections of the data. Gustavo Magaña López, Laurence Calzone, Andrei Yu. Zinovyev, Loïc Paulevé |
PLoS Comput. Biol. | 3 |
| 2023 | Multiscale model of the different modes of cancer cell invasionabstractMOTIVATION: Mathematical models of biological processes altered in cancer are built using the knowledge of complex networks of signaling pathways, detailing the molecular regulations inside different cell types, such as tumor cells, immune and other stromal cells. If these models mainly focus on intracellular information, they often omit a description of the spatial organization among cells and their interactions, and with the tumoral microenvironment. RESULTS: We present here a model of tumor cell invasion simulated with PhysiBoSS, a multiscale framework, which combines agent-based modeling and continuous time Markov processes applied on Boolean network models. With this model, we aim to study the different modes of cell migration and to predict means to block it by considering not only spatial information obtained from the agent-based simulation but also intracellular regulation obtained from the Boolean model. Our multiscale model integrates the impact of gene mutations with the perturbation of the environmental conditions and allows the visualization of the results with 2D and 3D representations. The model successfully reproduces single and collective migration processes and is validated on published experiments on cell invasion. In silico experiments are suggested to search for possible targets that can block the more invasive tumoral phenotypes. AVAILABILITY AND IMPLEMENTATION: https://github.com/sysbio-curie/Invasion_model_PhysiBoSS. Marco Ruscone, Arnau Montagud, Philippe Chavrier, Olivier Destaing, Isabelle Bonnet, Andrei Yu. Zinovyev, Emmanuel Barillot, Vincent Noel, Laurence Calzone |
Bioinform. | 6 |
| 2022 | Quasi-orthogonality and intrinsic dimensions as measures of learning and generalisationabstractFinding the best architectures for learning machines, such as deep neural networks, is a well-known technical and theoretical challenge. Recent work by Mellor et al [1] showed that there may exist correlations between the accuracies of trained networks and the values of some easily computable measures defined on randomly initialised networks which may enable the search of tens of thousands of neural architectures without training. Mellor et al [1] used the Hamming distance evaluated over all the ReLU neurons as such a measure. Motivated by these findings, in our work, we ask the question of the existence of other and perhaps more principled measures which could be used as determinants of the potential success of a given neural architecture. In particular, we examine if the dimensionality and quasi-orthogonality of neural networks' feature space could be correlated with the network's performance after training. We showed, using the setup as in Mellor et al [1], that dimensionality and quasi-orthogonality may jointly serve as networks' performance discriminants. In addition to offering new opportunities to accelerate neural architecture search, our findings suggest important relationships between the networks' final performance and properties of their randomly initialised feature spaces: data dimension and quasi-orthogonality. Alexander N. Gorban, Eugenij Moiseevich Mirkes, Jonathan Bac, Andrei Yu. Zinovyev, Ivan Tyukin |
IJCNN | 5 |
| 2022 | Hubness reduction improves clustering and trajectory inference in single-cell transcriptomic dataabstractMOTIVATION: Single-cell RNA-seq (scRNAseq) datasets are characterized by large ambient dimensionality, and their analyses can be affected by various manifestations of the dimensionality curse. One of these manifestations is the hubness phenomenon, i.e. existence of data points with surprisingly large incoming connectivity degree in the datapoint neighbourhood graph. Conventional approach to dampen the unwanted effects of high dimension consists in applying drastic dimensionality reduction. It remains unexplored if this step can be avoided thus retaining more information than contained in the low-dimensional projections, by correcting directly hubness. RESULTS: We investigated hubness in scRNAseq data. We show that hub cells do not represent any visible technical or biological bias. The effect of various hubness reduction methods is investigated with respect to the clustering, trajectory inference and visualization tasks in scRNAseq datasets. We show that hubness reduction generates neighbourhood graphs with properties more suitable for applying machine learning methods; and that it outperforms other state-of-the-art methods for improving neighbourhood graphs. As a consequence, clustering, trajectory inference and visualization perform better, especially for datasets characterized by large intrinsic dimensionality. Hubness is an important phenomenon characterizing data point neighbourhood graphs computed for various types of sequencing datasets. Reducing hubness can be beneficial for the analysis of scRNAseq data with large intrinsic dimensionality in which case it can be an alternative to drastic dimensionality reduction. AVAILABILITY AND IMPLEMENTATION: The code used to analyze the datasets and produce the figures of this article is available from https://github.com/sysbio-curie/schubness. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Elise Amblard, Jonathan Bac, Alexander Chervov, Vassili Soumelis, Andrei Yu. Zinovyev |
Bioinform. | 5 |
| 2022 | BIODICA: a computational environment for Independent Component Analysis of omics dataabstractSUMMARY: We developed BIODICA, an integrated computational environment for application of independent component analysis (ICA) to bulk and single-cell molecular profiles, interpretation of the results in terms of biological functions and correlation with metadata. The computational core is the novel Python package stabilized-ica which provides interface to several ICA algorithms, a stabilization procedure, meta-analysis and component interpretation tools. BIODICA is equipped with a user-friendly graphical user interface, allowing non-experienced users to perform the ICA-based omics data analysis. The results are provided in interactive ways, thus facilitating communication with biology experts. AVAILABILITY AND IMPLEMENTATION: BIODICA is implemented in Java, Python and JavaScript. The source code is freely available on GitHub under the MIT and the GNU LGPL licenses. BIODICA is supported on all major operating systems. URL: https://sysbio-curie.github.io/biodica-environment/. Nicolas Captier, Jane Merlevede, Askhat Molkenov, Ainur Seisenova, Altynbek Zhubanchaliyev, Petr V. Nazarov, Emmanuel Barillot, Ulykbek Kairov, Andrei Yu. Zinovyev |
Bioinform. | 9 |
| 2022 | Coloring Panchromatic Nighttime Satellite Images: Comparing the Performance of Several Machine Learning MethodsabstractArtificial light-at-night (ALAN), emitted from the ground and visible from space, marks human presence on earth. Since the launch of the Suomi National Polar Partnership satellite with the Visible Infrared Imaging Radiometer Suite Day–Night Band (VIIRS/DNB) onboard, global nighttime images have significantly improved; however, they remained panchromatic. Although multispectral images are also available, they are either commercial or free of charge, but sporadic. In this article, we use several machine learning techniques, such as linear, kernel, random forest regressions, and elastic map approach, to transform panchromatic VIIRS/DBN into red, green, blue (RGB) images. To validate the proposed approach, we analyze RGB images for eight urban areas worldwide. We link RGB values, obtained from ISS photographs, to panchromatic ALAN intensities, their pixel-wise differences, and several land-use-type proxies. Each dataset is used for model training, while other datasets are used for model validation. The analysis shows that model-estimated RGB images demonstrate a high degree of correspondence with the original RGB images from the ISS database. Yet, estimates, based on linear, kernel, and random forest regressions, provide better correlations, contrast similarity, and lower WMSEs levels, while RGB images, generated using elastic map approach, provide higher consistency of predictions. Natalya A. Rybnikova, Boris A. Portnov, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev, Anna Brook, Alexander N. Gorban |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Clinical trajectories estimated from bulk tumoral molecular profiles using elastic principal treesabstractClinical trajectory is a clinically relevant sequence of ordered patient phenotypes representing consecutive states of a developing disease and leading to some final state. Extracting trajectories from large scale medical data is of great interest for dynamical phenotyping of various diseases but remains a challenge for machine learning methods, especially in the case of synchronic (with short follow up) observations. Here we describe an approach for trajectory-based analysis of cancer data using elastic principal trees and test it on a large collection of molecular tumoral profiles for breast cancer. We show that the disease progress quantified with pseudotime (the geodesic distance from the root) along a particular trajectory can serve as a significant prognostic factor, not redundant with gene expression-based predictors. We conclude that application of the elastic principal trees to transcriptomic data can be of interest for clinical applications. Alexander Chervov, Andrei Yu. Zinovyev |
IJCNN | 2 |
| 2020 | Local intrinsic dimensionality estimators based on concentration of measureabstractIntrinsic dimensionality (ID) is one of the most fundamental characteristics of multi-dimensional data point clouds. Knowing ID is crucial to choose the appropriate machine learning approach as well as to understand its behavior and validate it. ID can be computed globally for the whole data point distribution, or computed locally in different regions of the data space. In this paper, we introduce new local estimators of ID based on linear separability of multi-dimensional data point clouds, which is one of the manifestations of concentration of measure. We empirically study the properties of these estimators and compare them with other recently introduced ID estimators exploiting various effects of measure concentration. Observed differences between estimators can be used to anticipate their behaviour in practical applications. Jonathan Bac, Andrei Yu. Zinovyev |
IJCNN | 2 |
| 2020 | cd2sbgnml: bidirectional conversion between CellDesigner and SBGN formatsabstractMOTIVATION: CellDesigner is a well-established biological map editor used in many large-scale scientific efforts. However, the interoperability between the Systems Biology Graphical Notation (SBGN) Markup Language (SBGN-ML) and the CellDesigner's proprietary Systems Biology Markup Language (SBML) extension formats remains a challenge due to the proprietary extensions used in CellDesigner files. RESULTS: We introduce a library named cd2sbgnml and an associated web service for bidirectional conversion between CellDesigner's proprietary SBML extension and SBGN-ML formats. We discuss the functionality of the cd2sbgnml converter, which was successfully used for the translation of comprehensive large-scale diagrams such as the RECON Human Metabolic network and the complete Atlas of Cancer Signalling Network, from the CellDesigner file format into SBGN-ML. AVAILABILITY AND IMPLEMENTATION: The cd2sbgnml conversion library and the web service were developed in Java, and distributed under the GNU Lesser General Public License v3.0. The sources along with a set of examples are available on GitHub (https://github.com/sbgn/cd2sbgnml and https://github.com/sbgn/cd2sbgnml-webservice, respectively). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Irina Balaur, Ludovic Roy, Alexander Mazein, S. Gökberk Karaca, Ugur Dogrusoz, Emmanuel Barillot, Andrei Yu. Zinovyev |
Bioinform. | 7 |
| 2020 | cd2sbgnml: bidirectional conversion between CellDesigner and SBGN formatsabstractBioinformatics (2019) doi: 10.1093/bioinformatics/btz969. The following funding acknowledgement was omitted from the above article: This work was also supported by the Innovative Medicines Initiative Joint Undertaking under grant agreement no. IMI 115446 (eTRIKS) to Charles Auffray and Rudi Balling, resources of which are composed of financial contributions from the European Union’s Seventh Framework Programme (2007-2013) and EFPIA companies. This has now been corrected. Irina Balaur, Ludovic Roy, Alexander Mazein, S. Gökberk Karaca, Ugur Dogrusoz, Emmanuel Barillot, Andrei Yu. Zinovyev |
Bioinform. | 7 |
| 2020 | Exact solving and sensitivity analysis of stochastic continuous time Boolean modelsabstractBACKGROUND: Solutions to stochastic Boolean models are usually estimated by Monte Carlo simulations, but as the state space of these models can be enormous, there is an inherent uncertainty about the accuracy of Monte Carlo estimates and whether simulations have reached all attractors. Moreover, these models have timescale parameters (transition rates) that the probability values of stationary solutions depend on in complex ways, raising the necessity of parameter sensitivity analysis. We address these two issues by an exact calculation method for this class of models. RESULTS: We show that the stationary probability values of the attractors of stochastic (asynchronous) continuous time Boolean models can be exactly calculated. The calculation does not require Monte Carlo simulations, instead it uses graph theoretical and matrix calculation methods previously applied in the context of chemical kinetics. In this version of the asynchronous updating framework the states of a logical model define a continuous time Markov chain and for a given initial condition the stationary solution is fully defined by the right and left nullspace of the master equation's kinetic matrix. We use topological sorting of the state transition graph and the dependencies between the nullspaces and the kinetic matrix to derive the stationary solution without simulations. We apply this calculation to several published Boolean models to analyze the under-explored question of the effect of transition rates on the stationary solutions and show they can be sensitive to parameter changes. The analysis distinguishes processes robust or, alternatively, sensitive to parameter values, providing both methodological and biological insights. CONCLUSION: Up to an intermediate size (the biggest model analyzed is 23 nodes) stochastic Boolean models can be efficiently solved by an exact matrix method, without using Monte Carlo simulations. Sensitivity analysis with respect to the model's timescale parameters often reveals a small subset of all parameters that primarily determine the stationary probability of attractor states. Mihály Koltai, Vincent Noel, Andrei Yu. Zinovyev, Laurence Calzone, Emmanuel Barillot |
BMC Bioinform. | 3 |
| 2020 | Collective intelligence defines biological functions in Wikipedia as communities in the hidden protein connection networkabstractEnglish Wikipedia, containing more than five millions articles, has approximately eleven thousands web pages devoted to proteins or genes most of which were generated by the Gene Wiki project. These pages contain information about interactions between proteins and their functional relationships. At the same time, they are interconnected with other Wikipedia pages describing biological functions, diseases, drugs and other topics curated by independent, not coordinated collective efforts. Therefore, Wikipedia contains a directed network of protein functional relations or physical interactions embedded into the global network of the encyclopedia terms, which defines hidden (indirect) functional proximity between proteins. We applied the recently developed reduced Google Matrix (REGOMAX) algorithm in order to extract the network of hidden functional connections between proteins in Wikipedia. In this network we discovered tight communities which reflect areas of interest in molecular biology or medicine and can be considered as definitions of biological functions shaped by collective intelligence. Moreover, by comparing two snapshots of Wikipedia graph (from years 2013 and 2017), we studied the evolution of the network of direct and hidden protein connections. We concluded that the hidden connections are more dynamic compared to the direct ones and that the size of the hidden interaction communities grows with time. We recapitulate the results of Wikipedia protein community analysis and annotation in the form of an interactive online map, which can serve as a portal to the Gene Wiki project. Andrei Yu. Zinovyev, Urszula Czerwinska, Laura Cantini, Emmanuel Barillot, Klaus M. Frahm, Dima Shepelyansky |
PLoS Comput. Biol. | 1 |
| 2019 | Synthesis of Boolean Networks from Biological Dynamical Constraints using Answer-Set ProgrammingabstractBoolean networks model finite discrete dynamical systems with complex behaviours. The state of each component is determined by a Boolean function of the state of (a subset of) the components of the network. This paper addresses the synthesis of these Boolean functions from constraints on their domain and emerging dynamical properties of the resulting network. The dynamical properties relate to the existence and absence of trajectories between partially observed configurations, and to the stable behaviours (fixpoints and cyclic attractors). The synthesis is expressed as a Boolean satisfiability problem relying on Answer-Set Programming with a parametrized complexity, and leads to a complete non-redundant characterization of the set of solutions. Considered constraints are particularly suited to address the synthesis of models of cellular differentiation processes, as illustrated on a case study. The scalability of the approach is demonstrated on random networks with scale-free structures up to 100 to 1,000 nodes depending on the type of constraints. Stéphanie Chevalier, Christine Froidevaux, Loïc Paulevé, Andrei Yu. Zinovyev |
ICTAI | 4 |
| 2019 | Estimating the effective dimension of large biological datasets using Fisher separability analysisabstractModern large-scale datasets are frequently said to be high-dimensional. However, their data point clouds frequently possess structures, significantly decreasing their intrinsic dimensionality (ID) due to the presence of clusters, points being located close to low-dimensional varieties or fine-grained lumping. We introduce and test a dimensionality estimator, based on analysing the separability properties of data points, on several benchmarks and real biological datasets. We show that the introduced measure of ID has performance competitive with state-of-the-art measures, being efficient across a wide range of dimensions and performing better in the case of noisy samples. Moreover, it allows estimating the intrinsic dimension in situations where the intrinsic manifold assumption is not valid. Luca Albergante, Jonathan Bac, Andrei Yu. Zinovyev |
IJCNN | 3 |
| 2019 | Application of Atlas of Cancer Signalling Network in preclinical studiesabstractCancer initiation and progression are associated with multiple molecular mechanisms. The knowledge of these mechanisms is expanding and should be converted into guidelines for tackling the disease. Here, we discuss the formalization of biological knowledge into a comprehensive resource: the Atlas of Cancer Signalling Network (ACSN) and the Google Maps-based tool NaviCell, which supports map navigation. The application of ACSN for omics data visualization, in the context of signalling maps, is possible via the NaviCell Web Service module and through the NaviCom tool. It allows generation of network-based molecular portraits of cancer using multilevel omics data. We review how these resources and tools are applied for cancer preclinical studies. Structural analysis of the maps together with omics data helps to rationalize the synergistic effects of drugs and allows design of complex disease stage-specific druggable interventions. The use of ACSN modules and maps as signatures of biological functions can help in cancer data analysis and interpretation. In addition, they empowered finding of associations between perturbations in particular molecular mechanisms and the risk to develop a specific type of cancer. These approaches are helpful, among others, to study the interplay between molecular mechanisms of cancer. It opens an opportunity to decipher how gene interactions govern the hallmarks of cancer in specific contexts. We discuss a perspective to develop a flexible methodology and a pipeline to enable systematic omics data analysis in the context of signalling network maps, for stratifying patients and suggesting interventions points and drug repositioning in cancer and other diseases. L. Cristobal Monraz Gomez, Maria Kondratova, Jean-Marie Ravel, Emmanuel Barillot, Andrei Yu. Zinovyev, Inna Kuperstein |
Briefings Bioinform. | 5 |
| 2019 | Conceptual and computational framework for logical modelling of biological networks deregulated in diseasesabstractMathematical models can serve as a tool to formalize biological knowledge from diverse sources, to investigate biological questions in a formal way, to test experimental hypotheses, to predict the effect of perturbations and to identify underlying mechanisms. We present a pipeline of computational tools that performs a series of analyses to explore a logical model's properties. A logical model of initiation of the metastatic process in cancer is used as a transversal example. We start by analysing the structure of the interaction network constructed from the literature or existing databases. Next, we show how to translate this network into a mathematical object, specifically a logical model, and how robustness analyses can be applied to it. We explore the visualization of the stable states, defined as specific attractors of the model, and match them to cellular fates or biological read-outs. With the different tools we present here, we explain how to assign to each solution of the model a probability and how to identify genetic interactions using mutant phenotype probabilities. Finally, we connect the model to relevant experimental data: we present how some data analyses can direct the construction of the network, and how the solutions of a mathematical model can also be compared with experimental data, with a particular focus on high-throughput data in cancer biology. A step-by-step tutorial is provided as a Supplementary Material and all models, tools and scripts are provided on an accompanying website: https://github.com/sysbio-curie/Logical_modelling_pipeline. Arnau Montagud, Pauline Traynard, Loredana Martignetti, Eric Bonnet, Emmanuel Barillot, Andrei Yu. Zinovyev, Laurence Calzone |
Briefings Bioinform. | 6 |
| 2019 | Community-driven roadmap for integrated disease mapsabstractThe Disease Maps Project builds on a network of scientific and clinical groups that exchange best practices, share information and develop systems biomedicine tools. The project aims for an integrated, highly curated and user-friendly platform for disease-related knowledge. The primary focus of disease maps is on interconnected signaling, metabolic and gene regulatory network pathways represented in standard formats. The involvement of domain experts ensures that the key disease hallmarks are covered and relevant, up-to-date knowledge is adequately represented. Expert-curated and computer readable, disease maps may serve as a compendium of knowledge, allow for data-supported hypothesis generation or serve as a scaffold for the generation of predictive mathematical models. This article summarizes the 2nd Disease Maps Community meeting, highlighting its important topics and outcomes. We outline milestones on the roadmap for the future development of disease maps, including creating and maintaining standardized disease maps; sharing parts of maps that encode common human disease mechanisms; providing technical solutions for complexity management of maps; and Web tools for in-depth exploration of such maps. A dedicated discussion was focused on mathematical modeling approaches, as one of the main goals of disease map development is the generation of mathematically interpretable representations to predict disease comorbidity or drug response and to suggest drug repositioning, altogether supporting clinical decisions. Marek Ostaszewski, Stephan Gebel, Inna Kuperstein, Alexander Mazein, Andrei Yu. Zinovyev, Ugur Dogrusoz, Jan Hasenauer, Ronan M. T. Fleming, Nicolas Le Novère, Piotr Gawron, Thomas S. Ligon, Anna Niarakis, David P. Nickerson, Daniel Weindl, Rudi Balling, Emmanuel Barillot, Charles Auffray, Reinhard Schneider 0002 |
Briefings Bioinform. | 5 |
| 2019 | Assessing reproducibility of matrix factorization methods in independent transcriptomesabstractMOTIVATION: Matrix factorization (MF) methods are widely used in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). MF algorithms have never been compared based on the between-datasets reproducibility of their outputs in similar independent datasets. Lack of this knowledge might have a crucial impact when generalizing the predictions made in a study to others. RESULTS: We systematically test widely used MF methods on several transcriptomic datasets collected from the same cancer type (14 colorectal, 8 breast and 4 ovarian cancer transcriptomic datasets). Inspired by concepts of evolutionary bioinformatics, we design a novel framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the MF methods for their ability to produce generalizable components. We show that a particular protocol of application of independent component analysis (ICA), accompanied by a stabilization procedure, leads to a significant increase in the between-datasets reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other standard methods. We developed a user-friendly tool for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors associated to biological processes or to technological artifacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping. AVAILABILITY AND IMPLEMENTATION: The RBH construction tool is available from http://goo.gl/DzpwYp. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Laura Cantini, Ulykbek Kairov, Aurélien de Reyniès, Emmanuel Barillot, François Radvanyi, Andrei Yu. Zinovyev |
Bioinform. | 6 |
| 2019 | PhysiBoSS: a multi-scale agent-based modelling framework integrating physical dimension and cell signallingabstractMOTIVATION: Due to the complexity and heterogeneity of multicellular biological systems, mathematical models that take into account cell signalling, cell population behaviour and the extracellular environment are particularly helpful. We present PhysiBoSS, an open source software which combines intracellular signalling using Boolean modelling (MaBoSS) and multicellular behaviour using agent-based modelling (PhysiCell). RESULTS: PhysiBoSS provides a flexible and computationally efficient framework to explore the effect of environmental and genetic alterations of individual cells at the population level, bridging the critical gap from single-cell genotype to single-cell phenotype and emergent multicellular behaviour. PhysiBoSS thus becomes very useful when studying heterogeneous population response to treatment, mutation effects, different modes of invasion or isomorphic morphogenesis events. To concretely illustrate a potential use of PhysiBoSS, we studied heterogeneous cell fate decisions in response to TNF treatment. We explored the effect of different treatments and the behaviour of several resistant mutants. We highlighted the importance of spatial information on the population dynamics by considering the effect of competition for resources like oxygen. AVAILABILITY AND IMPLEMENTATION: PhysiBoSS is freely available on GitHub (https://github.com/sysbio-curie/PhysiBoSS), with a Docker image (https://hub.docker.com/r/gletort/physiboss/). It is distributed as open source under the BSD 3-clause license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gaëlle Letort, Arnau Montagud, Gautier Stoll, Randy W. Heiland, Emmanuel Barillot, Paul Macklin, Andrei Yu. Zinovyev, Laurence Calzone |
Bioinform. | 7 |
| 2019 | Metabolic and signalling network maps integration: application to cross-talk studies and omics data analysis in cancerabstractBACKGROUND: The interplay between metabolic processes and signalling pathways remains poorly understood. Global, detailed and comprehensive reconstructions of human metabolism and signalling pathways exist in the form of molecular maps, but they have never been integrated together. We aim at filling in this gap by integrating of both signalling and metabolic pathways allowing a visual exploration of multi-level omics data and study of cross-regulatory circuits between these processes in health and in disease. RESULTS: We combined two comprehensive manually curated network maps. Atlas of Cancer Signalling Network (ACSN), containing mechanisms frequently implicated in cancer; and ReconMap 2.0, a comprehensive reconstruction of human metabolic network. We linked ACSN and ReconMap 2.0 maps via common players and represented the two maps as interconnected layers using the NaviCell platform for maps exploration ( https://navicell.curie.fr/pages/maps_ReconMap%202.html ). In addition, proteins catalysing metabolic reactions in ReconMap 2.0 were not previously visually represented on the map canvas. This precluded visualisation of omics data in the context of ReconMap 2.0. We suggested a solution for displaying protein nodes on the ReconMap 2.0 map in the vicinity of the corresponding reaction or process nodes. This permits multi-omics data visualisation in the context of both map layers. Exploration and shuttling between the two map layers is possible using Google Maps-like features of NaviCell. The integrated networks ACSN-ReconMap 2.0 are accessible online and allows data visualisation through various modes such as markers, heat maps, bar-plots, glyphs and map staining. The integrated networks were applied for comparison of immunoreactive and proliferative ovarian cancer subtypes using transcriptomic, copy number and mutation multi-omics data. A certain number of metabolic and signalling processes specifically deregulated in each of the ovarian cancer sub-types were identified. CONCLUSIONS: As knowledge evolves and new omics data becomes more heterogeneous, gathering together existing domains of biology under common platforms is essential. We believe that an integrated ACSN-ReconMap 2.0 networks will help in understanding various disease mechanisms and discovery of new interactions at the intersection of cell signalling and metabolism. In addition, the successful integration of metabolic and signalling networks allows broader systems biology approach application for data interpretation and retrieval of intervention points to tackle simultaneously the key players coordinating signalling and metabolism in human diseases. Nicolas Sompairac, Jennifer Modamio, Emmanuel Barillot, Ronan M. T. Fleming, Andrei Yu. Zinovyev, Inna Kuperstein |
BMC Bioinform. | 5 |
| 2018 | Data analysis with arbitrary error measures approximated by piece-wise quadratic PQSQ functionsabstractDefining an error function (a measure of deviation of a model prediction from the data) is a critical step in any optimization-based data analysis method, including regression, clustering and dimension reduction. Usual quadratic error function in case of real-life high-dimensional and noisy data suffers from non-robustness to presence of outliers. Therefore, using non-quadratic error functions in data analysis and machine learning (such as L1 norm-based) is an active field of modern research but the majority of methods suggested are either slow or imprecise (use arbitrary heuristics). We suggest a flexible and highly performant approach to generalize most of existing data analysis methods to an arbitrary error function of subquadratic growth. For this purpose, we exploit PQSQ functions (piece-wise quadratic of subquadratic growth), which can be minimized by a simple and fast splitting-based iterative algorithm. The theoretical basis of the PQSQ approach is an application of min-plus (idempotent) algebra to data approximation. We introduce the general idea of the approach and illustrate it on four standard tools of machine learning: simple regression, regularized regression, k-mean clustering and principal component analysis. In all cases, PQSQ-based methods achieve better robustness with respect to the presence of strong noise in the data compared to the standard methods. Alexander N. Gorban, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev |
IJCNN | 3 |
| 2017 | MaBoSS 2.0: an environment for stochastic Boolean modelingabstractMOTIVATION: Modeling of signaling pathways is an important step towards the understanding and the treatment of diseases such as cancers, HIV or auto-immune diseases. MaBoSS is a software that allows to simulate populations of cells and to model stochastically the intracellular mechanisms that are deregulated in diseases. MaBoSS provides an output of a Boolean model in the form of time-dependent probabilities, for all biological entities (genes, proteins, phenotypes, etc.) of the model. RESULTS: We present a new version of MaBoSS (2.0), including an updated version of the core software and an environment. With this environment, the needs for modeling signaling pathways are facilitated, including model construction, visualization, simulations of mutations, drug treatments and sensitivity analyses. It offers a framework for automated production of theoretical predictions. AVAILABILITY AND IMPLEMENTATION: MaBoSS software can be found at https://maboss.curie.fr , including tutorials on existing models and examples of models. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gautier Stoll, Barthélémy Caron, Eric Viara, Aurélien Dugourd, Andrei Yu. Zinovyev, Aurélien Naldi, Guido Kroemer, Emmanuel Barillot, Laurence Calzone |
Bioinform. | 5 |
| 2017 | NetNorM: Capturing cancer-relevant information in somatic exome mutation data with gene networks for cancer stratification and prognosisabstractGenome-wide somatic mutation profiles of tumours can now be assessed efficiently and promise to move precision medicine forward. Statistical analysis of mutation profiles is however challenging due to the low frequency of most mutations, the varying mutation rates across tumours, and the presence of a majority of passenger events that hide the contribution of driver events. Here we propose a method, NetNorM, to represent whole-exome somatic mutation data in a form that enhances cancer-relevant information using a gene network as background knowledge. We evaluate its relevance for two tasks: survival prediction and unsupervised patient stratification. Using data from 8 cancer types from The Cancer Genome Atlas (TCGA), we show that it improves over the raw binary mutation data and network diffusion for these two tasks. In doing so, we also provide a thorough assessment of somatic mutations prognostic power which has been overlooked by previous studies because of the sparse and binary nature of mutations. Marine Le Morvan, Andrei Yu. Zinovyev, Jean-Philippe Vert |
PLoS Comput. Biol. | 2 |
| 2016 | Piece-wise quadratic approximations of arbitrary error functions for fast and robust machine learning
Alexander N. Gorban, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev |
Neural Networks | 3 |
| 2015 | Fast and user-friendly non-linear principal manifold learning by method of elastic mapsabstractMethod of elastic maps allows fast learning of nonlinear principal manifolds for large datasets. We present user-friendly implementation of the method in ViDaExpert software. Equipped with several dialogs for configuring data point representations (size, shape, color) and fast 3D viewer, ViDaExpert is a handy tool allowing to construct an interactive 3D-scene representing a table of data in multidimensional space and perform its quick and insightfull statistical analysis, from basic to advanced methods. We list several recent application examples of manifold learning by method of elastic maps in various fields of life sciences. Alexander N. Gorban, Andrei Yu. Zinovyev |
DSAA | 2 |
| 2015 | Mathematical Modelling of Molecular Pathways Enabling Tumour Cell Invasion and MigrationabstractUnderstanding the etiology of metastasis is very important in clinical perspective, since it is estimated that metastasis accounts for 90% of cancer patient mortality. Metastasis results from a sequence of multiple steps including invasion and migration. The early stages of metastasis are tightly controlled in normal cells and can be drastically affected by malignant mutations; therefore, they might constitute the principal determinants of the overall metastatic rate even if the later stages take long to occur. To elucidate the role of individual mutations or their combinations affecting the metastatic development, a logical model has been constructed that recapitulates published experimental results of known gene perturbations on local invasion and migration processes, and predict the effect of not yet experimentally assessed mutations. The model has been validated using experimental data on transcriptome dynamics following TGF-β-dependent induction of Epithelial to Mesenchymal Transition in lung cancer cell lines. A method to associate gene expression profiles with different stable state solutions of the logical model has been developed for that purpose. In addition, we have systematically predicted alleviating (masking) and synergistic pairwise genetic interactions between the genes composing the model with respect to the probability of acquiring the metastatic phenotype. We focused on several unexpected synergistic genetic interactions leading to theoretically very high metastasis probability. Among them, the synergistic combination of Notch overexpression and p53 deletion shows one of the strongest effects, which is in agreement with a recent published experiment in a mouse model of gut cancer. The mathematical model can recapitulate experimental mutations in both cell line and mouse models. Furthermore, the model predicts new gene perturbations that affect the early steps of metastasis underlying potential intervention points for innovative therapeutic strategies in oncology. David P. A. Cohen, Loredana Martignetti, Sylvie Robine, Emmanuel Barillot, Andrei Yu. Zinovyev, Laurence Calzone |
PLoS Comput. Biol. | 5 |
| 2013 | OCSANA: optimal combinations of interventions from network analysisabstractUNLABELLED: Targeted therapies interfering with specifically one protein activity are promising strategies in the treatment of diseases like cancer. However, accumulated empirical experience has shown that targeting multiple proteins in signaling networks involved in the disease is often necessary. Thus, one important problem in biomedical research is the design and prioritization of optimal combinations of interventions to repress a pathological behavior, while minimizing side-effects. OCSANA (optimal combinations of interventions from network analysis) is a new software designed to identify and prioritize optimal and minimal combinations of interventions to disrupt the paths between source nodes and target nodes. When specified by the user, OCSANA seeks to additionally minimize the side effects that a combination of interventions can cause on specified off-target nodes. With the crucial ability to cope with very large networks, OCSANA includes an exact solution and a novel selective enumeration approach for the combinatorial interventions' problem. AVAILABILITY: The latest version of OCSANA, implemented as a plugin for Cytoscape and distributed under LGPL license, is available together with source code at http://bioinfo.curie.fr/projects/ocsana. Paola Vera-Licona, Eric Bonnet, Emmanuel Barillot, Andrei Yu. Zinovyev |
Bioinform. | 4 |
| 2013 | Synthetic Lethality between Gene Defects Affecting a Single Non-essential Molecular Pathway with Reversible StepsabstractSystematic analysis of synthetic lethality (SL) constitutes a critical tool for systems biology to decipher molecular pathways. The most accepted mechanistic explanation of SL is that the two genes function in parallel, mutually compensatory pathways, known as between-pathway SL. However, recent genome-wide analyses in yeast identified a significant number of within-pathway negative genetic interactions. The molecular mechanisms leading to within-pathway SL are not fully understood. Here, we propose a novel mechanism leading to within-pathway SL involving two genes functioning in a single non-essential pathway. This type of SL termed within-reversible-pathway SL involves reversible pathway steps, catalyzed by different enzymes in the forward and backward directions, and kinetic trapping of a potentially toxic intermediate. Experimental data with recombinational DNA repair genes validate the concept. Mathematical modeling recapitulates the possibility of kinetic trapping and revealed the potential contributions of synthetic, dosage-lethal interactions in such a genetic system as well as the possibility of within-pathway positive masking interactions. Analysis of yeast gene interaction and pathway data suggests broad applicability of this novel concept. These observations extend the canonical interpretation of synthetic-lethal or synthetic-sick interactions with direct implications to reconstruct molecular pathways and improve therapeutic approaches to diseases such as cancer. Andrei Yu. Zinovyev, Inna Kuperstein, Emmanuel Barillot, Wolf-Dietrich Heyer |
PLoS Comput. Biol. | 1 |
| 2011 | Control-free calling of copy number alterations in deep-sequencing data using GC-content normalizationabstractSUMMARY: We present a tool for control-free copy number alteration (CNA) detection using deep-sequencing data, particularly useful for cancer studies. The tool deals with two frequent problems in the analysis of cancer deep-sequencing data: absence of control sample and possible polyploidy of cancer cells. FREEC (control-FREE Copy number caller) automatically normalizes and segments copy number profiles (CNPs) and calls CNAs. If ploidy is known, FREEC assigns absolute copy number to each predicted CNA. To normalize raw CNPs, the user can provide a control dataset if available; otherwise GC content is used. We demonstrate that for Illumina single-end, mate-pair or paired-end sequencing, GC-contentr normalization provides smooth profiles that can be further segmented and analyzed in order to predict CNAs. AVAILABILITY: Source code and sample data are available at http://bioinfo-out.curie.fr/projects/freec/. Valentina Boeva, Andrei Yu. Zinovyev, Kevin Bleakley, Jean-Philippe Vert, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot |
Bioinform. | 2 |
| 2010 | Principal Manifolds and Graphs in Practice: from Molecular Biology to Dynamical SystemsabstractWe present several applications of non-linear data modeling, using principal manifolds and principal graphs constructed using the metaphor of elasticity (elastic principal graph approach). These approaches are generalizations of the Kohonen's self-organizing maps, a class of artificial neural networks. On several examples we show advantages of using non-linear objects for data approximation in comparison to the linear ones. We propose four numerical criteria for comparing linear and non-linear mappings of datasets into the spaces of lower dimension. The examples are taken from comparative political science, from analysis of high-throughput data in molecular biology, from analysis of dynamical systems. Alexander N. Gorban, Andrei Yu. Zinovyev |
Int. J. Neural Syst. | 2 |
| 2010 | Mathematical Modelling of Cell-Fate Decision in Response to Death Receptor EngagementabstractCytokines such as TNF and FASL can trigger death or survival depending on cell lines and cellular conditions. The mechanistic details of how a cell chooses among these cell fates are still unclear. The understanding of these processes is important since they are altered in many diseases, including cancer and AIDS. Using a discrete modelling formalism, we present a mathematical model of cell fate decision recapitulating and integrating the most consistent facts extracted from the literature. This model provides a generic high-level view of the interplays between NFkappaB pro-survival pathway, RIP1-dependent necrosis, and the apoptosis pathway in response to death receptor-mediated signals. Wild type simulations demonstrate robust segregation of cellular responses to receptor engagement. Model simulations recapitulate documented phenotypes of protein knockdowns and enable the prediction of the effects of novel knockdowns. In silico experiments simulate the outcomes following ligand removal at different stages, and suggest experimental approaches to further validate and specialise the model for particular cell types. We also propose a reduced conceptual model implementing the logic of the decision process. This analysis gives specific predictions regarding cross-talks between the three pathways, as well as the transient role of RIP1 protein in necrosis, and confirms the phenotypes of novel perturbations. Our wild type and mutant simulations provide novel insights to restore apoptosis in defective cells. The model analysis expands our understanding of how cell fate decision is made. Moreover, our current model can be used to assess contradictory or controversial data from the literature. Ultimately, it constitutes a valuable reasoning tool to delineate novel experiments. Laurence Calzone, Laurent Tournier, Simon Fourquet, Denis Thieffry, Boris Zhivotovsky, Emmanuel Barillot, Andrei Yu. Zinovyev |
PLoS Comput. Biol. | 7 |
| 2008 | BiNoM: a Cytoscape plugin for manipulating and analyzing biological networksabstractUNLABELLED: BiNoM (Biological Network Manager) is a new bioinformatics software that significantly facilitates the usage and the analysis of biological networks in standard systems biology formats (SBML, SBGN, BioPAX). BiNoM implements a full-featured BioPAX editor and a method of 'interfaces' for accessing BioPAX content. BiNoM is able to work with huge BioPAX files such as whole pathway databases. In addition, BiNoM allows the analysis of networks created with CellDesigner software and their conversion into BioPAX format. BiNoM comes as a library and as a Cytoscape plugin which adds a rich set of operations to Cytoscape such as path and cycle analysis, clustering sub-networks, decomposition of network into modules, clipboard operations and others. AVAILABILITY: Last version of BiNoM distributed under the LGPL licence together with documentation, source code and API are available at http://bioinfo.curie.fr/projects/binom Andrei Yu. Zinovyev, Eric Viara, Laurence Calzone, Emmanuel Barillot |
Bioinform. | 1 |
| 2007 | Branching Principal Components: Elastic Graphs, Topological Grammars and Metro MapsabstractTo approximate complex data, we propose new type of low-dimensional "principal object":principal cubic complex. This complex is a generalization of linear and nonlinear principal manifolds and includes them as a particular case. To construct such an object, we combine the method oftopological grammarswith the minimization of elastic energy defined for its embedment into multidimensional data space. The whole complex is presented as a system of nodes and springs and as a product of one-dimensional continua (represented by graphs), and the grammars describe how these continua transform during the process of optimal complex construction. The simplest case of a topological grammar ("add a node or bisect an edge") produces "principal trees" that are useful in many practical applications. We demonstrate how this can be applied to the analysis of bacterial genomes and for visualization of microarray data using "metro map" visual representation. Alexander N. Gorban, Neil R. Sumner, Andrei Yu. Zinovyev |
IJCNN | 3 |
| 2007 | Classification of microarray data using gene networksabstractBACKGROUND: Microarrays have become extremely useful for analysing genetic phenomena, but establishing a relation between microarray analysis results (typically a list of genes) and their biological significance is often difficult. Currently, the standard approach is to map a posteriori the results onto gene networks in order to elucidate the functions perturbed at the level of pathways. However, integrating a priori knowledge of the gene networks could help in the statistical analysis of gene expression data and in their biological interpretation. RESULTS: We propose a method to integrate a priori the knowledge of a gene network in the analysis of gene expression data. The approach is based on the spectral decomposition of gene expression profiles with respect to the eigenfunctions of the graph, resulting in an attenuation of the high-frequency components of the expression profiles with respect to the topology of the graph. We show how to derive unsupervised and supervised classification algorithms of expression profiles, resulting in classifiers with biological relevance. We illustrate the method with the analysis of a set of expression profiles from irradiated and non-irradiated yeast strains. CONCLUSION: Including a priori knowledge of a gene network for the analysis of gene expression data leads to good classification performance and improved interpretability of the results. Franck Rapaport, Andrei Yu. Zinovyev, Marie Dutreix, Emmanuel Barillot, Jean-Philippe Vert |
BMC Bioinform. | 2 |
| 2003 | Application of the method of elastic maps in analysis of genetic textsabstractMethod of elastic maps allows to construct efficiently 1D, 2D and 3D nonlinear approximations to the principal manifolds with different topology (piece of plane, sphere, torus etc.) and to project data onto it. We describe the idea of the method and demonstrate its applications in analysis of genetic sequences. Alexander N. Gorban, Andrei Yu. Zinovyev, Donald C. Wunsch II |
IJCNN | 2 |
| 2003 | Codon adaptation index as a measure of dominating codon biasabstractUNLABELLED: We propose a simple algorithm to detect dominating synonymous codon usage bias in genomes. The algorithm is based on a precise mathematical formulation of the problem that lead us to use the Codon Adaptation Index (CAI) as a 'universal' measure of codon bias. This measure has been previously employed in the specific context of translational bias. With the set of coding sequences as a sole source of biological information, the algorithm provides a reference set of genes which is highly representative of the bias. This set can be used to compute the CAI of genes of prokaryotic and eukaryotic organisms, including those whose functional annotation is not yet available. An important application concerns the detection of a reference set characterizing translational bias which is known to correlate to expression levels; in this case, the algorithm becomes a key tool to predict gene expression levels, to guide regulatory circuit reconstruction, and to compare species. The algorithm detects also leading-lagging strands bias, GC-content bias, GC3 bias, and horizontal gene transfer. The approach is validated on 12 slow-growing and fast-growing bacteria, Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. AVAILABILITY: http://www.ihes.fr/~materials. Alessandra Carbone, Andrei Yu. Zinovyev, François Képès |
Bioinform. | 2 |