EDBT 2026 Demo / reviewers in the wild / expert
Julio Saez-Rodriguez
dblp:35/1131 · also Julio Sáez-Rodríguez
· DBLP profile ↗
49ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-8552-8976ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 46 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CIRCE: a scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility dataabstractMOTIVATION: Chromatin 3D folding creates numerous DNA interactions, participating in gene expression regulation. Single-cell chromatin-accessibility assays now profile hundreds of thousands of cells, challenging existing methods for mapping cis-regulatory interactions. RESULTS: We present CIRCE, a fast and scalable Python package to predict cis-regulatory DNA interactions from single-cell chromatin accessibility data. CIRCE re-implements the Cicero workflow to analyse single-cell atlases, cutting runtime and memory use by several orders of magnitude. We also provide new options to compute metacells, grouping similar cells to reduce data sparsity. We benchmarked CIRCE against Cicero on two datasets of different sizes and demonstrated the improvement from CIRCE's metacells' strategy with promoter capture Hi-C data. We also evaluated how DNA interaction predictions are impacted by different pre-processing. We observed a negative impact of Cicero's count normalization, and the best performance was obtained with the single-cell count matrix directly. Finally, we demonstrated the scalability of CIRCE by processing a dataset of more than 700 000 cells and 1 million DNA regions in less than an hour. CIRCE should greatly facilitate the prediction of DNA region interactions for scverse and Python users, while providing new and up-to-date pre-processing insights. AVAILABILITY AND IMPLEMENTATION: CIRCE is released as an open-source software under the AGPL-3.0 licence. The package source code is available on GitHub at https://github.com/cantinilab/CIRCE, and its documentation is accessible at https://circe.readthedocs.io. The code to reproduce the presented results is available as a Snakemake pipeline at https://github.com/cantinilab/circe_reproducibility.s. Remi Trimbour, Julio Saez-Rodriguez, Laura Cantini |
Bioinform. | 2 |
| 2025 | NetworkCommons: bridging data, knowledge, and methods to build and evaluate context-specific biological networksabstractSUMMARY: We present NetworkCommons, a platform for integrating prior knowledge, omics data, and network inference methods, facilitating their usage and evaluation. NetworkCommons aims to be an infrastructure for the network biology community that supports the development of better methods and benchmarks, by enhancing interoperability and integration. AVAILABILITY AND IMPLEMENTATION: NetworkCommons is implemented in Python and offers programmatic access to multiple omics datasets, network inference methods, and benchmarking setups. It is a free software, available at https://github.com/saezlab/networkcommons, and deposited in Zenodo at https://doi.org/10.5281/zenodo.14719118. Victor Paton, Dénes Türei, Olga Ivanova, Sophia Müller-Dott, Pablo Rodríguez-Mier, Veronica Venafra, Livia Perfetto, Martín Garrido-Rodriguez, Julio Saez-Rodriguez |
Bioinform. | 9 |
| 2025 | RIDDEN: Data-driven inference of receptor activity from transcriptomic dataabstractIntracellular signaling initiated from ligand-bound receptors plays a fundamental role in both physiological regulation and development of disease states, making receptors one of the most frequent drug targets. Systems level analysis of receptor activity can help to identify cell and disease type-specific receptor activity alterations. While several computational methods have been developed to analyze ligand-receptor interactions based on transcriptomics data, none of them focuses directly on the receptor side of these interactions. Also, most of the methods use directly the expression of ligands and receptors to infer active interaction, while co-expression of genes does not necessarily indicate functional interactions or activated state. To address these problems, we developed RIDDEN (Receptor actIvity Data Driven inferENce), a computational tool, which predicts receptor activities from the receptor-regulated gene expression profiles, and not from the expressions of ligand and receptor genes. We collected 14463 perturbation gene expression profiles for 229 different receptors. Using these data, we trained the RIDDEN model, which can effectively predict receptor activity for new bulk and single-cell transcriptomics datasets. We validated RIDDEN's performance on independent in vitro and in vivo receptor perturbation data, showing that RIDDEN's model weights correspond to known regulatory interactions between receptors and transcription factors, and that predicted receptor activities correlate with receptor and ligand expressions in in vivo datasets. We also show that RIDDEN can be used to identify mechanistic biomarkers in an immune checkpoint blockade-treated cancer patient cohort. RIDDEN, the largest transcriptomics-based receptor activity inference model, can be used to identify cell populations with altered receptor activity and, in turn, foster the study of cell-cell communication using transcriptomics data. Szilvia Barsi, Eszter Varga, Daniel Dimitrov, Julio Saez-Rodriguez, László Hunyady, Bence Szalai |
PLoS Comput. Biol. | 4 |
| 2025 | NeKo: A tool for automatic network construction from prior knowledgeabstractBiological networks provide a structured framework for analyzing the dynamic interplay and interactions between molecular entities, facilitating deeper insights into cellular functions and biological processes. Network construction often requires extensive manual curation based on scientific literature and public databases, a time-consuming and laborious task. To address this challenge, we introduce NeKo, a Python package to automate the construction of biological networks by integrating and prioritizing molecular interactions from various databases. NeKo allows users to provide their molecules of interest (e.g., genes, proteins or phosphosites), select interaction resources and apply flexible strategies to build networks based on prior knowledge. Users can filter interactions by various criteria, such as direct or indirect links and signed or unsigned interactions, to tailor the network to their needs and downstream analysis. We demonstrate some of NeKo's capabilities in two use cases: first we construct a network based on transcriptomics from medulloblastoma; in the second, we model drug synergies. NeKo streamlines the network-building process, making it more accessible and efficient for researchers. Marco Ruscone, Eirini Tsirvouli, Andrea Checcoli, Dénes Türei, Emmanuel Barillot, Julio Saez-Rodriguez, Loredana Martignetti, Åsmund Flobak, Laurence Calzone |
PLoS Comput. Biol. | 6 |
| 2024 | MetalinksDB: a flexible and contextualizable resource of metabolite-protein interactionsabstractFrom the catalytic breakdown of nutrients to signaling, interactions between metabolites and proteins play an essential role in cellular function. An important case is cell-cell communication, where metabolites, secreted into the microenvironment, initiate signaling cascades by binding to intra- or extracellular receptors of neighboring cells. Protein-protein cell-cell communication interactions are routinely predicted from transcriptomic data. However, inferring metabolite-mediated intercellular signaling remains challenging, partially due to the limited size of intercellular prior knowledge resources focused on metabolites. Here, we leverage knowledge-graph infrastructure to integrate generalistic metabolite-protein with curated metabolite-receptor resources to create MetalinksDB. MetalinksDB is an order of magnitude larger than existing metabolite-receptor resources and can be tailored to specific biological contexts, such as diseases, pathways, or tissue/cellular locations. We demonstrate MetalinksDB's utility in identifying deregulated processes in renal cancer using multi-omics bulk data. Furthermore, we infer metabolite-driven intercellular signaling in acute kidney injury using spatial transcriptomics data. MetalinksDB is a comprehensive and customizable database of intercellular metabolite-protein interactions, accessible via a web interface (https://metalinks.omnipathdb.org/) and programmatically as a knowledge graph (https://github.com/biocypher/metalinks). We anticipate that by enabling diverse analyses tailored to specific biological contexts, MetalinksDB will facilitate the discovery of disease-relevant metabolite-mediated intercellular signaling processes. Elias B. Farr, Daniel Dimitrov, Christina Schmidt, Dénes Türei, Sebastian Lobentanzer, Aurélien Dugourd, Julio Saez-Rodriguez |
Briefings Bioinform. | 7 |
| 2024 | PhosX: data-driven kinase activity inference from phosphoproteomics experimentsabstractSUMMARY: The inference of kinase activity from phosphoproteomics data can point to causal mechanisms driving signalling processes and potential drug targets. Identifying the kinases whose change in activity explains the observed phosphorylation profiles, however, remains challenging, and constrained by the manually curated knowledge of kinase-substrate associations. Recently, experimentally determined substrate sequence specificities of human kinases have become available, but robust methods to exploit this new data for kinase activity inference are still missing. We present PhosX, a method to estimate differential kinase activity from phosphoproteomics data that combines state-of-the-art statistics in enrichment analysis with kinases' substrate sequence specificity information. Using a large phosphoproteomics dataset with known differentially regulated kinases we show that our method identifies upregulated and downregulated kinases by only relying on the input phosphopeptides' sequences and intensity changes. We find that PhosX outperforms the currently available approach for the same task, and performs better or similarly to state-of-the-art methods that rely on previously known kinase-substrate associations. We therefore recommend its use for data-driven kinase activity inference. AVAILABILITY AND IMPLEMENTATION: PhosX is implemented in Python, open-source under the Apache-2.0 licence, and distributed on the Python Package Index. The code is available on GitHub (https://github.com/alussana/phosx). Alessandro Lussana, Sophia Müller-Dott, Julio Saez-Rodriguez, Evangelia Petsalaki |
Bioinform. | 3 |
| 2024 | Multiscale networks in multiple sclerosisabstractComplex diseases such as Multiple Sclerosis (MS) cover a wide range of biological scales, from genes and proteins to cells and tissues, up to the full organism. In fact, any phenotype for an organism is dictated by the interplay among these scales. We conducted a multilayer network analysis and deep phenotyping with multi-omics data (genomics, phosphoproteomics and cytomics), brain and retinal imaging, and clinical data, obtained from a multicenter prospective cohort of 328 patients and 90 healthy controls. Multilayer networks were constructed using mutual information for topological analysis, and Boolean simulations were constructed using Pearson correlation to identified paths within and among all layers. The path more commonly found from the Boolean simulations connects protein MK03, with total T cells, the thickness of the retinal nerve fiber layer (RNFL), and the walking speed. This path contains nodes involved in protein phosphorylation, glial cell differentiation, and regulation of stress-activated MAPK cascade, among others. Specific paths identified were subsequently analyzed by flow cytometry at the single-cell level. Combinations of several proteins (GSK3AB, HSBP1 or RS6) and immune cells (Th17, Th1 non-classic, CD8, CD8 Treg, CD56 neg, and B memory) were part of the paths explaining the clinical phenotype. The advantage of the path identified from the Boolean simulations is that it connects information about these known biological pathways with the layers at higher scales (retina damage and disability). Overall, the identified paths provide a means to connect the molecular aspects of MS with the overall phenotype. Keith E. Kennedy, Nicole Kerlero de Rosbo, Antonio Uccelli, Maria Cellerino, Federico Ivaldi, Paola Contini, Raffaele De Palma, Hanne F. Harbo, Tone Berge, Steffan D. Bos, Einar A. Høgestøl, Synne Brune-Ingebretsen, Sigrid A. de Rodez Benavent, Friedemann Paul, Alexander U. Brandt, Priscilla Bäcker-Koduah, Janina Behrens, Joseph Kuchling, Susanna Asseyer, Michael Scheel, Claudia Chien, Hanna Gwendolyn Zimmermann, Seyedamirhosein Motamedi, Josef Kauer-Bonin, Julio Saez-Rodriguez, Melanie Rinas, Leonidas G. Alexopoulos, Magí Andorrà, Sara Llufriu, Albert Saiz, Yolanda Blanco, Eloy Martinez-Heras, Elisabeth Solana, Irene Pulido-Valdeolivas, Elena H. Martinez-Lapiscina, Jordi García-Ojalvo, Pablo Villoslada |
PLoS Comput. Biol. | 25 |
| 2022 | FUNKI: interactive functional footprint-based analysis of omics dataabstractMOTIVATION: Omics data are broadly used to get a snapshot of the molecular status of cells. In particular, changes in omics can be used to estimate the activity of pathways, transcription factors and kinases based on known regulated targets, that we call footprints. Then the molecular paths driving these activities can be estimated using causal reasoning on large signalling networks. RESULTS: We have developed FUNKI, a FUNctional toolKIt for footprint analysis. It provides a user-friendly interface for an easy and fast analysis of transcriptomics, phosphoproteomics and metabolomics data, either from bulk or single-cell experiments. FUNKI also features different options to visualize the results and run post-analyses, and is mirrored as a scripted version in R. AVAILABILITY AND IMPLEMENTATION: FUNKI is a free and open-source application built on R and Shiny, available at https://github.com/saezlab/ShinyFUNKI and https://saezlab.shinyapps.io/funki/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rosa D. Hernansaiz-Ballesteros, Christian H. Holland, Aurélien Dugourd, Julio Saez-Rodriguez |
Bioinform. | 4 |
| 2022 | Parallel ant colony optimization for the training of cell signaling networksabstractAcquiring a functional comprehension of the deregulation of cell signaling networks in disease allows progress in the development of new therapies and drugs. Computational models are becoming increasingly popular as a systematic tool to analyze the functioning of complex biochemical networks, such as those involved in cell signaling. CellNOpt is a framework to build predictive logic-based models of signaling pathways by training a prior knowledge network to biochemical data obtained from perturbation experiments. This training can be formulated as an optimization problem that can be solved using metaheuristics. However, the genetic algorithm used so far in CellNOpt presents limitations in terms of execution time and quality of solutions when applied to large instances. Thus, in order to overcome those issues, in this paper we propose the use of a method based on ant colony optimization, adapted to the problem at hand and parallelized using a hybrid approach. The performance of this novel method is illustrated with several challenging benchmark problems in the study of new therapies for liver cancer. Patricia González, Roberto Prado-Rodriguez, Attila Gábor, Julio Saez-Rodriguez, Julio R. Banga, Ramón Doallo |
Expert Syst. Appl. | 4 |
| 2022 | Computational drug repurposing against SARS-CoV-2 reveals plasma membrane cholesterol depletion as key factor of antiviral drug activityabstractComparing SARS-CoV-2 infection-induced gene expression signatures to drug treatment-induced gene expression signatures is a promising bioinformatic tool to repurpose existing drugs against SARS-CoV-2. The general hypothesis of signature-based drug repurposing is that drugs with inverse similarity to a disease signature can reverse disease phenotype and thus be effective against it. However, in the case of viral infection diseases, like SARS-CoV-2, infected cells also activate adaptive, antiviral pathways, so that the relationship between effective drug and disease signature can be more ambiguous. To address this question, we analysed gene expression data from in vitro SARS-CoV-2 infected cell lines, and gene expression signatures of drugs showing anti-SARS-CoV-2 activity. Our extensive functional genomic analysis showed that both infection and treatment with in vitro effective drugs leads to activation of antiviral pathways like NFkB and JAK-STAT. Based on the similarity-and not inverse similarity-between drug and infection-induced gene expression signatures, we were able to predict the in vitro antiviral activity of drugs. We also identified SREBF1/2, key regulators of lipid metabolising enzymes, as the most activated transcription factors by several in vitro effective antiviral drugs. Using a fluorescently labeled cholesterol sensor, we showed that these drugs decrease the cholesterol levels of plasma-membrane. Supplementing drug-treated cells with cholesterol reversed the in vitro antiviral effect, suggesting the depleting plasma-membrane cholesterol plays a key role in virus inhibitory mechanism. Our results can help to more effectively repurpose approved drugs against SARS-CoV-2, and also highlights key mechanisms behind their antiviral effect. Szilvia Barsi, Henrietta Papp, Alberto Valdeolivas, Dániel J. Tóth, Anett Kuczmog, Mónika Madai, László Hunyady, Péter Várnai, Julio Saez-Rodriguez, Ferenc Jakab, Bence Szalai |
PLoS Comput. Biol. | 9 |
| 2021 | Reusability and composability in process description maps: RAS-RAF-MEK-ERK signallingabstractDetailed maps of the molecular basis of the disease are powerful tools for interpreting data and building predictive models. Modularity and composability are considered necessary network features for large-scale collaborative efforts to build comprehensive molecular descriptions of disease mechanisms. An effective way to create and manage large systems is to compose multiple subsystems. Composable network components could effectively harness the contributions of many individuals and enable teams to seamlessly assemble many individual components into comprehensive maps. We examine manually built versions of the RAS-RAF-MEK-ERK cascade from the Atlas of Cancer Signalling Network, PANTHER and Reactome databases and review them in terms of their reusability and composability for assembling new disease models. We identify design principles for managing complex systems that could make it easier for investigators to share and reuse network components. We demonstrate the main challenges including incompatible levels of detail and ambiguous representation of complexes and highlight the need to address these challenges. Alexander Mazein, Adrien Rougny, Jonathan R. Karr, Julio Saez-Rodriguez, Marek Ostaszewski, Reinhard Schneider 0002 |
Briefings Bioinform. | 4 |
| 2021 | Setting the basis of best practices and standards for curation and annotation of logical models in biology - highlights of the [BC]2 2019 CoLoMoTo/SysMod WorkshopabstractThe fast accumulation of biological data calls for their integration, analysis and exploitation through more systematic approaches. The generation of novel, relevant hypotheses from this enormous quantity of data remains challenging. Logical models have long been used to answer a variety of questions regarding the dynamical behaviours of regulatory networks. As the number of published logical models increases, there is a pressing need for systematic model annotation, referencing and curation in community-supported and standardised formats. This article summarises the key topics and future directions of a meeting entitled 'Annotation and curation of computational models in biology', organised as part of the 2019 [BC]2 conference. The purpose of the meeting was to develop and drive forward a plan towards the standardised annotation of logical models, review and connect various ongoing projects of experts from different communities involved in the modelling and annotation of molecular biological entities, interactions, pathways and models. This article defines a roadmap towards the annotation and curation of logical models, including milestones for best practices and minimum standard requirements. Anna Niarakis, Martin Kuiper, Marek Ostaszewski, Rahuman S. Malik-Sheriff, Cristina Casals-Casas, Denis Thieffry, Tom C. Freeman, Paul D. Thomas, Vasundra Touré, Vincent Noel, Gautier Stoll, Julio Saez-Rodriguez, Aurélien Naldi, Eugenia Oshurko, Ioannis Xenarios, Sylvain Soliman, Claudine Chaouiya, Tomás Helikar, Laurence Calzone |
Briefings Bioinform. | 12 |
| 2021 | SysMod: the ISCB community for data-driven computational modelling and multi-scale analysis of biological systemsabstractComputational models of biological systems can exploit a broad range of rapidly developing approaches, including novel experimental approaches, bioinformatics data analysis, emerging modelling paradigms, data standards and algorithms. A discussion about the most recent advances among experts from various domains is crucial to foster data-driven computational modelling and its growing use in assessing and predicting the behaviour of biological systems. Intending to encourage the development of tools, approaches and predictive models, and to deepen our understanding of biological systems, the Community of Special Interest (COSI) was launched in Computational Modelling of Biological Systems (SysMod) in 2016. SysMod's main activity is an annual meeting at the Intelligent Systems for Molecular Biology (ISMB) conference, which brings together computer scientists, biologists, mathematicians, engineers, computational and systems biologists. In the five years since its inception, SysMod has evolved into a dynamic and expanding community, as the increasing number of contributions and participants illustrate. SysMod maintains several online resources to facilitate interaction among the community members, including an online forum, a calendar of relevant meetings and a YouTube channel with talks and lectures of interest for the modelling community. For more than half a decade, the growing interest in computational systems modelling and multi-scale data integration has inspired and supported the SysMod community. Its members get progressively more involved and actively contribute to the annual COSI meeting and several related community workshops and meetings, focusing on specific topics, including particular techniques for computational modelling or standardisation efforts. Andreas Dräger, Tomás Helikar, Matteo Barberis, Marc R. Birtwistle, Laurence Calzone, Claudine Chaouiya, Jan Hasenauer, Jonathan R. Karr, Anna Niarakis, María Rodríguez Martínez, Julio Saez-Rodriguez, Juilee Thakar |
Bioinform. | 11 |
| 2021 | The Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST)abstractMOTIVATION: A large variety of molecular interactions occurs between biomolecular components in cells. When a molecular interaction results in a regulatory effect, exerted by one component onto a downstream component, a so-called 'causal interaction' takes place. Causal interactions constitute the building blocks in our understanding of larger regulatory networks in cells. These causal interactions and the biological processes they enable (e.g. gene regulation) need to be described with a careful appreciation of the underlying molecular reactions. A proper description of this information enables archiving, sharing and reuse by humans and for automated computational processing. Various representations of causal relationships between biological components are currently used in a variety of resources. RESULTS: Here, we propose a checklist that accommodates current representations, called the Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST). This checklist defines both the required core information, as well as a comprehensive set of other contextual details valuable to the end user and relevant for reusing and reproducing causal molecular interaction information. The MI2CAST checklist can be used as reporting guidelines when annotating and curating causal statements, while fostering uniformity and interoperability of the data across resources. AVAILABILITY AND IMPLEMENTATION: The checklist together with examples is accessible at https://github.com/MI2CAST/MI2CAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vasundra Touré, Steven Vercruysse, Marcio Luis Acencio, Ruth C. Lovering, Sandra E. Orchard, Glyn Bradley, Cristina Casals-Casas, Claudine Chaouiya, Noemi del-Toro, Åsmund Flobak, Pascale Gaudet, Henning Hermjakob, Charles Tapley Hoyt, Luana Licata, Astrid Lægreid, Chris Mungall, Anne Niknejad, Simona Panni, Livia Perfetto, Pablo Porras, Dexter Pratt, Julio Saez-Rodriguez, Denis Thieffry, Paul D. Thomas, Dénes Türei, Martin Kuiper |
Bioinform. | 22 |
| 2021 | A comprehensive database for integrated analysis of omics data in autoimmune diseasesabstractBACKGROUND: Autoimmune diseases are heterogeneous pathologies with difficult diagnosis and few therapeutic options. In the last decade, several omics studies have provided significant insights into the molecular mechanisms of these diseases. Nevertheless, data from different cohorts and pathologies are stored independently in public repositories and a unified resource is imperative to assist researchers in this field. RESULTS: Here, we present Autoimmune Diseases Explorer ( https://adex.genyo.es ), a database that integrates 82 curated transcriptomics and methylation studies covering 5609 samples for some of the most common autoimmune diseases. The database provides, in an easy-to-use environment, advanced data analysis and statistical methods for exploring omics datasets, including meta-analysis, differential expression or pathway analysis. CONCLUSIONS: This is the first omics database focused on autoimmune diseases. This resource incorporates homogeneously processed data to facilitate integrative analyses among studies. Jordi Martorell-Marugan, Raúl López-Domínguez, Adrián García-Moreno, Daniel Toro-Domínguez, Juan Antonio Villatoro-García, Guillermo Barturen, Adoración Martín-Gómez, Kevin Troulé, Gonzalo Gómez-López, Fátima Al-Shahrour, Víctor González-Rumayor, María Peña-Chilet, Joaquín Dopazo, Julio Saez-Rodriguez, Marta E. Alarcón-Riquelme, Pedro Carmona-Saez |
BMC Bioinform. | 14 |
| 2020 | Bringing data from curated pathway resources to Cytoscape with OmniPathabstractSUMMARY: Multiple databases provide valuable information about curated pathways and other resources that can be used to build and analyze networks. OmniPath combines 61 (and continuously growing) network resources into a comprehensive collection, with over 120 000 interactions. We present here the OmniPath App, a Cytoscape plugin to flexibly import data from OmniPath via a simple and intuitive interface. Thus, it makes possible to directly access the large body of high-quality knowledge provided by OmniPath within Cytoscape for inspection and further use with other tools. AVAILABILITY AND IMPLEMENTATION: The OmniPath App has been developed for Cytoscape 3 in the Java programing language. The latest source code and the plugin can be found at: https://github.com/saezlab/Omnipath_Cytoscape and http://apps.cytoscape.org/apps/omnipath, respectively. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Francesco Ceccarelli, Dénes Türei, Attila Gábor, Julio Saez-Rodriguez |
Bioinform. | 4 |
| 2020 | Converting networks to predictive logic models from perturbation signalling data with CellNOptabstractSUMMARY: The molecular changes induced by perturbations such as drugs and ligands are highly informative of the intracellular wiring. Our capacity to generate large datasets is increasing steadily. A useful way to extract mechanistic insight from the data is by integrating them with a prior knowledge network of signalling to obtain dynamic models. CellNOpt is a collection of Bioconductor R packages for building logic models from perturbation data and prior knowledge of signalling networks. We have recently developed new components and refined the existing ones to keep up with the computational demand of increasingly large datasets, including (i) an efficient integer linear programming, (ii) a probabilistic logic implementation for semi-quantitative datasets, (iii) the integration of a stochastic Boolean simulator, (iv) a tool to identify missing links, (v) systematic post-hoc analyses and (vi) an R-Shiny tool to run CellNOpt interactively. AVAILABILITY AND IMPLEMENTATION: R-package(s): https://github.com/saezlab/cellnopt. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Enio Gjerga, Panuwat Trairatphisan, Attila Gábor, Hermann Koch, Céline Chevalier, Franceco Ceccarelli, Aurélien Dugourd, Alexander Mitsos, Julio Saez-Rodriguez, Jonathan D. Wren |
Bioinform. | 9 |
| 2020 | BIAS: Transparent reporting of biomedical image analysis challengesabstractThe number of biomedical image analysis challenges organized per year is steadily increasing. These international competitions have the purpose of benchmarking algorithms on common data sets, typically to identify the best method for a given problem. Recent research, however, revealed that common practice related to challenge reporting does not allow for adequate interpretation and reproducibility of results. To address the discrepancy between the impact of challenges and the quality (control), the Biomedical Image Analysis ChallengeS (BIAS) initiative developed a set of recommendations for the reporting of challenges. The BIAS statement aims to improve the transparency of the reporting of a biomedical image analysis challenge regardless of field of application, image modality or task category assessed. This article describes how the BIAS statement was developed and presents a checklist which authors of biomedical image analysis challenges are encouraged to include in their submission when giving a paper on a challenge into review. The purpose of the checklist is to standardize and facilitate the review process and raise interpretability and reproducibility of challenge results by making relevant information explicit. Lena Maier-Hein, Annika Reinke, Michal Kozubek 0001, Anne L. Martel, Tal Arbel, Matthias Eisenmann, Allan Hanbury, Pierre Jannin, Henning Müller, Sinan Onogur, Julio Saez-Rodriguez, Bram van Ginneken, Annette Kopp-Schneider, Bennett A. Landman |
Medical Image Anal. | 11 |
| 2019 | Hybrid parallel multimethod hyperheuristic for mixed-integer dynamic optimization problems in computational systems biology
Patricia González, Pablo Argüeso-Alejandro, David R. Penas, Xoán C. Pardo, Julio Saez-Rodriguez, Julio R. Banga, Ramón Doallo |
J. Supercomput. | 5 |
| 2018 | GDSCTools for mining pharmacogenomic interactions in cancerabstractMotivation: Large pharmacogenomic screenings integrate heterogeneous cancer genomic datasets as well as anti-cancer drug responses on thousand human cancer cell lines. Mining this data to identify new therapies for cancer sub-populations would benefit from common data structures, modular computational biology tools and user-friendly interfaces. Results: We have developed GDSCTools: a software aimed at the identification of clinically relevant genomic markers of drug response. The Genomics of Drug Sensitivity in Cancer (GDSC) database (www.cancerRxgene.org) integrates heterogeneous cancer genomic datasets as well as anti-cancer drug responses on a thousand cancer cell lines. Including statistical tools (analysis of variance) and predictive methods (Elastic Net), as well as common data structures, GDSCTools allows users to reproduce published results from GDSC and to implement new analytical methods. In addition, non-GDSC data resources can also be analysed since drug responses and genomic features can be encoded as CSV files. Contact: [email protected] or saezrodriguez.rwth-aachen.de or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Thomas Cokelaer, Elisabeth Chen, Francesco Iorio, Michael P. Menden, Howard Lightfoot, Julio Saez-Rodriguez, Mathew Garnett |
Bioinform. | 6 |
| 2018 | A systematic atlas of chaperome deregulation topologies across the human cancer landscapeabstractProteome balance is safeguarded by the proteostasis network (PN), an intricately regulated network of conserved processes that evolved to maintain native function of the diverse ensemble of protein species, ensuring cellular and organismal health. Proteostasis imbalances and collapse are implicated in a spectrum of human diseases, from neurodegeneration to cancer. The characteristics of PN disease alterations however have not been assessed in a systematic way. Since the chaperome is among the central components of the PN, we focused on the chaperome in our study by utilizing a curated functional ontology of the human chaperome that we connect in a high-confidence physical protein-protein interaction network. Challenged by the lack of a systems-level understanding of proteostasis alterations in the heterogeneous spectrum of human cancers, we assessed gene expression across more than 10,000 patient biopsies covering 22 solid cancers. We derived a novel customized Meta-PCA dimension reduction approach yielding M-scores as quantitative indicators of disease expression changes to condense the complexity of cancer transcriptomics datasets into quantitative functional network topographies. We confirm upregulation of the HSP90 family and also highlight HSP60s, Prefoldins, HSP100s, ER- and mitochondria-specific chaperones as pan-cancer enriched. Our analysis also reveals a surprisingly consistent strong downregulation of small heat shock proteins (sHSPs) and we stratify two cancer groups based on the preferential upregulation of ATP-dependent chaperones. Strikingly, our analyses highlight similarities between stem cell and cancer proteostasis, and diametrically opposed chaperome deregulation between cancers and neurodegenerative diseases. We developed a web-based Proteostasis Profiler tool (Pro2) enabling intuitive analysis and visual exploration of proteostasis disease alterations using gene expression data. Our study showcases a comprehensive profiling of chaperome shifts in human cancers and sets the stage for a systematic global analysis of PN alterations across the human diseasome towards novel hypotheses for therapeutic network re-adjustment in proteostasis disorders. Ali Hadizadeh Esfahani, Angelina Sverchkova, Julio Saez-Rodriguez, Andreas Schuppert, Marc Brehme |
PLoS Comput. Biol. | 3 |
| 2018 | Computational discovery of dynamic cell line specific Boolean networks from multiplex time-course dataabstractProtein signaling networks are static views of dynamic processes where proteins go through many biochemical modifications such as ubiquitination and phosphorylation to propagate signals that regulate cells and can act as feed-back systems. Understanding the precise mechanisms underlying protein interactions can elucidate how signaling and cell cycle progression occur within cells in different diseases such as cancer. Large-scale protein signaling networks contain an important number of experimentally verified protein relations but lack the capability to predict the outcomes of the system, and therefore to be trained with respect to experimental measurements. Boolean Networks (BNs) are a simple yet powerful framework to study and model the dynamics of the protein signaling networks. While many BN approaches exist to model biological systems, they focus mainly on system properties, and few exist to integrate experimental data in them. In this work, we show an application of a method conceived to integrate time series phosphoproteomic data into protein signaling networks. We use a large-scale real case study from the HPN-DREAM Breast Cancer challenge. Our efficient and parameter-free method combines logic programming and model-checking to infer a family of BNs from multiple perturbation time series data of four breast cancer cell lines given a prior protein signaling network. Because each predicted BN family is cell line specific, our method highlights commonalities and discrepancies between the four cell lines. Our models have a Root Mean Square Error (RMSE) of 0.31 with respect to the testing data, while the best performant method of this HPN-DREAM challenge had a RMSE of 0.47. To further validate our results, BNs are compared with the canonical mTOR pathway showing a comparable AUROC score (0.77) to the top performing HPN-DREAM teams. In addition, our approach can also be used as a complementary method to identify erroneous experiments. These results prove our methodology as an efficient dynamic model discovery method in multiple perturbation time course experimental data of large-scale signaling networks. The software and data are publicly available at https://github.com/misbahch6/caspo-ts. Misbah Razzaq, Loïc Paulevé, Anne Siegel, Julio Saez-Rodriguez, Jérémie Bourdon, Carito Guziolowski |
PLoS Comput. Biol. | 4 |
| 2017 | Benchmarking substrate-based kinase activity inference using phosphoproteomic dataabstractMOTIVATION: Phosphoproteomic experiments are increasingly used to study the changes in signaling occurring across different conditions. It has been proposed that changes in phosphorylation of kinase target sites can be used to infer when a kinase activity is under regulation. However, these approaches have not yet been benchmarked due to a lack of appropriate benchmarking strategies. RESULTS: We used curated phosphoproteomic experiments and a gold standard dataset containing a total of 184 kinase-condition pairs where regulation is expected to occur to benchmark and compare different kinase activity inference strategies: Z-test, Kolmogorov Smirnov test, Wilcoxon rank sum test, gene set enrichment analysis (GSEA), and a multiple linear regression model. We also tested weighted variants of the Z-test and GSEA that include information on kinase sequence specificity as proxy for affinity. Finally, we tested how the number of known substrates and the type of evidence ( in vivo , in vitro or in silico ) supporting these influence the predictions. CONCLUSIONS: Most models performed well with the Z-test and the GSEA performing best as determined by the area under the ROC curve (Mean AUC = 0.722). Weighting kinase targets by the kinase target sequence preference improves the results marginally. However, the number of known substrates and the evidence supporting the interactions has a strong effect on the predictions. AVAILABILITY AND IMPLEMENTATION: The KSEA implementation is available in https://github.com/ evocellnet/ksea. Additional data is available in http://phosfate.com. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Claudia Hernandez-Armenta, David Ochoa, Emanuel J. V. Gonçalves, Julio Saez-Rodriguez, Pedro Beltrão |
Bioinform. | 4 |
| 2017 | caspo: a toolbox for automated reasoning on the response of logical signaling networks familiesabstractSummary: We introduce the caspo toolbox, a python package implementing a workflow for reasoning on logical networks families. Our software allows researchers to (i) a family of logical networks derived from a given topology and explaining the experimental response to various perturbations; (ii) all logical networks in a given family by their input-output behaviors; (iii) the response of the system to every possible perturbation based on the ensemble of predictions; (iv) new experimental perturbations to discriminate among a family of logical networks; and (v) a family of logical networks by finding all interventions strategies forcing a set of targets into a desired steady state. Availability and Implementation: caspo is open-source software distributed under the GPLv3 license. Source code is publicly hosted at http://github.com/bioasp/caspo . Contact: [email protected]. Santiago Videla, Julio Saez-Rodriguez, Carito Guziolowski, Anne Siegel |
Bioinform. | 2 |
| 2017 | Systematic Analysis of Transcriptional and Post-transcriptional Regulation of Metabolism in YeastabstractCells react to extracellular perturbations with complex and intertwined responses. Systematic identification of the regulatory mechanisms that control these responses is still a challenge and requires tailored analyses integrating different types of molecular data. Here we acquired time-resolved metabolomics measurements in yeast under salt and pheromone stimulation and developed a machine learning approach to explore regulatory associations between metabolism and signal transduction. Existing phosphoproteomics measurements under the same conditions and kinase-substrate regulatory interactions were used to in silico estimate the enzymatic activity of signalling kinases. Our approach identified informative associations between kinases and metabolic enzymes capable of predicting metabolic changes. We extended our analysis to two studies containing transcriptomics, phosphoproteomics and metabolomics measurements across a comprehensive panel of kinases/phosphatases knockouts and time-resolved perturbations to the nitrogen metabolism. Changes in activity of transcription factors, kinases and phosphatases were estimated in silico and these were capable of building predictive models to infer the metabolic adaptations of previously unseen conditions across different dynamic experiments. Time-resolved experiments were significantly more informative than genetic perturbations to infer metabolic adaptation. This difference may be due to the indirect nature of the associations and of general cellular states that can hinder the identification of causal relationships. This work provides a novel genome-scale integrative analysis to propose putative transcriptional and post-translational regulatory mechanisms of metabolic processes. Emanuel J. V. Gonçalves, Zrinka Raguz Nakic, Mattia Zampieri, Omar Wagih, David Ochoa, Uwe Sauer, Pedro Beltrão, Julio Saez-Rodriguez |
PLoS Comput. Biol. | 8 |
| 2017 | Data-driven reverse engineering of signaling pathways using ensembles of dynamic modelsabstractDespite significant efforts and remarkable progress, the inference of signaling networks from experimental data remains very challenging. The problem is particularly difficult when the objective is to obtain a dynamic model capable of predicting the effect of novel perturbations not considered during model training. The problem is ill-posed due to the nonlinear nature of these systems, the fact that only a fraction of the involved proteins and their post-translational modifications can be measured, and limitations on the technologies used for growing cells in vitro, perturbing them, and measuring their variations. As a consequence, there is a pervasive lack of identifiability. To overcome these issues, we present a methodology called SELDOM (enSEmbLe of Dynamic lOgic-based Models), which builds an ensemble of logic-based dynamic models, trains them to experimental data, and combines their individual simulations into an ensemble prediction. It also includes a model reduction step to prune spurious interactions and mitigate overfitting. SELDOM is a data-driven method, in the sense that it does not require any prior knowledge of the system: the interaction networks that act as scaffolds for the dynamic models are inferred from data using mutual information. We have tested SELDOM on a number of experimental and in silico signal transduction case-studies, including the recent HPN-DREAM breast cancer challenge. We found that its performance is highly competitive compared to state-of-the-art methods for the purpose of recovering network topology. More importantly, the utility of SELDOM goes beyond basic network inference (i.e. uncovering static interaction networks): it builds dynamic (based on ordinary differential equation) models, which can be used for mechanistic interpretations and reliable dynamic predictions in new experimental conditions (i.e. not used in the training). For this task, SELDOM's ensemble prediction is not only consistently better than predictions from individual models, but also often outperforms the state of the art represented by the methods used in the HPN-DREAM challenge. David Henriques, Alejandro Fernández Villaverde, Miguel Rocha 0001, Julio Saez-Rodriguez, Julio R. Banga |
PLoS Comput. Biol. | 4 |
| 2016 | Efficient randomization of biological networks while preserving functional characterization of individual nodesabstractBACKGROUND: Networks are popular and powerful tools to describe and model biological processes. Many computational methods have been developed to infer biological networks from literature, high-throughput experiments, and combinations of both. Additionally, a wide range of tools has been developed to map experimental data onto reference biological networks, in order to extract meaningful modules. Many of these methods assess results' significance against null distributions of randomized networks. However, these standard unconstrained randomizations do not preserve the functional characterization of the nodes in the reference networks (i.e. their degrees and connection signs), hence including potential biases in the assessment. RESULTS: Building on our previous work about rewiring bipartite networks, we propose a method for rewiring any type of unweighted networks. In particular we formally demonstrate that the problem of rewiring a signed and directed network preserving its functional connectivity (F-rewiring) reduces to the problem of rewiring two induced bipartite networks. Additionally, we reformulate the lower bound to the iterations' number of the switching-algorithm to make it suitable for the F-rewiring of networks of any size. Finally, we present BiRewire3, an open-source Bioconductor package enabling the F-rewiring of any type of unweighted network. We illustrate its application to a case study about the identification of modules from gene expression data mapped on protein interaction networks, and a second one focused on building logic models from more complex signed-directed reference signaling networks and phosphoproteomic data. CONCLUSIONS: BiRewire3 it is freely available at https://www.bioconductor.org/packages/BiRewire/ , and it should have a broad application as it allows an efficient and analytically derived statistical assessment of results from any network biology tool. Francesco Iorio, Marti Bernardo-Faura, Andrea Gobbi, Thomas Cokelaer, Giuseppe Jurman, Julio Saez-Rodriguez |
BMC Bioinform. | 6 |
| 2016 | A computational method for designing diverse linear epitopes including citrullinated peptides with desired binding affinities to intravenous immunoglobulinabstractBACKGROUND: Understanding the interactions between antibodies and the linear epitopes that they recognize is an important task in the study of immunological diseases. We present a novel computational method for the design of linear epitopes of specified binding affinity to Intravenous Immunoglobulin (IVIg). RESULTS: We show that the method, called Pythia-design can accurately design peptides with both high-binding affinity and low binding affinity to IVIg. To show this, we experimentally constructed and tested the computationally constructed designs. We further show experimentally that these designed peptides are more accurate that those produced by a recent method for the same task. Pythia-design is based on combining random walks with an ensemble of probabilistic support vector machines (SVM) classifiers, and we show that it produces a diverse set of designed peptides, an important property to develop robust sets of candidates for construction. We show that by combining Pythia-design and the method of (PloS ONE 6(8):23616, 2011), we are able to produce an even more accurate collection of designed peptides. Analysis of the experimental validation of Pythia-design peptides indicates that binding of IVIg is favored by epitopes that contain trypthophan and cysteine. CONCLUSIONS: Our method, Pythia-design, is able to generate a diverse set of binding and non-binding peptides, and its designs have been experimentally shown to be accurate. Rob Patro, Raquel Norel, Robert J. Prill, Julio Saez-Rodriguez, Peter Lorenz, Felix Steinbeck, Bjoern Ziems, Mitja Lustrek, Nicola Barbarini, Alessandra Tiengo, Riccardo Bellazzi, Hans-Jürgen Thiesen, Gustavo Stolovitzky, Carl Kingsford |
BMC Bioinform. | 4 |
| 2015 | Reverse engineering of logic-based differential equation models using a mixed-integer dynamic optimization approachabstractMOTIVATION: Systems biology models can be used to test new hypotheses formulated on the basis of previous knowledge or new experimental data, contradictory with a previously existing model. New hypotheses often come in the shape of a set of possible regulatory mechanisms. This search is usually not limited to finding a single regulation link, but rather a combination of links subject to great uncertainty or no information about the kinetic parameters. RESULTS: In this work, we combine a logic-based formalism, to describe all the possible regulatory structures for a given dynamic model of a pathway, with mixed-integer dynamic optimization (MIDO). This framework aims to simultaneously identify the regulatory structure (represented by binary parameters) and the real-valued parameters that are consistent with the available experimental data, resulting in a logic-based differential equation model. The alternative to this would be to perform real-valued parameter estimation for each possible model structure, which is not tractable for models of the size presented in this work. The performance of the method presented here is illustrated with several case studies: a synthetic pathway problem of signaling regulation, a two-component signal transduction pathway in bacterial homeostasis, and a signaling network in liver cancer cells. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected]. David Henriques, Miguel Rocha 0001, Julio Saez-Rodriguez, Julio R. Banga |
Bioinform. | 3 |
| 2015 | Cooperative development of logical modelling standards and tools with CoLoMoToabstractThe identification of large regulatory and signalling networks involved in the control of crucial cellular processes calls for proper modelling approaches. Indeed, models can help elucidate properties of these networks, understand their behaviour and provide (testable) predictions by performing in silico experiments. In this context, qualitative, logical frameworks have emerged as relevant approaches, as demonstrated by a growing number of published models, along with new methodologies and software tools. This productive activity now requires a concerted effort to ensure model reusability and interoperability between tools. Following an outline of the logical modelling framework, we present the most important achievements of the Consortium for Logical Models and Tools, along with future objectives. Our aim is to advertise this open community, which welcomes contributions from all researchers interested in logical modelling or in related mathematical and computational developments. Aurélien Naldi, Pedro T. Monteiro 0001, Christoph Müssel, Hans A. Kestler, Denis Thieffry, Ioannis Xenarios, Julio Saez-Rodriguez, Tomás Helikar, Claudine Chaouiya |
Bioinform. | 7 |
| 2015 | Extended notions of sign consistency to relate experimental data to signaling and regulatory network topologiesabstractBACKGROUND: A rapidly growing amount of knowledge about signaling and gene regulatory networks is available in databases such as KEGG, Reactome, or RegulonDB. There is an increasing need to relate this knowledge to high-throughput data in order to (in)validate network topologies or to decide which interactions are present or inactive in a given cell type under a particular environmental condition. Interaction graphs provide a suitable representation of cellular networks with information flows and methods based on sign consistency approaches have been shown to be valuable tools to (i) predict qualitative responses, (ii) to test the consistency of network topologies and experimental data, and (iii) to apply repair operations to the network model suggesting missing or wrong interactions. RESULTS: We present a framework to unify different notions of sign consistency and propose a refined method for data discretization that considers uncertainties in experimental profiles. We furthermore introduce a new constraint to filter undesired model behaviors induced by positive feedback loops. Finally, we generalize the way predictions can be made by the sign consistency approach. In particular, we distinguish strong predictions (e.g. increase of a node level) and weak predictions (e.g., node level increases or remains unchanged) enlarging the overall predictive power of the approach. We then demonstrate the applicability of our framework by confronting a large-scale gene regulatory network model of Escherichia coli with high-throughput transcriptomic measurements. CONCLUSION: Overall, our work enhances the flexibility and power of the sign consistency approach for the prediction of the behavior of signaling and gene regulatory networks and, more generally, for the validation and inference of these networks. Sven Thiele, Luca Cerone, Julio Saez-Rodriguez, Anne Siegel, Carito Guziolowski, Steffen Klamt |
BMC Bioinform. | 3 |
| 2015 | Learning Boolean logic models of signaling networks with ASP
Santiago Videla, Carito Guziolowski, Federica Eduati, Sven Thiele, Martin Gebser, Jacques Nicolas, Julio Saez-Rodriguez, Torsten Schaub, Anne Siegel |
Theor. Comput. Sci. | 7 |
| 2014 | Fast randomization of large genomic datasets while preserving alteration countsabstractMOTIVATION: Studying combinatorial patterns in cancer genomic datasets has recently emerged as a tool for identifying novel cancer driver networks. Approaches have been devised to quantify, for example, the tendency of a set of genes to be mutated in a 'mutually exclusive' manner. The significance of the proposed metrics is usually evaluated by computing P-values under appropriate null models. To this end, a Monte Carlo method (the switching-algorithm) is used to sample simulated datasets under a null model that preserves patient- and gene-wise mutation rates. In this method, a genomic dataset is represented as a bipartite network, to which Markov chain updates (switching-steps) are applied. These steps modify the network topology, and a minimal number of them must be executed to draw simulated datasets independently under the null model. This number has previously been deducted empirically to be a linear function of the total number of variants, making this process computationally expensive. RESULTS: We present a novel approximate lower bound for the number of switching-steps, derived analytically. Additionally, we have developed the R package BiRewire, including new efficient implementations of the switching-algorithm. We illustrate the performances of BiRewire by applying it to large real cancer genomics datasets. We report vast reductions in time requirement, with respect to existing implementations/bounds and equivalent P-value computations. Thus, we propose BiRewire to study statistical properties in genomic datasets, and other data that can be modeled as bipartite networks. AVAILABILITY AND IMPLEMENTATION: BiRewire is available on BioConductor at http://www.bioconductor.org/packages/2.13/bioc/html/BiRewire.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andrea Gobbi, Francesco Iorio, Kevin J. Dawson, David C. Wedge, David Tamborero, Ludmil B. Alexandrov, Núria López-Bigas, Mathew Garnett, Giuseppe Jurman, Julio Saez-Rodriguez |
Bioinform. | 10 |
| 2014 | Exhaustively characterizing feasible logic models of a signaling network using Answer Set Programming
Carito Guziolowski, Santiago Videla, Federica Eduati, Sven Thiele, Thomas Cokelaer, Anne Siegel, Julio Saez-Rodriguez |
Bioinform. | 7 |
| 2014 | MEIGO: an open-source software suite based on metaheuristics for global optimization in systems biology and bioinformaticsabstractBACKGROUND: Optimization is the key to solving many problems in computational biology. Global optimization methods, which provide a robust methodology, and metaheuristics in particular have proven to be the most efficient methods for many applications. Despite their utility, there is a limited availability of metaheuristic tools. RESULTS: We present MEIGO, an R and Matlab optimization toolbox (also available in Python via a wrapper of the R version), that implements metaheuristics capable of solving diverse problems arising in systems biology and bioinformatics. The toolbox includes the enhanced scatter search method (eSS) for continuous nonlinear programming (cNLP) and mixed-integer programming (MINLP) problems, and variable neighborhood search (VNS) for Integer Programming (IP) problems. Additionally, the R version includes BayesFit for parameter estimation by Bayesian inference. The eSS and VNS methods can be run on a single-thread or in parallel using a cooperative strategy. The code is supplied under GPLv3 and is available at http://www.iim.csic.es/~gingproc/meigo.html. Documentation and examples are included. The R package has been submitted to BioConductor. We evaluate MEIGO against optimization benchmarks, and illustrate its applicability to a series of case studies in bioinformatics and systems biology where it outperforms other state-of-the-art methods. CONCLUSIONS: MEIGO provides a free, open-source platform for optimization that can be applied to multiple domains of systems biology and bioinformatics. It includes efficient state of the art metaheuristics, and its open and modular structure allows the addition of further methods. Jose A. Egea, David Henriques, Thomas Cokelaer, Alejandro Fernández Villaverde, Aidan MacNamara, Diana-Patricia Danciu, Julio R. Banga, Julio Saez-Rodriguez |
BMC Bioinform. | 8 |
| 2013 | BioServices: a common Python package to access biological Web Services programmaticallyabstractAbstract Motivation: Web interfaces provide access to numerous biological databases. Many can be accessed to in a programmatic way thanks to Web Services. Building applications that combine several of them would benefit from a single framework. Results: BioServices is a comprehensive Python framework that provides programmatic access to major bioinformatics Web Services (e.g. KEGG, UniProt, BioModels, ChEMBLdb). Wrapping additional Web Services based either on Representational State Transfer or Simple Object Access Protocol/Web Services Description Language technologies is eased by the usage of object-oriented programming. Availability and implementation: BioServices releases and documentation are available at http://pypi.python.org/pypi/bioservices under a GPL-v3 license. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Thomas Cokelaer, Dennis Pultz, Lea M. Harder, Jordi Serra-Musach, Julio Saez-Rodriguez |
Bioinform. | 5 |
| 2013 | Exhaustively characterizing feasible logic models of a signaling network using Answer Set ProgrammingabstractMOTIVATION: Logic modeling is a useful tool to study signal transduction across multiple pathways. Logic models can be generated by training a network containing the prior knowledge to phospho-proteomics data. The training can be performed using stochastic optimization procedures, but these are unable to guarantee a global optima or to report the complete family of feasible models. This, however, is essential to provide precise insight in the mechanisms underlaying signal transduction and generate reliable predictions. RESULTS: We propose the use of Answer Set Programming to explore exhaustively the space of feasible logic models. Toward this end, we have developed caspo, an open-source Python package that provides a powerful platform to learn and characterize logic models by leveraging the rich modeling language and solving technologies of Answer Set Programming. We illustrate the usefulness of caspo by revisiting a model of pro-growth and inflammatory pathways in liver cells. We show that, if experimental error is taken into account, there are thousands (11 700) of models compatible with the data. Despite the large number, we can extract structural features from the models, such as links that are always (or never) present or modules that appear in a mutual exclusive fashion. To further characterize this family of models, we investigate the input-output behavior of the models. We find 91 behaviors across the 11 700 models and we suggest new experiments to discriminate among them. Our results underscore the importance of characterizing in a global and exhaustive manner the family of feasible models, with important implications for experimental design. AVAILABILITY: caspo is freely available for download (license GPLv3) and as a web service at http://caspo.genouest.org/. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. CONTACT: [email protected]. Carito Guziolowski, Santiago Videla, Federica Eduati, Sven Thiele, Thomas Cokelaer, Anne Siegel, Julio Saez-Rodriguez |
Bioinform. | 7 |
| 2013 | DvD: An R/Cytoscape pipeline for drug repurposing using public repositories of gene expression dataabstractSUMMARY: Drug versus Disease (DvD) provides a pipeline, available through R or Cytoscape, for the comparison of drug and disease gene expression profiles from public microarray repositories. Negatively correlated profiles can be used to generate hypotheses of drug-repurposing, whereas positively correlated profiles may be used to infer side effects of drugs. DvD allows users to compare drug and disease signatures with dynamic access to databases Array Express, Gene Expression Omnibus and data from the Connectivity Map. AVAILABILITY AND IMPLEMENTATION: R package (submitted to Bioconductor) under GPL 3 and Cytoscape plug-in freely available for download at www.ebi.ac.uk/saezrodriguez/DVD/. Clare Pacini, Francesco Iorio, Emanuel J. V. Gonçalves, Murat Iskar, Thomas Klabunde, Peer Bork, Julio Saez-Rodriguez |
Bioinform. | 7 |
| 2013 | CySBGN: A Cytoscape plug-in to integrate SBGN mapsabstractBACKGROUND: A standard graphical notation is essential to facilitate exchange of network representations of biological processes. Towards this end, the Systems Biology Graphical Notation (SBGN) has been proposed, and it is already supported by a number of tools. However, support for SBGN in Cytoscape, one of the most widely used platforms in biology to visualise and analyse networks, is limited, and in particular it is not possible to import SBGN diagrams. RESULTS: We have developed CySBGN, a Cytoscape plug-in that extends the use of Cytoscape visualisation and analysis features to SBGN maps. CySBGN adds support for Cytoscape users to visualize any of the three complementary SBGN languages: Process Description, Entity Relationship, and Activity Flow. The interoperability with other tools (CySBML plug-in and Systems Biology Format Converter) was also established allowing an automated generation of SBGN diagrams based on previously imported SBML models. The plug-in was tested using a suite of 53 different test cases that covers almost all possible entities, shapes, and connections. A rendering comparison with other tools that support SBGN was performed. To illustrate the interoperability with other Cytoscape functionalities, we present two analysis examples, shortest path calculation, and motif identification in a metabolic network. CONCLUSIONS: CySBGN imports, modifies and analyzes SBGN diagrams in Cytoscape, and thus allows the application of the large palette of tools and plug-ins in this platform to networks and pathways in SBGN format. Emanuel J. V. Gonçalves, Martijn P. van Iersel, Julio Saez-Rodriguez |
BMC Bioinform. | 3 |
| 2012 | Integrating literature-constrained and data-driven inference of signalling networksabstractMOTIVATION: Recent developments in experimental methods facilitate increasingly larger signal transduction datasets. Two main approaches can be taken to derive a mathematical model from these data: training a network (obtained, e.g., from literature) to the data, or inferring the network from the data alone. Purely data-driven methods scale up poorly and have limited interpretability, whereas literature-constrained methods cannot deal with incomplete networks. RESULTS: We present an efficient approach, implemented in the R package CNORfeeder, to integrate literature-constrained and data-driven methods to infer signalling networks from perturbation experiments. Our method extends a given network with links derived from the data via various inference methods, and uses information on physical interactions of proteins to guide and validate the integration of links. We apply CNORfeeder to a network of growth and inflammatory signalling. We obtain a model with superior data fit in the human liver cancer HepG2 and propose potential missing pathways. AVAILABILITY: CNORfeeder is in the process of being submitted to Bioconductor and in the meantime available at www.cellnopt.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Federica Eduati, Javier De Las Rivas, Barbara Di Camillo, Gianna Toffolo, Julio Saez-Rodriguez |
Bioinform. | 5 |
| 2011 | Training Signaling Pathway Maps to Biochemical Data with Constrained Fuzzy Logic: Quantitative Analysis of Liver Cell Responses to Inflammatory StimuliabstractPredictive understanding of cell signaling network operation based on general prior knowledge but consistent with empirical data in a specific environmental context is a current challenge in computational biology. Recent work has demonstrated that Boolean logic can be used to create context-specific network models by training proteomic pathway maps to dedicated biochemical data; however, the Boolean formalism is restricted to characterizing protein species as either fully active or inactive. To advance beyond this limitation, we propose a novel form of fuzzy logic sufficiently flexible to model quantitative data but also sufficiently simple to efficiently construct models by training pathway maps on dedicated experimental measurements. Our new approach, termed constrained fuzzy logic (cFL), converts a prior knowledge network (obtained from literature or interactome databases) into a computable model that describes graded values of protein activation across multiple pathways. We train a cFL-converted network to experimental data describing hepatocytic protein activation by inflammatory cytokines and demonstrate the application of the resultant trained models for three important purposes: (a) generating experimentally testable biological hypotheses concerning pathway crosstalk, (b) establishing capability for quantitative prediction of protein activity, and (c) prediction and understanding of the cytokine release phenotypic response. Our methodology systematically and quantitatively trains a protein pathway map summarizing curated literature to context-specific biochemical data. This process generates a computable model yielding successful prediction of new test data and offering biological insight into complex datasets that are difficult to fully analyze by intuition alone. Melody K. Morris, Julio Saez-Rodriguez, David C. Clarke, Peter K. Sorger, Douglas A. Lauffenburger |
PLoS Comput. Biol. | 2 |
| 2009 | Fuzzy Logic Analysis of Kinase Pathway Crosstalk in TNF/EGF/Insulin-Induced SignalingabstractWhen modeling cell signaling networks, a balance must be struck between mechanistic detail and ease of interpretation. In this paper we apply a fuzzy logic framework to the analysis of a large, systematic dataset describing the dynamics of cell signaling downstream of TNF, EGF, and insulin receptors in human colon carcinoma cells. Simulations based on fuzzy logic recapitulate most features of the data and generate several predictions involving pathway crosstalk and regulation. We uncover a relationship between MK2 and ERK pathways that might account for the previously identified pro-survival influence of MK2. We also find unexpected inhibition of IKK following EGF treatment, possibly due to down-regulation of autocrine signaling. More generally, fuzzy logic models are flexible, able to incorporate qualitative and noisy data, and powerful enough to produce quantitative predictions and new biological insights about the operation of signaling networks. Bree B. Aldridge, Julio Saez-Rodriguez, Jeremy Muhlich, Peter K. Sorger, Douglas A. Lauffenburger |
PLoS Comput. Biol. | 2 |
| 2009 | Identifying Drug Effects via Pathway Alterations using an Integer Linear Programming Optimization Formulation on Phosphoproteomic DataabstractUnderstanding the mechanisms of cell function and drug action is a major endeavor in the pharmaceutical industry. Drug effects are governed by the intrinsic properties of the drug (i.e., selectivity and potency) and the specific signaling transduction network of the host (i.e., normal vs. diseased cells). Here, we describe an unbiased, phosphoproteomic-based approach to identify drug effects by monitoring drug-induced topology alterations. With our proposed method, drug effects are investigated under diverse stimulations of the signaling network. Starting with a generic pathway made of logical gates, we build a cell-type specific map by constraining it to fit 13 key phopshoprotein signals under 55 experimental conditions. Fitting is performed via an Integer Linear Program (ILP) formulation and solution by standard ILP solvers; a procedure that drastically outperforms previous fitting schemes. Then, knowing the cell's topology, we monitor the same key phosphoprotein signals under the presence of drug and we re-optimize the specific map to reveal drug-induced topology alterations. To prove our case, we make a topology for the hepatocytic cell-line HepG2 and we evaluate the effects of 4 drugs: 3 selective inhibitors for the Epidermal Growth Factor Receptor (EGFR) and a non-selective drug. We confirm effects easily predictable from the drugs' main target (i.e., EGFR inhibitors blocks the EGFR pathway) but we also uncover unanticipated effects due to either drug promiscuity or the cell's specific topology. An interesting finding is that the selective EGFR inhibitor Gefitinib inhibits signaling downstream the Interleukin-1alpha (IL1alpha) pathway; an effect that cannot be extracted from binding affinity-based approaches. Our method represents an unbiased approach to identify drug effects on small to medium size pathways which is scalable to larger topologies with any type of signaling interventions (small molecules, RNAi, etc). The method can reveal drug effects on pathways, the cornerstone for identifying mechanisms of drug's efficacy. Alexander Mitsos, Ioannis N. Melas, Paraskeuas Siminelakis, Aikaterini D. Chairakaki, Julio Saez-Rodriguez, Leonidas G. Alexopoulos |
PLoS Comput. Biol. | 5 |
| 2009 | The Logic of EGFR/ErbB Signaling: Theoretical Properties and Analysis of High-Throughput DataabstractThe epidermal growth factor receptor (EGFR) signaling pathway is probably the best-studied receptor system in mammalian cells, and it also has become a popular example for employing mathematical modeling to cellular signaling networks. Dynamic models have the highest explanatory and predictive potential; however, the lack of kinetic information restricts current models of EGFR signaling to smaller sub-networks. This work aims to provide a large-scale qualitative model that comprises the main and also the side routes of EGFR/ErbB signaling and that still enables one to derive important functional properties and predictions. Using a recently introduced logical modeling framework, we first examined general topological properties and the qualitative stimulus-response behavior of the network. With species equivalence classes, we introduce a new technique for logical networks that reveals sets of nodes strongly coupled in their behavior. We also analyzed a model variant which explicitly accounts for uncertainties regarding the logical combination of signals in the model. The predictive power of this model is still high, indicating highly redundant sub-structures in the network. Finally, one key advance of this work is the introduction of new techniques for assessing high-throughput data with logical models (and their underlying interaction graph). By employing these techniques for phospho-proteomic data from primary hepatocytes and the HepG2 cell line, we demonstrate that our approach enables one to uncover inconsistencies between experimental results and our current qualitative knowledge and to generate new hypotheses and conclusions. Our results strongly suggest that the Rac/Cdc42 induced p38 and JNK cascades are independent of PI3K in both primary hepatocytes and HepG2. Furthermore, we detected that the activation of JNK in response to neuregulin follows a PI3K-dependent signaling pathway. Regina Samaga, Julio Saez-Rodriguez, Leonidas G. Alexopoulos, Peter K. Sorger, Steffen Klamt |
PLoS Comput. Biol. | 2 |
| 2008 | Flexible informatics for linking experimental data to mathematical models via DataRailabstractMOTIVATION: Linking experimental data to mathematical models in biology is impeded by the lack of suitable software to manage and transform data. Model calibration would be facilitated and models would increase in value were it possible to preserve links to training data along with a record of all normalization, scaling, and fusion routines used to assemble the training data from primary results. RESULTS: We describe the implementation of DataRail, an open source MATLAB-based toolbox that stores experimental data in flexible multi-dimensional arrays, transforms arrays so as to maximize information content, and then constructs models using internal or external tools. Data integrity is maintained via a containment hierarchy for arrays, imposition of a metadata standard based on a newly proposed MIDAS format, assignment of semantically typed universal identifiers, and implementation of a procedure for storing the history of all transformations with the array. We illustrate the utility of DataRail by processing a newly collected set of approximately 22 000 measurements of protein activities obtained from cytokine-stimulated primary and transformed human liver cells. AVAILABILITY: DataRail is distributed under the GNU General Public License and available at http://code.google.com/p/sbpipeline/ Julio Saez-Rodriguez, Arthur Goldsipe, Jeremy Muhlich, Leonidas G. Alexopoulos, Bjorn Millard, Douglas A. Lauffenburger, Peter K. Sorger |
Bioinform. | 1 |
| 2007 | A Logical Model Provides Insights into T Cell Receptor SignalingabstractCellular decisions are determined by complex molecular interaction networks. Large-scale signaling networks are currently being reconstructed, but the kinetic parameters and quantitative data that would allow for dynamic modeling are still scarce. Therefore, computational studies based upon the structure of these networks are of great interest. Here, a methodology relying on a logical formalism is applied to the functional analysis of the complex signaling network governing the activation of T cells via the T cell receptor, the CD4/CD8 co-receptors, and the accessory signaling receptor CD28. Our large-scale Boolean model, which comprises 94 nodes and 123 interactions and is based upon well-established qualitative knowledge from primary T cells, reveals important structural features (e.g., feedback loops and network-wide dependencies) and recapitulates the global behavior of this network for an array of published data on T cell activation in wild-type and knock-out conditions. More importantly, the model predicted unexpected signaling events after antibody-mediated perturbation of CD28 and after genetic knockout of the kinase Fyn that were subsequently experimentally validated. Finally, we show that the logical model reveals key elements and potential failure modes in network functioning and provides candidates for missing links. In summary, our large-scale logical model for T cell activation proved to be a promising in silico tool, and it inspires immunologists to ask new questions. We think that it holds valuable potential in foreseeing the effects of drugs and network modifications. Julio Saez-Rodriguez, Luca Simeoni, Jonathan A. Lindquist, Rebecca Hemenway, Ursula Bommhardt, Boerge Arndt, Utz-Uwe Haus, Robert Weismantel, Ernst Dieter Gilles, Steffen Klamt, Burkhart Schraven |
PLoS Comput. Biol. | 1 |
| 2006 | A domain-oriented approach to the reduction of combinatorial complexity in signal transduction networksabstractBACKGROUND: Receptors and scaffold proteins possess a number of distinct domains and bind multiple partners. A common problem in modeling signaling systems arises from a combinatorial explosion of different states generated by feasible molecular species. The number of possible species grows exponentially with the number of different docking sites and can easily reach several millions. Models accounting for this combinatorial variety become impractical for many applications. RESULTS: Our results show that under realistic assumptions on domain interactions, the dynamics of signaling pathways can be exactly described by reduced, hierarchically structured models. The method presented here provides a rigorous way to model a large class of signaling networks using macro-states (macroscopic quantities such as the levels of occupancy of the binding domains) instead of micro-states (concentrations of individual species). The method is described using generic multidomain proteins and is applied to the molecule LAT. CONCLUSION: The presented method is a systematic and powerful tool to derive reduced model structures describing the dynamics of multiprotein complex formation accurately. Holger Conzelmann, Julio Saez-Rodriguez, Thomas Sauter, Boris N. Kholodenko, Ernst Dieter Gilles |
BMC Bioinform. | 2 |
| 2006 | A methodology for the structural and functional analysis of signaling and regulatory networksabstractBACKGROUND: Structural analysis of cellular interaction networks contributes to a deeper understanding of network-wide interdependencies, causal relationships, and basic functional capabilities. While the structural analysis of metabolic networks is a well-established field, similar methodologies have been scarcely developed and applied to signaling and regulatory networks. RESULTS: We propose formalisms and methods, relying on adapted and partially newly introduced approaches, which facilitate a structural analysis of signaling and regulatory networks with focus on functional aspects. We use two different formalisms to represent and analyze interaction networks: interaction graphs and (logical) interaction hypergraphs. We show that, in interaction graphs, the determination of feedback cycles and of all the signaling paths between any pair of species is equivalent to the computation of elementary modes known from metabolic networks. Knowledge on the set of signaling paths and feedback loops facilitates the computation of intervention strategies and the classification of compounds into activators, inhibitors, ambivalent factors, and non-affecting factors with respect to a certain species. In some cases, qualitative effects induced by perturbations can be unambiguously predicted from the network scheme. Interaction graphs however, are not able to capture AND relationships which do frequently occur in interaction networks. The consequent logical concatenation of all the arcs pointing into a species leads to Boolean networks. For a Boolean representation of cellular interaction networks we propose a formalism based on logical (or signed) interaction hypergraphs, which facilitates in particular a logical steady state analysis (LSSA). LSSA enables studies on the logical processing of signals and the identification of optimal intervention points (targets) in cellular networks. LSSA also reveals network regions whose parametrization and initial states are crucial for the dynamic behavior. We have implemented these methods in our software tool CellNetAnalyzer (successor of FluxAnalyzer) and illustrate their applicability using a logical model of T-Cell receptor signaling providing non-intuitive results regarding feedback loops, essential elements, and (logical) signal processing upon different stimuli. CONCLUSION: The methods and formalisms we propose herein are another step towards the comprehensive functional analysis of cellular interaction networks. Their potential, shown on a realistic T-cell signaling model, makes them a promising tool. Steffen Klamt, Julio Saez-Rodriguez, Jonathan A. Lindquist, Luca Simeoni, Ernst Dieter Gilles |
BMC Bioinform. | 2 |
| 2006 | Visual setup of logical models of signaling and regulatory networks with ProMoTabstractBACKGROUND: The analysis of biochemical networks using a logical (Boolean) description is an important approach in Systems Biology. Recently, new methods have been proposed to analyze large signaling and regulatory networks using this formalism. Even though there is a large number of tools to set up models describing biological networks using a biochemical (kinetic) formalism, however, they do not support logical models. RESULTS: Herein we present a flexible framework for setting up large logical models in a visual manner with the software tool ProMoT. An easily extendible library, ProMoT's inherent modularity and object-oriented concept as well as adaptive visualization techniques provide a versatile environment. Both the graphical and the textual description of the logical model can be exported to different formats. CONCLUSION: New features of ProMoT facilitate an efficient set-up of large Boolean models of biochemical interaction networks. The modeling environment is flexible; it can easily be adapted to specific requirements, and new extensions can be introduced. ProMoT is freely available from http://www.mpi-magdeburg.mpg.de/projects/promot/. Julio Saez-Rodriguez, Sebastian Mirschel, Rebecca Hemenway, Steffen Klamt, Ernst Dieter Gilles, Martin Ginkel |
BMC Bioinform. | 1 |