EDBT 2026 Demo / reviewers in the wild / expert
Anne Siegel
dblp:20/3925
· DBLP profile ↗
32ranked-venue papers
0as first author
9since 2021 · last 2024
0000-0001-6542-1568ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 8 since 2021Theory of computation · 7Artificial intelligence and machine learning · 5 · 1 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CEGAR-Based Approach for Solving Combinatorial Optimization Modulo Quantified Linear Arithmetics ProblemsabstractBioinformatics has always been a prolific domain for generating complex satisfiability and optimization problems. For instance, the synthesis of multi-scale models of biological networks has recently been associated with the resolution of optimization problems mixing Boolean logic and universally quantified linear constraints (OPT+qLP), which can be benchmarked on real-world models. In this paper, we introduce a Counter-Example-Guided Abstraction Refinement (CEGAR) to solve such problems efficiently. Our CEGAR exploits monotone properties inherent to linear optimization in order to generalize counter-examples of Boolean relaxations. We implemented our approach by extending Answer Set Programming (ASP) solver Clingo with a quantified linear constraints propagator. Our prototype enables exploiting independence of sub-formulas to further exploit the generalization of counter-examples. We evaluate the impact of refinement and partitioning on two sets of OPT+qLP problems inspired by system biology. Additionally, we conducted a comparison with the state-of-the-art ASP solver Clingo[lpx] that handles non-quantified linear constraints, showing the advantage of our CEGAR approach for solving large problems. Kerian Thuillier, Anne Siegel, Loïc Paulevé |
AAAI | 2 |
| 2024 | A rule-based multiscale model of hepatic stellate cell plasticity: Critical role of the inactivation loop in fibrosis progressionabstractHepatic stellate cells (HSC) are the source of extracellular matrix (ECM) whose overproduction leads to fibrosis, a condition that impairs liver functions in chronic liver diseases. Understanding the dynamics of HSCs will provide insights needed to develop new therapeutic approaches. Few models of hepatic fibrosis have been proposed, and none of them include the heterogeneity of HSC phenotypes recently highlighted by single-cell RNA sequencing analyses. Here, we developed rule-based models to study HSC dynamics during fibrosis progression and reversion. We used the Kappa graph rewriting language, for which we used tokens and counters to overcome temporal explosion. HSCs are modeled as agents that present seven physiological cellular states and that interact with (TGFβ1) molecules which regulate HSC activation and the secretion of type I collagen, the main component of the ECM. Simulation studies revealed the critical role of the HSC inactivation process during fibrosis progression and reversion. While inactivation allows elimination of activated HSCs during reversion steps, reactivation loops of inactivated HSCs (iHSCs) are required to sustain fibrosis. Furthermore, we demonstrated the model's sensitivity to (TGFβ1) parameters, suggesting its adaptability to a variety of pathophysiological conditions for which levels of (TGFβ1) production associated with the inflammatory response differ. Using new experimental data from a mouse model of CCl4-induced liver fibrosis, we validated the predicted ECM dynamics. Our model also predicts the accumulation of iHSCs during chronic liver disease. By analyzing RNA sequencing data from patients with non-alcoholic steatohepatitis (NASH) associated with liver fibrosis, we confirmed this accumulation, identifying iHSCs as novel markers of fibrosis progression. Overall, our study provides the first model of HSC dynamics in chronic liver disease that can be used to explore the regulatory role of iHSCs in liver homeostasis. Moreover, our model can also be generalized to fibroblasts during repair and fibrosis in other tissues. Matthieu Bouguéon, Vincent Legagneux, Octave Hazard, Jérémy Bomo, Anne Siegel, Jérôme Feret, Nathalie Théret |
PLoS Comput. Biol. | 5 |
| 2024 | Regulus infers signed regulatory relations from few samples' information using discretization and likelihood constraintsabstractMOTIVATION: Transcriptional regulation is performed by transcription factors (TF) binding to DNA in context-dependent regulatory regions and determines the activation or inhibition of gene expression. Current methods of transcriptional regulatory circuits inference, based on one or all of TF, regions and genes activity measurements require a large number of samples for ranking the candidate TF-gene regulation relations and rarely predict whether they are activations or inhibitions. We hypothesize that transcriptional regulatory circuits can be inferred from fewer samples by (1) fully integrating information on TF binding, gene expression and regulatory regions accessibility, (2) reducing data complexity and (3) using biology-based likelihood constraints to determine the global consistency between a candidate TF-gene relation and patterns of genes expressions and region activations, as well as qualify regulations as activations or inhibitions. RESULTS: We introduce Regulus, a method which computes TF-gene relations from gene expressions, regulatory region activities and TF binding sites data, together with the genomic locations of all entities. After aggregating gene expressions and region activities into patterns, data are integrated into a RDF (Resource Description Framework) endpoint. A dedicated SPARQL (SPARQL Protocol and RDF Query Language) query retrieves all potential relations between expressed TF and genes involving active regulatory regions. These TF-region-gene relations are then filtered using biological likelihood constraints allowing to qualify them as activation or inhibition. Regulus provides signed relations consistent with public databases and, when applied to biological data, identifies both known and potential new regulators. Regulus is devoted to context-specific transcriptional circuits inference in human settings where samples are scarce and cell populations are closely related, using discretization into patterns and likelihood reasoning to decipher the most robust regulatory relations. Marine Louarn, Guillaume Collet, Ève Barré, Thierry Fest, Olivier Dameron, Anne Siegel, Fabrice Chatonnet |
PLoS Comput. Biol. | 6 |
| 2024 | SPARTA: Interpretable functional classification of microbiomes and detection of hidden cumulative effectsabstractThe composition of the gut microbiota is a known factor in various diseases and has proven to be a strong basis for automatic classification of disease state. A need for a better understanding of microbiota data on the functional scale has since been voiced, as it would enhance these approaches' biological interpretability. In this paper, we have developed a computational pipeline for integrating the functional annotation of the gut microbiota into an automatic classification process and facilitating downstream interpretation of its results. The process takes as input taxonomic composition data, which can be built from 16S or whole genome sequencing, and links each component to its functional annotations through interrogation of the UniProt database. A functional profile of the gut microbiota is built from this basis. Both profiles, microbial and functional, are used to train Random Forest classifiers to discern unhealthy from control samples. SPARTA ensures full reproducibility and exploration of inherent variability by extending state-of-the-art methods in three dimensions: increased number of trained random forests, selection of important variables with an iterative process, repetition of full selection process from different seeds. This process shows that the translation of the microbiota into functional profiles gives non-significantly different performances when compared to microbial profiles on 5 of 6 datasets. This approach's main contribution however stems from its interpretability rather than its performance: through repetition, it also outputs a robust subset of discriminant variables. These selections were shown to be more consistent than those obtained by a state-of-the-art method, and their contents were validated through a manual bibliographic research. The interconnections between selected taxa and functional annotations were also analyzed and revealed that important annotations emerge from the cumulated influence of non-selected taxa. Baptiste Ruiz, Arnaud Belcour, Samuel Blanquart, Sylvie Buffet-Bataillon, Isabelle Le Huërou-Luron, Anne Siegel, Yann Le Cunff |
PLoS Comput. Biol. | 6 |
| 2022 | Addressing barriers in comprehensiveness, accessibility, reusability, interoperability and reproducibility of computational models in systems biologyabstractComputational models are often employed in systems biology to study the dynamic behaviours of complex systems. With the rise in the number of computational models, finding ways to improve the reusability of these models and their ability to reproduce virtual experiments becomes critical. Correct and effective model annotation in community-supported and standardised formats is necessary for this improvement. Here, we present recent efforts toward a common framework for annotated, accessible, reproducible and interoperable computational models in biology, and discuss key challenges of the field. Anna Niarakis, Dagmar Waltemath, James A. Glazier, Falk Schreiber, Sarah M. Keating, David P. Nickerson, Claudine Chaouiya, Anne Siegel, Vincent Noel, Henning Hermjakob, Tomás Helikar, Sylvain Soliman, Laurence Calzone |
Briefings Bioinform. | 8 |
| 2022 | MERRIN: MEtabolic regulation rule INference from time series dataabstractMOTIVATION: Many techniques have been developed to infer Boolean regulations from a prior knowledge network (PKN) and experimental data. Existing methods are able to reverse-engineer Boolean regulations for transcriptional and signaling networks, but they fail to infer regulations that control metabolic networks. RESULTS: We present a novel approach to infer Boolean rules for metabolic regulation from time-series data and a PKN. Our method is based on a combination of answer set programming and linear programming. By solving both combinatorial and linear arithmetic constraints, we generate candidate Boolean regulations that can reproduce the given data when coupled to the metabolic network. We evaluate our approach on a core regulated metabolic network and show how the quality of the predictions depends on the available kinetic, fluxomics or transcriptomics time-series data. AVAILABILITY AND IMPLEMENTATION: Software available at https://github.com/bioasp/merrin. SUPPLEMENTARY INFORMATION: Supplementary data are available at https://doi.org/10.5281/zenodo.6670164. Kerian Thuillier, Caroline Baroukh, Alexander Bockmayr, Ludovic Cottret, Loïc Paulevé, Anne Siegel |
Bioinform. | 6 |
| 2022 | Discrete modeling for integration and analysis of large-scale signaling networksabstractMost biological processes are orchestrated by large-scale molecular networks which are described in large-scale model repositories and whose dynamics are extremely complex. An observed phenotype is a state of this system that results from control mechanisms whose identification is key to its understanding. The Biological Pathway Exchange (BioPAX) format is widely used to standardize the biological information relative to regulatory processes. However, few modeling approaches developed so far enable for computing the events that control a phenotype in large-scale networks. Here we developed an integrated approach to build large-scale dynamic networks from BioPAX knowledge databases in order to analyse trajectories and to identify sets of biological entities that control a phenotype. The Cadbiom approach relies on the guarded transitions formalism, a discrete modeling approach which models a system dynamics by taking into account competition and cooperation events in chains of reactions. The method can be applied to every BioPAX (large-scale) model thanks to a specific package which automatically generates Cadbiom models from BioPAX files. The Cadbiom framework was applied to the BioPAX version of two resources (PID, KEGG) of the Pathway Commons database and to the Atlas of Cancer Signalling Network (ACSN). As a case-study, it was used to characterize sets of biological entities implicated in the epithelial-mesenchymal transition. Our results highlight the similarities between the PID and ACSN resources in terms of biological content, and underline the heterogeneity of usage of the BioPAX semantics limiting the fusion of models that require curation. Causality analyses demonstrate the smart complementarity of the databases in terms of combinatorics of controllers that explain a phenotype. From a biological perspective, our results show the specificity of controllers for epithelial and mesenchymal phenotypes that are consistent with the literature and identify a novel signature for intermediate states. Pierre Vignet, Jean Coquet, Sébastien Auber, Matéo Boudet, Anne Siegel, Nathalie Théret |
PLoS Comput. Biol. | 5 |
| 2021 | PAX2GRAPHML: a python library for large-scale regulation network analysis using BioPAXabstractSUMMARY: PAX2GRAPHML is an open-source Python library that allows to easily manipulate BioPAX source files as regulated reaction graphs described in.graphml format. The concept of regulated reactions, which allows connecting regulatory, signaling and metabolic levels, has been used. Biochemical reactions and regulatory interactions are homogeneously described by regulated reactions involving substrates, products, activators and inhibitors as elements. PAX2GRAPHML is highly flexible and allows generating graphs of regulated reactions from a single BioPAX source or by combining and filtering BioPAX sources. Supported by the graph exchange format .graphml, the large-scale graphs produced from one or more data sources can be further analyzed with PAX2GRAPHML or standard Python and R graph libraries. AVAILABILITY AND IMPLEMENTATION: https://pax2graphml.genouest.org. François Moreews, Hugo Simon, Anne Siegel, Florence Gondret, Emmanuelle Becker |
Bioinform. | 3 |
| 2021 | Constructing xenobiotic maps of metabolism to predict enzymes catalyzing metabolites capable of binding to DNAabstractBACKGROUND: The liver plays a major role in the metabolic activation of xenobiotics (drugs, chemicals such as pollutants, pesticides, food additives...). Among environmental contaminants of concern, heterocyclic aromatic amines (HAA) are xenobiotics classified by IARC as possible or probable carcinogens (2A or 2B). There exist little information about the effect of these HAA in humans. While HAA is a family of more than thirty identified chemicals, the metabolic activation and possible DNA adduct formation have been fully characterized in human liver for only a few of them (MeIQx, PhIP, A[Formula: see text]C). RESULTS: We have developed a modeling approach in order to predict all the possible metabolites of a xenobiotic and enzymatic profiles that are linked to the production of metabolites able to bind DNA. Our prediction of metabolites approach relies on the construction of an enriched and annotated map of metabolites from an input metabolite.The pipeline assembles reaction prediction tools (SyGMa), sites of metabolism prediction tools (Way2Drug, SOMP and Fame 3), a tool to estimate the ability of a xenobotics to form DNA adducts (XenoSite Reactivity V1), and a filtering procedure based on Bayesian framework. This prediction pipeline was evaluated using caffeine and then applied to HAA. The method was applied to determine enzymes profiles associated with the maximization of metabolites derived from each HAA which are able to bind to DNA. The classification of HAA according to enzymatic profiles was consistent with their chemical structures. CONCLUSIONS: Overall, a predictive toxicological model based on an in silico systems biology approach opens perspectives to estimate the genotoxicity of various chemical classes of environmental contaminants. Moreover, our approach based on enzymes profile determination opens the possibility of predicting various xenobiotics metabolites susceptible to bind to DNA in both normal and physiopathological situations. Maël Conan, Nathalie Théret, Sophie Langouet, Anne Siegel |
BMC Bioinform. | 4 |
| 2019 | Increasing Life Science Resources Re-Usability using Semantic Web TechnologiesabstractIn life sciences, current standardization and integration efforts are directed towards reference data and knowledge bases. However, original studies results are generally provided in non standardized and specific formats. In addition, the only formalization of analysis pipelines is often limited to textual descriptions in the method sections. Both factors impair the results reproducibility, their maintenance and their reuse for advancing other studies. Semantic Web technologies have proven their efficiency for facilitating the integration and reuse of reference data and knowledge bases. We thus hypothesize that Semantic Web technologies also facilitate reproducibility and reuse of life sciences studies involving pipelines that compute associations between entities according to intermediary relations and dependencies. In order to assess this hypothesis, we considered a case-study in systems biology (http://regulatorycircuits.org), which provides tissue-specific regulatory interaction networks to elucidate perturbations across complex diseases. Our approach consisted in surveying the complete set of provided supplementary files to reveal the underlying structure between the biological entities described in the data. We relied on this structure and used Semantic Web technologies (i) to integrate the Regulatory Circuits data, and (ii) to formalize the analysis pipeline as SPARQL queries. Our result was a 335,429,988 triples dataset on which two SPARQL queries were sufficient to extract each single tissuespecific regulatory network. Marine Louarn, Fabrice Chatonnet, Xavier Garnier, Thierry Fest, Anne Siegel, Olivier Dameron |
eScience | 5 |
| 2019 | Hybrid metabolic network completionabstractAbstract Metabolic networks play a crucial role in biology since they capture all chemical reactions in an organism. While there are networks of high quality for many model organisms, networks for less studied organisms are often of poor quality and suffer from incompleteness. To this end, we introduced in previous work an answer set programming (ASP)-based approach to metabolic network completion. Although this qualitative approach allows for restoring moderately degraded networks, it fails to restore highly degraded ones. This is because it ignores quantitative constraints capturing reaction rates. To address this problem, we propose a hybrid approach to metabolic network completion that integrates our qualitative ASP approach with quantitative means for capturing reaction rates. We begin by formally reconciling existing stoichiometric and topological approaches to network completion in a unified formalism. With it, we develop a hybrid ASP encoding and rely upon the theory reasoning capacities of the ASP system clingo for solving the resulting logic program with linear constraints over reals. We empirically evaluate our approach by means of the metabolic network of Escherichia coli . Our analysis shows that our novel approach yields greatly superior results than obtainable from purely qualitative or quantitative approaches. Clémence Frioux, Torsten Schaub, Sebastian Schellhorn, Anne Siegel, Philipp Wanko |
Theory Pract. Log. Program. | 4 |
| 2018 | Scalable and exhaustive screening of metabolic functions carried out by microbial consortiaabstractMotivation: The selection of species exhibiting metabolic behaviors of interest is a challenging step when switching from the investigation of a large microbiota to the study of functions effectiveness. Approaches based on a compartmentalized framework are not scalable. The output of scalable approaches based on a non-compartmentalized modeling may be so large that it has neither been explored nor handled so far. Results: We present the Miscoto tool to facilitate the selection of a community optimizing a desired function in a microbiome by reporting several possibilities which can be then sorted according to biological criteria. Communities are exhaustively identified using logical programming and by combining the non-compartmentalized and the compartmentalized frameworks. The benchmarking of 4.9 million metabolic functions associated with the Human Microbiome Project, shows that Miscoto is suited to screen and classify metabolic producibility in terms of feasibility, functional redundancy and cooperation processes involved. As an illustration of a host-microbial system, screening the Recon 2.2 human metabolism highlights the role of different consortia within a family of 773 intestinal bacteria. Availability and implementation: Miscoto source code, instructions for use and examples are available at: https://github.com/cfrioux/miscoto. Clémence Frioux, Enora Fremy, Camille Trottier, Anne Siegel |
Bioinform. | 4 |
| 2018 | Traceability, reproducibility and wiki-exploration for "à-la-carte" reconstructions of genome-scale metabolic modelsabstractGenome-scale metabolic models have become the tool of choice for the global analysis of microorganism metabolism, and their reconstruction has attained high standards of quality and reliability. Improvements in this area have been accompanied by the development of some major platforms and databases, and an explosion of individual bioinformatics methods. Consequently, many recent models result from "à la carte" pipelines, combining the use of platforms, individual tools and biological expertise to enhance the quality of the reconstruction. Although very useful, introducing heterogeneous tools, that hardly interact with each other, causes loss of traceability and reproducibility in the reconstruction process. This represents a real obstacle, especially when considering less studied species whose metabolic reconstruction can greatly benefit from the comparison to good quality models of related organisms. This work proposes an adaptable workspace, AuReMe, for sustainable reconstructions or improvements of genome-scale metabolic models involving personalized pipelines. At each step, relevant information related to the modifications brought to the model by a method is stored. This ensures that the process is reproducible and documented regardless of the combination of tools used. Additionally, the workspace establishes a way to browse metabolic models and their metadata through the automatic generation of ad-hoc local wikis dedicated to monitoring and facilitating the process of reconstruction. AuReMe supports exploration and semantic query based on RDF databases. We illustrate how this workspace allowed handling, in an integrated way, the metabolic reconstructions of non-model organisms such as an extremophile bacterium or eukaryote algae. Among relevant applications, the latter reconstruction led to putative evolutionary insights of a metabolic pathway. Méziane Aite, Marie Chevallier, Clémence Frioux, Camille Trottier, Jeanne Got, María Paz Cortés, Sebastián Nelson Mendoza, Grégory Carrier, Olivier Dameron, Nicolas Guillaudeux, Mauricio Latorre, Nicolás Loira, Gabriel V. Markov, Alejandro Maass, Anne Siegel |
PLoS Comput. Biol. | 15 |
| 2018 | Computational discovery of dynamic cell line specific Boolean networks from multiplex time-course dataabstractProtein signaling networks are static views of dynamic processes where proteins go through many biochemical modifications such as ubiquitination and phosphorylation to propagate signals that regulate cells and can act as feed-back systems. Understanding the precise mechanisms underlying protein interactions can elucidate how signaling and cell cycle progression occur within cells in different diseases such as cancer. Large-scale protein signaling networks contain an important number of experimentally verified protein relations but lack the capability to predict the outcomes of the system, and therefore to be trained with respect to experimental measurements. Boolean Networks (BNs) are a simple yet powerful framework to study and model the dynamics of the protein signaling networks. While many BN approaches exist to model biological systems, they focus mainly on system properties, and few exist to integrate experimental data in them. In this work, we show an application of a method conceived to integrate time series phosphoproteomic data into protein signaling networks. We use a large-scale real case study from the HPN-DREAM Breast Cancer challenge. Our efficient and parameter-free method combines logic programming and model-checking to infer a family of BNs from multiple perturbation time series data of four breast cancer cell lines given a prior protein signaling network. Because each predicted BN family is cell line specific, our method highlights commonalities and discrepancies between the four cell lines. Our models have a Root Mean Square Error (RMSE) of 0.31 with respect to the testing data, while the best performant method of this HPN-DREAM challenge had a RMSE of 0.47. To further validate our results, BNs are compared with the canonical mTOR pathway showing a comparable AUROC score (0.77) to the top performing HPN-DREAM teams. In addition, our approach can also be used as a complementary method to identify erroneous experiments. These results prove our methodology as an efficient dynamic model discovery method in multiple perturbation time course experimental data of large-scale signaling networks. The software and data are publicly available at https://github.com/misbahch6/caspo-ts. Misbah Razzaq, Loïc Paulevé, Anne Siegel, Julio Saez-Rodriguez, Jérémie Bourdon, Carito Guziolowski |
PLoS Comput. Biol. | 3 |
| 2017 | Hybrid Metabolic Network Completion
Clémence Frioux, Torsten Schaub, Sebastian Schellhorn, Anne Siegel, Philipp Wanko |
LPNMR | 4 |
| 2017 | caspo: a toolbox for automated reasoning on the response of logical signaling networks familiesabstractSummary: We introduce the caspo toolbox, a python package implementing a workflow for reasoning on logical networks families. Our software allows researchers to (i) a family of logical networks derived from a given topology and explaining the experimental response to various perturbations; (ii) all logical networks in a given family by their input-output behaviors; (iii) the response of the system to every possible perturbation based on the ensemble of predictions; (iv) new experimental perturbations to discriminate among a family of logical networks; and (v) a family of logical networks by finding all interventions strategies forcing a set of targets into a desired steady state. Availability and Implementation: caspo is open-source software distributed under the GPLv3 license. Source code is publicly hosted at http://github.com/bioasp/caspo . Contact: [email protected]. Santiago Videla, Julio Saez-Rodriguez, Carito Guziolowski, Anne Siegel |
Bioinform. | 4 |
| 2017 | Meneco, a Topology-Based Gap-Filling Tool Applicable to Degraded Genome-Wide Metabolic NetworksabstractIncreasing amounts of sequence data are becoming available for a wide range of non-model organisms. Investigating and modelling the metabolic behaviour of those organisms is highly relevant to understand their biology and ecology. As sequences are often incomplete and poorly annotated, draft networks of their metabolism largely suffer from incompleteness. Appropriate gap-filling methods to identify and add missing reactions are therefore required to address this issue. However, current tools rely on phenotypic or taxonomic information, or are very sensitive to the stoichiometric balance of metabolic reactions, especially concerning the co-factors. This type of information is often not available or at least prone to errors for newly-explored organisms. Here we introduce Meneco, a tool dedicated to the topological gap-filling of genome-scale draft metabolic networks. Meneco reformulates gap-filling as a qualitative combinatorial optimization problem, omitting constraints raised by the stoichiometry of a metabolic network considered in other methods, and solves this problem using Answer Set Programming. Run on several artificial test sets gathering 10,800 degraded Escherichia coli networks Meneco was able to efficiently identify essential reactions missing in networks at high degradation rates, outperforming the stoichiometry-based tools in scalability. To demonstrate the utility of Meneco we applied it to two case studies. Its application to recent metabolic networks reconstructed for the brown algal model Ectocarpus siliculosus and an associated bacterium Candidatus Phaeomarinobacter ectocarpi revealed several candidate metabolic pathways for algal-bacterial interactions. Then Meneco was used to reconstruct, from transcriptomic and metabolomic data, the first metabolic network for the microalga Euglena mutabilis. These two case studies show that Meneco is a versatile tool to complete draft genome-scale metabolic networks produced from heterogeneous data, and to suggest relevant reactions that explain the metabolic capacity of a biological system. Sylvain Prigent 0001, Clémence Frioux, Simon M. Dittami, Sven Thiele, Abdelhalim Larhlimi, Guillaume Collet, Fabien Gutknecht, Jeanne Got, Damien Eveillard, Jérémie Bourdon, Frédéric Plewniak, Thierry Tonon, Anne Siegel |
PLoS Comput. Biol. | 13 |
| 2016 | Deciphering transcriptional regulations coordinating the response to environmental changesabstractBACKGROUND: Gene co-expression evidenced as a response to environmental changes has shown that transcriptional activity is coordinated, which pinpoints the role of transcriptional regulatory networks (TRNs). Nevertheless, the prediction of TRNs based on the affinity of transcription factors (TFs) with binding sites (BSs) generally produces an over-estimation of the observable TF/BS relations within the network and therefore many of the predicted relations are spurious. RESULTS: We present LOMBARDE, a bioinformatics method that extracts from a TRN determined from a set of predicted TF/BS affinities a subnetwork explaining a given set of observed co-expressions by choosing the TFs and BSs most likely to be involved in the co-regulation. LOMBARDE solves an optimization problem which selects confident paths within a given TRN that join a putative common regulator with two co-expressed genes via regulatory cascades. To evaluate the method, we used public data of Escherichia coli to produce a regulatory network that explained almost all observed co-expressions while using only 19 % of the input TF/BS affinities but including about 66 % of the independent experimentally validated regulations in the input data. When all known validated TF/BS affinities were integrated into the input data the precision of LOMBARDE increased significantly. The topological characteristics of the subnetwork that was obtained were similar to the characteristics described for known validated TRNs. CONCLUSIONS: LOMBARDE provides a useful modeling scheme for deciphering the regulatory mechanisms that underlie the phenotypic responses of an organism to environmental challenges. The method can become a reliable tool for further research on genome-scale transcriptional regulation studies. Vicente Acuña, Andrés Aravena, Carito Guziolowski, Damien Eveillard, Anne Siegel, Alejandro Maass |
BMC Bioinform. | 5 |
| 2015 | Decidability Problems for Self-induced Systems Generated by a Substitution
Timo Jolivet, Anne Siegel |
MCU | 2 |
| 2015 | Extended notions of sign consistency to relate experimental data to signaling and regulatory network topologiesabstractBACKGROUND: A rapidly growing amount of knowledge about signaling and gene regulatory networks is available in databases such as KEGG, Reactome, or RegulonDB. There is an increasing need to relate this knowledge to high-throughput data in order to (in)validate network topologies or to decide which interactions are present or inactive in a given cell type under a particular environmental condition. Interaction graphs provide a suitable representation of cellular networks with information flows and methods based on sign consistency approaches have been shown to be valuable tools to (i) predict qualitative responses, (ii) to test the consistency of network topologies and experimental data, and (iii) to apply repair operations to the network model suggesting missing or wrong interactions. RESULTS: We present a framework to unify different notions of sign consistency and propose a refined method for data discretization that considers uncertainties in experimental profiles. We furthermore introduce a new constraint to filter undesired model behaviors induced by positive feedback loops. Finally, we generalize the way predictions can be made by the sign consistency approach. In particular, we distinguish strong predictions (e.g. increase of a node level) and weak predictions (e.g., node level increases or remains unchanged) enlarging the overall predictive power of the approach. We then demonstrate the applicability of our framework by confronting a large-scale gene regulatory network model of Escherichia coli with high-throughput transcriptomic measurements. CONCLUSION: Overall, our work enhances the flexibility and power of the sign consistency approach for the prediction of the behavior of signaling and gene regulatory networks and, more generally, for the validation and inference of these networks. Sven Thiele, Luca Cerone, Julio Saez-Rodriguez, Anne Siegel, Carito Guziolowski, Steffen Klamt |
BMC Bioinform. | 4 |
| 2015 | Learning Boolean logic models of signaling networks with ASP
Santiago Videla, Carito Guziolowski, Federica Eduati, Sven Thiele, Martin Gebser, Jacques Nicolas, Julio Saez-Rodriguez, Torsten Schaub, Anne Siegel |
Theor. Comput. Sci. | 9 |
| 2014 | Modeling Parsimonious Putative Regulatory Networks: Complexity and Heuristic Approach
Vicente Acuña, Andrés Aravena, Alejandro Maass, Anne Siegel |
VMCAI | 4 |
| 2014 | Exhaustively characterizing feasible logic models of a signaling network using Answer Set Programming
Carito Guziolowski, Santiago Videla, Federica Eduati, Sven Thiele, Thomas Cokelaer, Anne Siegel, Julio Saez-Rodriguez |
Bioinform. | 6 |
| 2013 | An ASP Application in Integrative Biology: Identification of Functional Gene Units
Philippe Bordron, Damien Eveillard, Alejandro Maass, Anne Siegel, Sven Thiele |
LPNMR | 4 |
| 2013 | Extending the Metabolic Network of Ectocarpus Siliculosus Using Answer Set Programming
Guillaume Collet, Damien Eveillard, Martin Gebser, Sylvain Prigent 0001, Torsten Schaub, Anne Siegel, Sven Thiele |
LPNMR | 6 |
| 2013 | Exhaustively characterizing feasible logic models of a signaling network using Answer Set ProgrammingabstractMOTIVATION: Logic modeling is a useful tool to study signal transduction across multiple pathways. Logic models can be generated by training a network containing the prior knowledge to phospho-proteomics data. The training can be performed using stochastic optimization procedures, but these are unable to guarantee a global optima or to report the complete family of feasible models. This, however, is essential to provide precise insight in the mechanisms underlaying signal transduction and generate reliable predictions. RESULTS: We propose the use of Answer Set Programming to explore exhaustively the space of feasible logic models. Toward this end, we have developed caspo, an open-source Python package that provides a powerful platform to learn and characterize logic models by leveraging the rich modeling language and solving technologies of Answer Set Programming. We illustrate the usefulness of caspo by revisiting a model of pro-growth and inflammatory pathways in liver cells. We show that, if experimental error is taken into account, there are thousands (11 700) of models compatible with the data. Despite the large number, we can extract structural features from the models, such as links that are always (or never) present or modules that appear in a mutual exclusive fashion. To further characterize this family of models, we investigate the input-output behavior of the models. We find 91 behaviors across the 11 700 models and we suggest new experiments to discriminate among them. Our results underscore the importance of characterizing in a global and exhaustive manner the family of feasible models, with important implications for experimental design. AVAILABILITY: caspo is freely available for download (license GPLv3) and as a web service at http://caspo.genouest.org/. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online. CONTACT: [email protected]. Carito Guziolowski, Santiago Videla, Federica Eduati, Sven Thiele, Thomas Cokelaer, Anne Siegel, Julio Saez-Rodriguez |
Bioinform. | 6 |
| 2013 | Minimal intervention strategies in logical signaling networks with ASPabstractAbstract Proposing relevant perturbations to biological signaling networks is central to many problems in biology and medicine because it allows for enabling or disabling certain biological outcomes. In contrast to quantitative methods that permit fine-grained (kinetic) analysis, qualitative approaches allow for addressing large-scale networks. This is accomplished by more abstract representations such as logical networks. We elaborate upon such a qualitative approach aiming at the computation of minimal interventions in logical signaling networks relying on Kleene's three-valued logic and fixpoint semantics. We address this problem within answer set programming and show that it greatly outperforms previous work using dedicated algorithms. Roland Kaminski, Torsten Schaub, Anne Siegel, Santiago Videla |
Theory Pract. Log. Program. | 3 |
| 2011 | Integrating Quantitative Knowledge into a Qualitative Gene Regulatory NetworkabstractDespite recent improvements in molecular techniques, biological knowledge remains incomplete. Any theorizing about living systems is therefore necessarily based on the use of heterogeneous and partial information. Much current research has focused successfully on the qualitative behaviors of macromolecular networks. Nonetheless, it is not capable of taking into account available quantitative information such as time-series protein concentration variations. The present work proposes a probabilistic modeling framework that integrates both kinds of information. Average case analysis methods are used in combination with Markov chains to link qualitative information about transcriptional regulations to quantitative information about protein concentrations. The approach is illustrated by modeling the carbon starvation response in Escherichia coli. It accurately predicts the quantitative time-series evolution of several protein concentrations using only knowledge of discrete gene interactions and a small number of quantitative observations on a single protein concentration. From this, the modeling technique also derives a ranking of interactions with respect to their importance during the experiment considered. Such a classification is confirmed by the literature. Therefore, our method is principally novel in that it allows (i) a hybrid model that integrates both qualitative discrete model and quantities to be built, even using a small amount of quantitative information, (ii) new quantitative predictions to be derived, (iii) the robustness and relevance of interactions with respect to phenotypic criteria to be precisely quantified, and (iv) the key features of the model to be extracted that can be used as a guidance to design future experiments. Jérémie Bourdon, Damien Eveillard, Anne Siegel |
PLoS Comput. Biol. | 3 |
| 2011 | Designing Logical Rules to Model the Response of Biomolecular Networks with Complex Interactions: An Application to Cancer ModelingabstractWe discuss the propagation of constraints in eukaryotic interaction networks in relation to model prediction and the identification of critical pathways. In order to cope with posttranslational interactions, we consider two types of nodes in the network, corresponding to proteins and to RNA. Microarray data provides very lacunar information for such types of networks because protein nodes, although needed in the model, are not observed. Propagation of observations in such networks leads to poor and nonsignificant model predictions, mainly because rules used to propagate information--usually disjunctive constraints--are weak. Here, we propose a new, stronger type of logical constraints that allow us to strengthen the analysis of the relation between microarray and interaction data. We use these rules to identify the nodes which are responsible for a phenotype, in particular for cell cycle progression. As the benchmark, we use an interaction network describing major pathways implied in Ewing's tumor development. The Python library used to obtain our results is publicly available on our supplementary web page. Carito Guziolowski, Sylvain Blachon, Tatiana Baumuratova, Gautier Stoll, Ovidiu Radulescu, Anne Siegel |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2010 | Repair and Prediction (under Inconsistency) in Large Biological Networks with Answer Set Programming
Martin Gebser, Carito Guziolowski, Mihail Ivanchev, Torsten Schaub, Anne Siegel, Sven Thiele, Philippe Veber |
KR | 5 |
| 2008 | Inferring the role of transcription factors in regulatory networksabstractBACKGROUND: Expression profiles obtained from multiple perturbation experiments are increasingly used to reconstruct transcriptional regulatory networks, from well studied, simple organisms up to higher eukaryotes. Admittedly, a key ingredient in developing a reconstruction method is its ability to integrate heterogeneous sources of information, as well as to comply with practical observability issues: measurements can be scarce or noisy. In this work, we show how to combine a network of genetic regulations with a set of expression profiles, in order to infer the functional effect of the regulations, as inducer or repressor. Our approach is based on a consistency rule between a network and the signs of variation given by expression arrays. RESULTS: We evaluate our approach in several settings of increasing complexity. First, we generate artificial expression data on a transcriptional network of E. coli extracted from the literature (1529 nodes and 3802 edges), and we estimate that 30% of the regulations can be annotated with about 30 profiles. We additionally prove that at most 40.8% of the network can be inferred using our approach. Second, we use this network in order to validate the predictions obtained with a compendium of real expression profiles. We describe a filtering algorithm that generates particularly reliable predictions. Finally, we apply our inference approach to S. cerevisiae transcriptional network (2419 nodes and 4344 interactions), by combining ChIP-chip data and 15 expression profiles. We are able to detect and isolate inconsistencies between the expression profiles and a significant portion of the model (15% of all the interactions). In addition, we report predictions for 14.5% of all interactions. CONCLUSION: Our approach does not require accurate expression levels nor times series. Nevertheless, we show on both data, real and artificial, that a relatively small number of perturbation experiments are enough to determine a significant portion of regulatory effects. This is a key practical asset compared to statistical methods for network reconstruction. We demonstrate that our approach is able to provide accurate predictions, even when the network is incomplete and the data is noisy. Philippe Veber, Carito Guziolowski, Michel Le Borgne, Ovidiu Radulescu, Anne Siegel |
BMC Bioinform. | 5 |
| 2004 | Two-dimensional iterated morphisms and discrete planes
Pierre Arnoux, Valérie Berthé, Anne Siegel |
Theor. Comput. Sci. | 3 |