Jan Hasenauer

dblp:87/8768 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
19since 2021 · last 2026
0000-0002-4935-3312ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 PEtab-GUI: a graphical user interface to create, edit, and inspect PEtab parameter estimation problems
abstract
MOTIVATION: Parameter estimation is a cornerstone of data-driven modeling in systems biology. Yet, constructing such problems in a reproducible and accessible manner remains challenging. The PEtab format has established itself as a powerful community standard to encode parameter estimation problems, promoting interoperability and reusability. However, its reliance on multiple interlinked files-often edited manually-can introduce inconsistencies, and new users often struggle to navigate them. Here, we present PEtab-GUI, an open-source Python application designed to streamline the creation, editing, and validation of PEtab problems through an intuitive graphical user interface. PEtab-GUI integrates all PEtab components, including SBML models and tabular files, into a single environment with live error checking and customizable defaults. Interactive visualization and simulation capabilities enable users to inspect the relationship between the model and the data. PEtab-GUI lowers the barrier to entry for specifying standardized parameter estimation problems, making dynamic modeling more accessible, especially in educational and interdisciplinary settings. AVAILABILITY AND IMPLEMENTATION: PEtab-GUI is implemented in Python, open-source under a 3-Clause BSD license. The code, designed to be modular and extensible, is hosted on https://github.com/PEtab-dev/PEtab-GUI, available as a Zenodo repository at https://doi.org/10.5281/zenodo.15355752, and can be installed from PyPI.
Paul J. Jost, Frank T. Bergmann, Daniel Weindl, Jan Hasenauer
Bioinform.4
2025 PEtab.jl: advancing the efficiency and utility of dynamic modelling
abstract
SUMMARY: Dynamic models represent a powerful tool for studying complex biological processes, ranging from cell signalling to cell differentiation. Building such models often requires computationally demanding modelling workflows, such as model exploration and parameter estimation. We developed two Julia-based tools: SBMLImporter.jl, an SBML importer, and PEtab.jl, an importer for parameter estimation problems in the PEtab format, designed to streamline modelling processes. These tools leverage Julia's high-performance computing capabilities, including symbolic pre-processing and advanced ODE solvers. PEtab.jl aims to be a Julia-accessible toolbox that supports the entire modelling pipeline from parameter estimation to identifiability analysis. AVAILABILITY AND IMPLEMENTATION: SBMLImporter.jl and PEtab.jl are implemented in the Julia programming language. Both packages are available on GitHub (github.com/sebapersson/SBMLImporter.jl and github.com/sebapersson/PEtab.jl) as officially registered Julia packages, installable via the Julia package manager. Each package is continuously tested and supported on Linux, macOS, and Windows.
Sebastian Persson, Fabian Fröhlich, Stephan Grein, Torkel Loman, Damiano Ognissanti, Viktor Hasselgren, Jan Hasenauer, Marija Cvijovic
Bioinform.7
2024 Data infrastructure for integrating clinical data in the large-scale international ORCHESTRA cohort: from data import to federated analysis
abstract
Large-scale international collaborations are increasingly managing large volumes of sensitive health data for research purposes. The use of such infrastructure requires fulfilling various technical and organizational requirements to ensure usability and security. Before data scientists can access the data, it must be imported into the infrastructure with high data security requirements. For federated analysis workflows, additional criteria, such as data harmonization and prevention of individual patient information disclosure, must also be met. This paper outlines the key components of the data infrastructure implemented in the European research project ORCHESTRA and elaborates on the methods that support federated analysis workflows in a heterogeneous legal environment. Special attention is given to data security, interoperability, and usability, from which data scientists and researchers would benefit. We demonstrate the usability of data infrastructure on a federated analysis and machine learning use cases on remote datasets which satisfies the previously mentioned requirements. Furthermore, we propose organizational measures to optimize the process, reducing the time between a data access request and granting access.
Miroslav Puskaric, Hammam Abu Attieh, Fabian Prasser, Roy Gusinow, Chiara Dellacasa, Elisa Rossi, Juan Mata Naranjo, Lorenzo Maria Canziani, Anna Górska, Jan Hasenauer
IEEE Big Data10
2024 An amortized approach to non-linear mixed-effects modeling based on neural posterior estimation
abstract
Non-linear mixed-effects models are a powerful tool for studying heterogeneous populations in various fields, including biology, medicine, economics, and engineering. Here, the aim is to find a distribution over the parameters that describe the whole population using a model that can generate simulations for an individual of that population. However, fitting these distributions to data is computationally challenging if the description of individuals is complex and the population is large. To address this issue, we propose a novel machine learning-based approach: We exploit neural density estimation based on conditional normalizing flows to approximate individual-specific posterior distributions in an amortized fashion, thereby allowing for efficient inference of population parameters. Applying this approach to problems from cell biology and pharmacology, we demonstrate its unseen flexibility and scalability to large data sets compared to established methods.
Jonas Arruda, Yannik Schälte, Clemens Peiter, Olga Teplytska, Ulrich Jaehde, Jan Hasenauer
ICML6
2024 Efficient parameter estimation for ODE models of cellular processes using semi-quantitative data
abstract
MOTIVATION: Quantitative dynamical models facilitate the understanding of biological processes and the prediction of their dynamics. The parameters of these models are commonly estimated from experimental data. Yet, experimental data generated from different techniques do not provide direct information about the state of the system but a nonlinear (monotonic) transformation of it. For such semi-quantitative data, when this transformation is unknown, it is not apparent how the model simulations and the experimental data can be compared. RESULTS: We propose a versatile spline-based approach for the integration of a broad spectrum of semi-quantitative data into parameter estimation. We derive analytical formulas for the gradients of the hierarchical objective function and show that this substantially increases the estimation efficiency. Subsequently, we demonstrate that the method allows for the reliable discovery of unknown measurement transformations. Furthermore, we show that this approach can significantly improve the parameter inference based on semi-quantitative data in comparison to available methods. AVAILABILITY AND IMPLEMENTATION: Modelers can easily apply our method by using our implementation in the open-source Python Parameter EStimation TOolbox (pyPESTO) available at https://github.com/ICB-DCM/pyPESTO.
Domagoj Doresic, Stephan Grein, Jan Hasenauer
Bioinform.3
2024 Missing data in amortized simulation-based neural posterior estimation
abstract
Amortized simulation-based neural posterior estimation provides a novel machine learning based approach for solving parameter estimation problems. It has been shown to be computationally efficient and able to handle complex models and data sets. Yet, the available approach cannot handle the in experimental studies ubiquitous case of missing data, and might provide incorrect posterior estimates. In this work, we discuss various ways of encoding missing data and integrate them into the training and inference process. We implement the approaches in the BayesFlow methodology, an amortized estimation framework based on invertible neural networks, and evaluate their performance on multiple test problems. We find that an approach in which the data vector is augmented with binary indicators of presence or absence of values performs the most robustly. Indeed, it improved the performance also for the simpler problem of data sets with variable length. Accordingly, we demonstrate that amortized simulation-based inference approaches are applicable even with missing data, and we provide a guideline for their handling, which is relevant for a broad spectrum of applications.
Jan Hasenauer, Yannik Schälte
PLoS Comput. Biol.2
2023 FitMultiCell: simulating and parameterizing computational models of multi-scale and multi-cellular processes
abstract
MOTIVATION: Biological tissues are dynamic and highly organized. Multi-scale models are helpful tools to analyse and understand the processes determining tissue dynamics. These models usually depend on parameters that need to be inferred from experimental data to achieve a quantitative understanding, to predict the response to perturbations, and to evaluate competing hypotheses. However, even advanced inference approaches such as approximate Bayesian computation (ABC) are difficult to apply due to the computational complexity of the simulation of multi-scale models. Thus, there is a need for a scalable pipeline for modeling, simulating, and parameterizing multi-scale models of multi-cellular processes. RESULTS: Here, we present FitMultiCell, a computationally efficient and user-friendly open-source pipeline that can handle the full workflow of modeling, simulating, and parameterizing for multi-scale models of multi-cellular processes. The pipeline is modular and integrates the modeling and simulation tool Morpheus and the statistical inference tool pyABC. The easy integration of high-performance infrastructure allows to scale to computationally expensive problems. The introduction of a novel standard for the formulation of parameter inference problems for multi-scale models additionally ensures reproducibility and reusability. By applying the pipeline to multiple biological problems, we demonstrate its broad applicability, which will benefit in particular image-based systems biology. AVAILABILITY AND IMPLEMENTATION: FitMultiCell is available open-source at https://gitlab.com/fitmulticell/fit.
Emad Alamoudi, Yannik Schälte, Jörn Starruß, Nils Bundgaard, Frederik Graw, Lutz Brusch, Jan Hasenauer
Bioinform.8
2023 Accessibility of covariance information creates vulnerability in Federated Learning frameworks
abstract
MOTIVATION: Federated Learning (FL) is gaining traction in various fields as it enables integrative data analysis without sharing sensitive data, such as in healthcare. However, the risk of data leakage caused by malicious attacks must be considered. In this study, we introduce a novel attack algorithm that relies on being able to compute sample means, sample covariances, and construct known linearly independent vectors on the data owner side. RESULTS: We show that these basic functionalities, which are available in several established FL frameworks, are sufficient to reconstruct privacy-protected data. Additionally, the attack algorithm is robust to defense strategies that involve adding random noise. We demonstrate the limitations of existing frameworks and propose potential defense strategies analyzing the implications of using differential privacy. The novel insights presented in this study will aid in the improvement of FL frameworks. AVAILABILITY AND IMPLEMENTATION: The code examples are provided at GitHub (https://github.com/manuhuth/Data-Leakage-From-Covariances.git). The CNSIM1 dataset, which we used in the manuscript, is available within the DSData R package (https://github.com/datashield/DSData/tree/main/data).
Manuel Huth, Jonas Arruda, Roy Gusinow, Lorenzo Contento, Evelina Tacconelli, Jan Hasenauer
Bioinform.6
2023 pyPESTO: a modular and scalable tool for parameter estimation for dynamic models
abstract
SUMMARY: Mechanistic models are important tools to describe and understand biological processes. However, they typically rely on unknown parameters, the estimation of which can be challenging for large and complex systems. pyPESTO is a modular framework for systematic parameter estimation, with scalable algorithms for optimization and uncertainty quantification. While tailored to ordinary differential equation problems, pyPESTO is broadly applicable to black-box parameter estimation problems. Besides own implementations, it provides a unified interface to various popular simulation and inference methods. AVAILABILITY AND IMPLEMENTATION: pyPESTO is implemented in Python, open-source under a 3-Clause BSD license. Code and documentation are available on GitHub (https://github.com/icb-dcm/pypesto).
Yannik Schälte, Fabian Fröhlich, Paul J. Jost, Jakob Vanhoefer, Dilan Pathirana, Paul Stapor, Polina A. Lakrisenko, Dantong Wang, Elba Raimúndez-Álvarez, Simon Merkt, Leonard Schmiester, Philipp Städter, Stephan Grein, Erika Dudkin, Domagoj Doresic, Daniel Weindl, Jan Hasenauer
Bioinform.17
2023 IntestLine: a shiny-based application to map the rolled intestinal tissue onto a line
abstract
SUMMARY: To allow the comprehensive histological analysis of the whole intestine, it is often rolled to a spiral before imaging. This Swiss-rolling technique facilitates robust experimental procedures, but it limits the possibilities to comprehend changes along the intestine. Here, we present IntestLine, a Shiny-based open-source application for processing imaging data of (rolled) intestinal tissues and subsequent mapping onto a line. The visualization of the mapped data facilitates the assessment of the whole intestine in both proximal-distal and serosa-luminal axis, and enables the observation of location-specific cell types and markers. Accordingly, IntestLine can serve as a tool to characterize the intestine in multi-modal imaging studies. AVAILABILITY AND IMPLEMENTATION: Source code can be found at Zenodo (https://doi.org/10.5281/zenodo.7081864) and GitHub (https://github.com/SchlitzerLab/IntestLine).
Altay Yuzeir, David Alejandro Bejarano, Stephan Grein, Jan Hasenauer, Andreas Schlitzer, Jiangyan Yu
Bioinform.4
2023 Validation of genetic variants from NGS data using deep convolutional neural networks
abstract
Accurate somatic variant calling from next-generation sequencing data is one most important tasks in personalised cancer therapy. The sophistication of the available technologies is ever-increasing, yet, manual candidate refinement is still a necessary step in state-of-the-art processing pipelines. This limits reproducibility and introduces a bottleneck with respect to scalability. We demonstrate that the validation of genetic variants can be improved using a machine learning approach resting on a Convolutional Neural Network, trained using existing human annotation. In contrast to existing approaches, we introduce a way in which contextual data from sequencing tracks can be included into the automated assessment. A rigorous evaluation shows that the resulting model is robust and performs on par with trained researchers following published standard operating procedure.
Marc Vaisband, Maria Schubert, Franz Josef Gassner, Roland Geisberger, Richard Greil, Nadja Zaborsky, Jan Hasenauer
BMC Bioinform.7
2023 Efficient computation of adjoint sensitivities at steady-state in ODE models of biochemical reaction networks
abstract
Dynamical models in the form of systems of ordinary differential equations have become a standard tool in systems biology. Many parameters of such models are usually unknown and have to be inferred from experimental data. Gradient-based optimization has proven to be effective for parameter estimation. However, computing gradients becomes increasingly costly for larger models, which are required for capturing the complex interactions of multiple biochemical pathways. Adjoint sensitivity analysis has been pivotal for working with such large models, but methods tailored for steady-state data are currently not available. We propose a new adjoint method for computing gradients, which is applicable if the experimental data include steady-state measurements. The method is based on a reformulation of the backward integration problem to a system of linear algebraic equations. The evaluation of the proposed method using real-world problems shows a speedup of total simulation time by a factor of up to 4.4. Our results demonstrate that the proposed approach can achieve a substantial improvement in computation time, in particular for large-scale models, where computational efficiency is critical.
Polina A. Lakrisenko, Paul Stapor, Stephan Grein, Lukasz Paszkowski, Dilan Pathirana, Fabian Fröhlich, Glenn Terje Lines, Daniel Weindl, Jan Hasenauer
PLoS Comput. Biol.9
2023 Assessment of Prediction Uncertainty Quantification Methods in Systems Biology
abstract
Biological processes are often modelled using ordinary differential equations. The unknown parameters of these models are estimated by optimizing the fit of model simulation and experimental data. The resulting parameter estimates inevitably possess some degree of uncertainty. In practical applications it is important to quantify these parameter uncertainties as well as the resulting prediction uncertainty, which are uncertainties of potentially time-dependent model characteristics. Unfortunately, estimating prediction uncertainties accurately is nontrivial, due to the nonlinear dependence of model characteristics on parameters. While a number of numerical approaches have been proposed for this task, their strengths and weaknesses have not been systematically assessed yet. To fill this knowledge gap, we apply four state of the art methods for uncertainty quantification to four case studies of different computational complexities. This reveals the trade-offs between their applicability and their statistical interpretability. Our results provide guidelines for choosing the most appropriate technique for a given problem and applying it successfully.
Alejandro Fernández Villaverde, Elba Raimúndez-Álvarez, Jan Hasenauer, Julio R. Banga
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 A protocol for dynamic model calibration
abstract
Ordinary differential equation models are nowadays widely used for the mechanistic description of biological processes and their temporal evolution. These models typically have many unknown and nonmeasurable parameters, which have to be determined by fitting the model to experimental data. In order to perform this task, known as parameter estimation or model calibration, the modeller faces challenges such as poor parameter identifiability, lack of sufficiently informative experimental data and the existence of local minima in the objective function landscape. These issues tend to worsen with larger model sizes, increasing the computational complexity and the number of unknown parameters. An incorrectly calibrated model is problematic because it may result in inaccurate predictions and misleading conclusions. For nonexpert users, there are a large number of potential pitfalls. Here, we provide a protocol that guides the user through all the steps involved in the calibration of dynamic models. We illustrate the methodology with two models and provide all the code required to reproduce the results and perform the same analysis on new models. Our protocol provides practitioners and researchers in biological modelling with a one-stop guide that is at the same time compact and sufficiently comprehensive to cover all aspects of the problem.
Alejandro Fernández Villaverde, Dilan Pathirana, Fabian Fröhlich, Jan Hasenauer, Julio R. Banga
Briefings Bioinform.4
2021 Closing the gap between formats for storing layout information in systems biology
abstract
The first version of this article listed one of its authors as Jan Hausenauer rather than Jan Hasenauer. This has now been corrected. The authors regret the error.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.4
2021 SysMod: the ISCB community for data-driven computational modelling and multi-scale analysis of biological systems
abstract
Computational models of biological systems can exploit a broad range of rapidly developing approaches, including novel experimental approaches, bioinformatics data analysis, emerging modelling paradigms, data standards and algorithms. A discussion about the most recent advances among experts from various domains is crucial to foster data-driven computational modelling and its growing use in assessing and predicting the behaviour of biological systems. Intending to encourage the development of tools, approaches and predictive models, and to deepen our understanding of biological systems, the Community of Special Interest (COSI) was launched in Computational Modelling of Biological Systems (SysMod) in 2016. SysMod's main activity is an annual meeting at the Intelligent Systems for Molecular Biology (ISMB) conference, which brings together computer scientists, biologists, mathematicians, engineers, computational and systems biologists. In the five years since its inception, SysMod has evolved into a dynamic and expanding community, as the increasing number of contributions and participants illustrate. SysMod maintains several online resources to facilitate interaction among the community members, including an online forum, a calendar of relevant meetings and a YouTube channel with talks and lectures of interest for the modelling community. For more than half a decade, the growing interest in computational systems modelling and multi-scale data integration has inspired and supported the SysMod community. Its members get progressively more involved and actively contribute to the annual COSI meeting and several related community workshops and meetings, focusing on specific topics, including particular techniques for computational modelling or standardisation efforts.
Andreas Dräger, Tomás Helikar, Matteo Barberis, Marc R. Birtwistle, Laurence Calzone, Claudine Chaouiya, Jan Hasenauer, Jonathan R. Karr, Anna Niarakis, María Rodríguez Martínez, Julio Saez-Rodriguez, Juilee Thakar
Bioinform.7
2021 AMICI: high-performance sensitivity analysis for large ordinary differential equation models
abstract
Ordinary differential equation models facilitate the understanding of cellular signal transduction and other biological processes. However, for large and comprehensive models, the computational cost of simulating or calibrating can be limiting. AMICI is a modular toolbox implemented in C++/Python/MATLAB that provides efficient simulation and sensitivity analysis routines tailored for scalable, gradient-based parameter estimation and uncertainty quantification. AMICI is published under the permissive BSD-3-Clause license with source code publicly available on https://github.com/AMICI-dev/AMICI. Citeable releases are archived on Zenodo.
Fabian Fröhlich, Daniel Weindl, Yannik Schälte, Dilan Pathirana, Lukasz Paszkowski, Glenn Terje Lines, Paul Stapor, Jan Hasenauer
Bioinform.8
2021 Efficient gradient-based parameter estimation for dynamic models using qualitative data
abstract
MOTIVATION: Unknown parameters of dynamical models are commonly estimated from experimental data. However, while various efficient optimization and uncertainty analysis methods have been proposed for quantitative data, methods for qualitative data are rare and suffer from bad scaling and convergence. RESULTS: Here, we propose an efficient and reliable framework for estimating the parameters of ordinary differential equation models from qualitative data. In this framework, we derive a semi-analytical algorithm for gradient calculation of the optimal scaling method developed for qualitative data. This enables the use of efficient gradient-based optimization algorithms. We demonstrate that the use of gradient information improves performance of optimization and uncertainty quantification on several application examples. On average, we achieve a speedup of more than one order of magnitude compared to gradient-free optimization. In addition, in some examples, the gradient-based approach yields substantially improved objective function values and quality of the fits. Accordingly, the proposed framework substantially improves the parameterization of models from qualitative data. AVAILABILITY AND IMPLEMENTATION: The proposed approach is implemented in the open-source Python Parameter EStimation TOolbox (pyPESTO). pyPESTO is available at https://github.com/ICB-DCM/pyPESTO. All application examples and code to reproduce this study are available at https://doi.org/10.5281/zenodo.4507613. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leonard Schmiester, Daniel Weindl, Jan Hasenauer
Bioinform.3
2021 PEtab - Interoperable specification of parameter estimation problems in systems biology
abstract
Reproducibility and reusability of the results of data-based modeling studies are essential. Yet, there has been-so far-no broadly supported format for the specification of parameter estimation problems in systems biology. Here, we introduce PEtab, a format which facilitates the specification of parameter estimation problems using Systems Biology Markup Language (SBML) models and a set of tab-separated value files describing the observation model and experimental data as well as parameters to be estimated. We already implemented PEtab support into eight well-established model simulation and parameter estimation toolboxes with hundreds of users in total. We provide a Python library for validation and modification of a PEtab problem and currently 20 example parameter estimation problems based on recent studies.
Leonard Schmiester, Yannik Schälte, Frank T. Bergmann, Tacio Camba, Erika Dudkin, Janine Egert, Fabian Fröhlich, Lara Fuhrmann, Adrian L. Hauber, Svenja Kemmer, Polina A. Lakrisenko, Carolin Loos, Simon Merkt, Wolfgang Müller 0001, Dilan Pathirana, Elba Raimúndez-Álvarez, Lukas Refisch, Marcus Rosenblatt, Paul Stapor, Philipp Städter, Dantong Wang, Franz-Georg Wieland, Julio R. Banga, Jens Timmer, Alejandro Fernández Villaverde, Sven Sahle, Clemens Kreutz, Jan Hasenauer, Daniel Weindl
PLoS Comput. Biol.28
2020 Closing the gap between formats for storing layout information in systems biology
abstract
The understanding of complex biological networks often relies on both a dedicated layout and a topology. Currently, there are three major competing layout-aware systems biology formats, but there are no software tools or software libraries supporting all of them. This complicates the management of molecular network layouts and hinders their reuse and extension. In this paper, we present a high-level overview of the layout formats in systems biology, focusing on their commonalities and differences, review their support in existing software tools, libraries and repositories and finally introduce a new conversion module within the MINERVA platform. The module is available via a REST API and offers, besides the ability to convert between layout-aware systems biology formats, the possibility to export layouts into several graphical formats. The module enables conversion of very large networks with thousands of elements, such as disease maps or metabolic reconstructions, rendering it widely applicable in systems biology.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.4
2020 Efficient exact inference for dynamical systems with noisy measurements using sequential approximate Bayesian computation
abstract
MOTIVATION: Approximate Bayesian computation (ABC) is an increasingly popular method for likelihood-free parameter inference in systems biology and other fields of research, as it allows analyzing complex stochastic models. However, the introduced approximation error is often not clear. It has been shown that ABC actually gives exact inference under the implicit assumption of a measurement noise model. Noise being common in biological systems, it is intriguing to exploit this insight. But this is difficult in practice, as ABC is in general highly computationally demanding. Thus, the question we want to answer here is how to efficiently account for measurement noise in ABC. RESULTS: We illustrate exemplarily how ABC yields erroneous parameter estimates when neglecting measurement noise. Then, we discuss practical ways of correctly including the measurement noise in the analysis. We present an efficient adaptive sequential importance sampling-based algorithm applicable to various model types and noise models. We test and compare it on several models, including ordinary and stochastic differential equations, Markov jump processes and stochastically interacting agents, and noise models including normal, Laplace and Poisson noise. We conclude that the proposed algorithm could improve the accuracy of parameter estimates for a broad spectrum of applications. AVAILABILITY AND IMPLEMENTATION: The developed algorithms are made publicly available as part of the open-source python toolbox pyABC (https://github.com/icb-dcm/pyabc). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yannik Schälte, Jan Hasenauer
Bioinform.2
2020 Efficient parameterization of large-scale dynamic models based on relative measurements
abstract
MOTIVATION: Mechanistic models of biochemical reaction networks facilitate the quantitative understanding of biological processes and the integration of heterogeneous datasets. However, some biological processes require the consideration of comprehensive reaction networks and therefore large-scale models. Parameter estimation for such models poses great challenges, in particular when the data are on a relative scale. RESULTS: Here, we propose a novel hierarchical approach combining (i) the efficient analytic evaluation of optimal scaling, offset and error model parameters with (ii) the scalable evaluation of objective function gradients using adjoint sensitivity analysis. We evaluate the properties of the methods by parameterizing a pan-cancer ordinary differential equation model (>1000 state variables, >4000 parameters) using relative protein, phosphoprotein and viability measurements. The hierarchical formulation improves optimizer performance considerably. Furthermore, we show that this approach allows estimating error model parameters with negligible computational overhead when no experimental estimates are available, providing an unbiased way to weight heterogeneous data. Overall, our hierarchical formulation is applicable to a wide range of models, and allows for the efficient parameterization of large-scale models based on heterogeneous relative measurements. AVAILABILITY AND IMPLEMENTATION: Supplementary code and data are available online at http://doi.org/10.5281/zenodo.3254429 and http://doi.org/10.5281/zenodo.3254441. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Leonard Schmiester, Yannik Schälte, Fabian Fröhlich, Jan Hasenauer, Daniel Weindl
Bioinform.4
2020 Model-based analysis of response and resistance factors of cetuximab treatment in gastric cancer cell lines
abstract
Targeted cancer therapies are powerful alternatives to chemotherapies or can be used complementary to these. Yet, the response to targeted treatments depends on a variety of factors, including mutations and expression levels, and therefore their outcome is difficult to predict. Here, we develop a mechanistic model of gastric cancer to study response and resistance factors for cetuximab treatment. The model captures the EGFR, ERK and AKT signaling pathways in two gastric cancer cell lines with different mutation patterns. We train the model using a comprehensive selection of time and dose response measurements, and provide an assessment of parameter and prediction uncertainties. We demonstrate that the proposed model facilitates the identification of causal differences between the cell lines. Furthermore, our study shows that the model provides predictions for the responses to different perturbations, such as knockdown and knockout experiments. Among other results, the model predicted the effect of MET mutations on cetuximab sensitivity. These predictive capabilities render the model a basis for the assessment of gastric cancer signaling and possibly for the development and discovery of predictive biomarkers.
Elba Raimúndez-Álvarez, Simone Keller, Gwen Zwingenberger, Karolin Ebert, Sabine Hug, Fabian J. Theis, Dieter Maier, Birgit Luber, Jan Hasenauer
PLoS Comput. Biol.9
2019 Community-driven roadmap for integrated disease maps
abstract
The Disease Maps Project builds on a network of scientific and clinical groups that exchange best practices, share information and develop systems biomedicine tools. The project aims for an integrated, highly curated and user-friendly platform for disease-related knowledge. The primary focus of disease maps is on interconnected signaling, metabolic and gene regulatory network pathways represented in standard formats. The involvement of domain experts ensures that the key disease hallmarks are covered and relevant, up-to-date knowledge is adequately represented. Expert-curated and computer readable, disease maps may serve as a compendium of knowledge, allow for data-supported hypothesis generation or serve as a scaffold for the generation of predictive mathematical models. This article summarizes the 2nd Disease Maps Community meeting, highlighting its important topics and outcomes. We outline milestones on the roadmap for the future development of disease maps, including creating and maintaining standardized disease maps; sharing parts of maps that encode common human disease mechanisms; providing technical solutions for complexity management of maps; and Web tools for in-depth exploration of such maps. A dedicated discussion was focused on mathematical modeling approaches, as one of the main goals of disease map development is the generation of mathematically interpretable representations to predict disease comorbidity or drug response and to suggest drug repositioning, altogether supporting clinical decisions.
Marek Ostaszewski, Stephan Gebel, Inna Kuperstein, Alexander Mazein, Andrei Yu. Zinovyev, Ugur Dogrusoz, Jan Hasenauer, Ronan M. T. Fleming, Nicolas Le Novère, Piotr Gawron, Thomas S. Ligon, Anna Niarakis, David P. Nickerson, Daniel Weindl, Rudi Balling, Emmanuel Barillot, Charles Auffray, Reinhard Schneider 0002
Briefings Bioinform.7
2019 Benchmark problems for dynamic modeling of intracellular processes
abstract
MOTIVATION: Dynamic models are used in systems biology to study and understand cellular processes like gene regulation or signal transduction. Frequently, ordinary differential equation (ODE) models are used to model the time and dose dependency of the abundances of molecular compounds as well as interactions and translocations. A multitude of computational approaches, e.g. for parameter estimation or uncertainty analysis have been developed within recent years. However, many of these approaches lack proper testing in application settings because a comprehensive set of benchmark problems is yet missing. RESULTS: We present a collection of 20 benchmark problems in order to evaluate new and existing methodologies, where an ODE model with corresponding experimental data is referred to as problem. In addition to the equations of the dynamical system, the benchmark collection provides observation functions as well as assumptions about measurement noise distributions and parameters. The presented benchmark models comprise problems of different size, complexity and numerical demands. Important characteristics of the models and methodological requirements are summarized, estimated parameters are provided, and some example studies were performed for illustrating the capabilities of the presented benchmark collection. AVAILABILITY AND IMPLEMENTATION: The models are provided in several standardized formats, including an easy-to-use human readable form and machine-readable SBML files. The data is provided as Excel sheets. All files are available at https://github.com/Benchmarking-Initiative/Benchmark-Models, including step-by-step explanations and MATLAB code to process and simulate the models. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Helge Hass, Carolin Loos, Elba Raimúndez-Álvarez, Jens Timmer, Jan Hasenauer, Clemens Kreutz
Bioinform.5
2019 Benchmarking optimization methods for parameter estimation in large kinetic models
abstract
MOTIVATION: Kinetic models contain unknown parameters that are estimated by optimizing the fit to experimental data. This task can be computationally challenging due to the presence of local optima and ill-conditioning. While a variety of optimization methods have been suggested to surmount these issues, it is difficult to choose the best one for a given problem a priori. A systematic comparison of parameter estimation methods for problems with tens to hundreds of optimization variables is currently missing, and smaller studies provided contradictory findings. RESULTS: We use a collection of benchmarks to evaluate the performance of two families of optimization methods: (i) multi-starts of deterministic local searches and (ii) stochastic global optimization metaheuristics; the latter may be combined with deterministic local searches, leading to hybrid methods. A fair comparison is ensured through a collaborative evaluation and a consideration of multiple performance metrics. We discuss possible evaluation criteria to assess the trade-off between computational efficiency and robustness. Our results show that, thanks to recent advances in the calculation of parametric sensitivities, a multi-start of gradient-based local methods is often a successful strategy, but a better performance can be obtained with a hybrid metaheuristic. The best performer combines a global scatter search metaheuristic with an interior point local method, provided with gradients estimated with adjoint-based sensitivities. We provide an implementation of this method to render it available to the scientific community. AVAILABILITY AND IMPLEMENTATION: The code to reproduce the results is provided as Supplementary Material and is available at Zenodo https://doi.org/10.5281/zenodo.1304034. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alejandro Fernández Villaverde, Fabian Fröhlich, Daniel Weindl, Jan Hasenauer, Julio R. Banga
Bioinform.4
2018 Bayesian parameter estimation for biochemical reaction networks using region-based adaptive parallel tempering
abstract
Motivation: Mathematical models have become standard tools for the investigation of cellular processes and the unraveling of signal processing mechanisms. The parameters of these models are usually derived from the available data using optimization and sampling methods. However, the efficiency of these methods is limited by the properties of the mathematical model, e.g. non-identifiabilities, and the resulting posterior distribution. In particular, multi-modal distributions with long valleys or pronounced tails are difficult to optimize and sample. Thus, the developement or improvement of optimization and sampling methods is subject to ongoing research. Results: We suggest a region-based adaptive parallel tempering algorithm which adapts to the problem-specific posterior distributions, i.e. modes and valleys. The algorithm combines several established algorithms to overcome their individual shortcomings and to improve sampling efficiency. We assessed its properties for established benchmark problems and two ordinary differential equation models of biochemical reaction networks. The proposed algorithm outperformed state-of-the-art methods in terms of calculation efficiency and mixing. Since the algorithm does not rely on a specific problem structure, but adapts to the posterior distribution, it is suitable for a variety of model classes. Availability and implementation: The code is available both as Supplementary Material and in a Git repository written in MATLAB. Supplementary information: Supplementary data are available at Bioinformatics online.
Benjamin Ballnus, Steffen Schaper, Fabian J. Theis, Jan Hasenauer
Bioinform.4
2018 pyABC: distributed, likelihood-free inference
abstract
Summary: Likelihood-free methods are often required for inference in systems biology. While approximate Bayesian computation (ABC) provides a theoretical solution, its practical application has often been challenging due to its high computational demands. To scale likelihood-free inference to computationally demanding stochastic models, we developed pyABC: a distributed and scalable ABC-Sequential Monte Carlo (ABC-SMC) framework. It implements a scalable, runtime-minimizing parallelization strategy for multi-core and distributed environments scaling to thousands of cores. The framework is accessible to non-expert users and also enables advanced users to experiment with and to custom implement many options of ABC-SMC schemes, such as acceptance threshold schedules, transition kernels and distance functions without alteration of pyABC's source code. pyABC includes a web interface to visualize ongoing and finished ABC-SMC runs and exposes an API for data querying and post-processing. Availability and Implementation: pyABC is written in Python 3 and is released under a 3-clause BSD license. The source code is hosted on https://github.com/icb-dcm/pyabc and the documentation on http://pyabc.readthedocs.io. It can be installed from the Python Package Index (PyPI). Supplementary information: Supplementary data are available at Bioinformatics online.
Emmanuel Klinger, Dennis Rickert, Jan Hasenauer
Bioinform.3
2018 GenSSI 2.0: multi-experiment structural identifiability analysis of SBML models
abstract
Motivation: Mathematical modeling using ordinary differential equations is used in systems biology to improve the understanding of dynamic biological processes. The parameters of ordinary differential equation models are usually estimated from experimental data. To analyze a priori the uniqueness of the solution of the estimation problem, structural identifiability analysis methods have been developed. Results: We introduce GenSSI 2.0, an advancement of the software toolbox GenSSI (Generating Series for testing Structural Identifiability). GenSSI 2.0 is the first toolbox for structural identifiability analysis to implement Systems Biology Markup Language import, state/parameter transformations and multi-experiment structural identifiability analysis. In addition, GenSSI 2.0 supports a range of MATLAB versions and is computationally more efficient than its previous version, enabling the analysis of more complex models. Availability and implementation: GenSSI 2.0 is an open-source MATLAB toolbox and available at https://github.com/genssi-developer/GenSSI. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Thomas S. Ligon, Fabian Fröhlich, Oana Chis, Julio R. Banga, Eva Balsa-Canto, Jan Hasenauer
Bioinform.6
2018 Hierarchical optimization for the efficient parametrization of ODE models
abstract
Motivation: Mathematical models are nowadays important tools for analyzing dynamics of cellular processes. The unknown model parameters are usually estimated from experimental data. These data often only provide information about the relative changes between conditions, hence, the observables contain scaling parameters. The unknown scaling parameters and corresponding noise parameters have to be inferred along with the dynamic parameters. The nuisance parameters often increase the dimensionality of the estimation problem substantially and cause convergence problems. Results: In this manuscript, we propose a hierarchical optimization approach for estimating the parameters for ordinary differential equation (ODE) models from relative data. Our approach restructures the optimization problem into an inner and outer subproblem. These subproblems possess lower dimensions than the original optimization problem, and the inner problem can be solved analytically. We evaluated accuracy, robustness and computational efficiency of the hierarchical approach by studying three signaling pathways. The proposed approach achieved better convergence than the standard approach and required a lower computation time. As the hierarchical optimization approach is widely applicable, it provides a powerful alternative to established approaches. Availability and implementation: The code is included in the MATLAB toolbox PESTO which is available at http://github.com/ICB-DCM/PESTO. Supplementary information: Supplementary data are available at Bioinformatics online.
Carolin Loos, Sabrina Krause, Jan Hasenauer
Bioinform.3
2018 Optimization and profile calculation of ODE models using second order adjoint sensitivity analysis
abstract
Motivation: Parameter estimation methods for ordinary differential equation (ODE) models of biological processes can exploit gradients and Hessians of objective functions to achieve convergence and computational efficiency. However, the computational complexity of established methods to evaluate the Hessian scales linearly with the number of state variables and quadratically with the number of parameters. This limits their application to low-dimensional problems. Results: We introduce second order adjoint sensitivity analysis for the computation of Hessians and a hybrid optimization-integration-based approach for profile likelihood computation. Second order adjoint sensitivity analysis scales linearly with the number of parameters and state variables. The Hessians are effectively exploited by the proposed profile likelihood computation approach. We evaluate our approaches on published biological models with real measurement data. Our study reveals an improved computational efficiency and robustness of optimization compared to established approaches, when using Hessians computed with adjoint sensitivity analysis. The hybrid computation method was more than 2-fold faster than the best competitor. Thus, the proposed methods and implemented algorithms allow for the improvement of parameter estimation for medium and large scale ODE models. Availability and implementation: The algorithms for second order adjoint sensitivity analysis are implemented in the Advanced MATLAB Interface to CVODES and IDAS (AMICI, https://github.com/ICB-DCM/AMICI/). The algorithm for hybrid profile likelihood computation is implemented in the parameter estimation toolbox (PESTO, https://github.com/ICB-DCM/PESTO/). Both toolboxes are freely available under the BSD license. Supplementary information: Supplementary data are available at Bioinformatics online.
Paul Stapor, Fabian Fröhlich, Jan Hasenauer
Bioinform.3
2018 PESTO: Parameter EStimation TOolbox
abstract
Summary: PESTO is a widely applicable and highly customizable toolbox for parameter estimation in MathWorks MATLAB. It offers scalable algorithms for optimization, uncertainty and identifiability analysis, which work in a very generic manner, treating the objective function as a black box. Hence, PESTO can be used for any parameter estimation problem, for which the user can provide a deterministic objective function in MATLAB. Availability and implementation: PESTO is a MATLAB toolbox, freely available under the BSD license. The source code, along with extensive documentation and example code, can be downloaded from https://github.com/ICB-DCM/PESTO/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Paul Stapor, Daniel Weindl, Benjamin Ballnus, Sabine Hug, Carolin Loos, Anna Fiedler, Sabrina Krause, Sabrina Hross, Fabian Fröhlich, Jan Hasenauer
Bioinform.10
2017 Parameter estimation for dynamical systems with discrete events and logical operations
abstract
Motivation: Ordinary differential equation (ODE) models are frequently used to describe the dynamic behaviour of biochemical processes. Such ODE models are often extended by events to describe the effect of fast latent processes on the process dynamics. To exploit the predictive power of ODE models, their parameters have to be inferred from experimental data. For models without events, gradient based optimization schemes perform well for parameter estimation, when sensitivity equations are used for gradient computation. Yet, sensitivity equations for models with parameter- and state-dependent events and event-triggered observations are not supported by existing toolboxes. Results: In this manuscript, we describe the sensitivity equations for differential equation models with events and demonstrate how to estimate parameters from event-resolved data using event-triggered observations in parameter estimation. We consider a model for GFP expression after transfection and a model for spiking neurons and demonstrate that we can improve computational efficiency and robustness of parameter estimation by using sensitivity equations for systems with events. Moreover, we demonstrate that, by using event-outputs, it is possible to consider event-resolved data, such as time-to-event data, for parameter estimation with ODE models. By providing a user-friendly, modular implementation in the toolbox AMICI, the developed methods are made publicly available and can be integrated in other systems biology toolboxes. Availability and Implementation: We implement the methods in the open-source toolbox Advanced MATLAB Interface for CVODES and IDAS (AMICI, https://github.com/ICB-DCM/AMICI ). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Fabian Fröhlich, Fabian J. Theis, Joachim O. Rädler, Jan Hasenauer
Bioinform.4
2017 A scalable moment-closure approximation for large-scale biochemical reaction networks
abstract
MOTIVATION: Stochastic molecular processes are a leading cause of cell-to-cell variability. Their dynamics are often described by continuous-time discrete-state Markov chains and simulated using stochastic simulation algorithms. As these stochastic simulations are computationally demanding, ordinary differential equation models for the dynamics of the statistical moments have been developed. The number of state variables of these approximating models, however, grows at least quadratically with the number of biochemical species. This limits their application to small- and medium-sized processes. RESULTS: In this article, we present a scalable moment-closure approximation (sMA) for the simulation of statistical moments of large-scale stochastic processes. The sMA exploits the structure of the biochemical reaction network to reduce the covariance matrix. We prove that sMA yields approximating models whose number of state variables depends predominantly on local properties, i.e. the average node degree of the reaction network, instead of the overall network size. The resulting complexity reduction is assessed by studying a range of medium- and large-scale biochemical reaction networks. To evaluate the approximation accuracy and the improvement in computational efficiency, we study models for JAK2/STAT5 signalling and NF κ B signalling. Our method is applicable to generic biochemical reaction networks and we provide an implementation, including an SBML interface, which renders the sMA easily accessible. AVAILABILITY AND IMPLEMENTATION: The sMA is implemented in the open-source MATLAB toolbox CERENA and is available from https://github.com/CERENADevelopers/CERENA . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Atefeh Kazeroonian, Fabian J. Theis, Jan Hasenauer
Bioinform.3
2017 Robust parameter estimation for dynamical systems from outlier-corrupted data
abstract
Motivation: Dynamics of cellular processes are often studied using mechanistic mathematical models. These models possess unknown parameters which are generally estimated from experimental data assuming normally distributed measurement noise. Outlier corruption of datasets often cannot be avoided. These outliers may distort the parameter estimates, resulting in incorrect model predictions. Robust parameter estimation methods are required which provide reliable parameter estimates in the presence of outliers. Results: In this manuscript, we propose and evaluate methods for estimating the parameters of ordinary differential equation models from outlier-corrupted data. As alternatives to the normal distribution as noise distribution, we consider the Laplace, the Huber, the Cauchy and the Student's t distribution. We assess accuracy, robustness and computational efficiency of estimators using these different distribution assumptions. To this end, we consider artificial data of a conversion process, as well as published experimental data for Epo-induced JAK/STAT signaling. We study how well the methods can compensate and discover artificially introduced outliers. Our evaluation reveals that using alternative distributions improves the robustness of parameter estimates. Availability and Implementation: The MATLAB implementation of the likelihood functions using the distribution assumptions is available at Bioinformatics online. Contact: [email protected]. Supplementary information: Supplementary material are available at Bioinformatics online.
Corinna Maier, Carolin Loos, Jan Hasenauer
Bioinform.3
2017 Scalable Parameter Estimation for Genome-Scale Biochemical Reaction Networks
abstract
Mechanistic mathematical modeling of biochemical reaction networks using ordinary differential equation (ODE) models has improved our understanding of small- and medium-scale biological processes. While the same should in principle hold for large- and genome-scale processes, the computational methods for the analysis of ODE models which describe hundreds or thousands of biochemical species and reactions are missing so far. While individual simulations are feasible, the inference of the model parameters from experimental data is computationally too intensive. In this manuscript, we evaluate adjoint sensitivity analysis for parameter estimation in large scale biochemical reaction networks. We present the approach for time-discrete measurement and compare it to state-of-the-art methods used in systems and computational biology. Our comparison reveals a significantly improved computational efficiency and a superior scalability of adjoint sensitivity analysis. The computational complexity is effectively independent of the number of parameters, enabling the analysis of large- and genome-scale models. Our study of a comprehensive kinetic model of ErbB signaling shows that parameter estimation using adjoint sensitivity analysis requires a fraction of the computation time of established methods. The proposed method will facilitate mechanistic modeling of genome-scale cellular processes, as required in the age of omics.
Fabian Fröhlich, Barbara Kaltenbacher, Fabian J. Theis, Jan Hasenauer
PLoS Comput. Biol.4
2016 MEMO: multi-experiment mixture model analysis of censored data
abstract
MOTIVATION: The statistical analysis of single-cell data is a challenge in cell biological studies. Tailored statistical models and computational methods are required to resolve the subpopulation structure, i.e. to correctly identify and characterize subpopulations. These approaches also support the unraveling of sources of cell-to-cell variability. Finite mixture models have shown promise, but the available approaches are ill suited to the simultaneous consideration of data from multiple experimental conditions and to censored data. The prevalence and relevance of single-cell data and the lack of suitable computational analytics make automated methods, that are able to deal with the requirements posed by these data, necessary. RESULTS: We present MEMO, a flexible mixture modeling framework that enables the simultaneous, automated analysis of censored and uncensored data acquired under multiple experimental conditions. MEMO is based on maximum-likelihood inference and allows for testing competing hypotheses. MEMO can be applied to a variety of different single-cell data types. We demonstrate the advantages of MEMO by analyzing right and interval censored single-cell microscopy data. Our results show that an examination of censoring and the simultaneous consideration of different experimental conditions are necessary to reveal biologically meaningful subpopulation structures. MEMO allows for a stringent analysis of single-cell data and enables researchers to avoid misinterpretation of censored data. Therefore, MEMO is a valuable asset for all fields that infer the characteristics of populations by looking at single individuals such as cell biology and medicine. AVAILABILITY AND IMPLEMENTATION: MEMO is implemented in MATLAB and freely available via github (https://github.com/MEMO-toolbox/MEMO). CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Eva-Maria Geissen, Jan Hasenauer, Stephanie Heinrich, Silke Hauf, Fabian J. Theis, Nicole Radde
Bioinform.2
2016 Analysis of CFSE time-series data using division-, age- and label-structured population models
abstract
MOTIVATION: In vitro and in vivo cell proliferation is often studied using the dye carboxyfluorescein succinimidyl ester (CFSE). The CFSE time-series data provide information about the proliferation history of populations of cells. While the experimental procedures are well established and widely used, the analysis of CFSE time-series data is still challenging. Many available analysis tools do not account for cell age and employ optimization methods that are inefficient (or even unreliable). RESULTS: We present a new model-based analysis method for CFSE time-series data. This method uses a flexible description of proliferating cell populations, namely, a division-, age- and label-structured population model. Efficient maximum likelihood and Bayesian estimation algorithms are introduced to infer the model parameters and their uncertainties. These methods exploit the forward sensitivity equations of the underlying partial differential equation model for efficient and accurate gradient calculation, thereby improving computational efficiency and reliability compared with alternative approaches and accelerating uncertainty analysis. The performance of the method is assessed by studying a dataset for immune cell proliferation. This revealed the importance of different factors on the proliferation rates of individual cells. Among others, the predominate effect of cell age on the division rate is found, which was not revealed by available computational methods. AVAILABILITY AND IMPLEMENTATION: The MATLAB source code implementing the models and algorithms is available from http://janhasenauer.github.io/ShAPE-DALSP/Contact: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sabrina Hross, Jan Hasenauer
Bioinform.2
2016 Inference for Stochastic Chemical Kinetics Using Moment Equations and System Size Expansion
abstract
Quantitative mechanistic models are valuable tools for disentangling biochemical pathways and for achieving a comprehensive understanding of biological systems. However, to be quantitative the parameters of these models have to be estimated from experimental data. In the presence of significant stochastic fluctuations this is a challenging task as stochastic simulations are usually too time-consuming and a macroscopic description using reaction rate equations (RREs) is no longer accurate. In this manuscript, we therefore consider moment-closure approximation (MA) and the system size expansion (SSE), which approximate the statistical moments of stochastic processes and tend to be more precise than macroscopic descriptions. We introduce gradient-based parameter optimization methods and uncertainty analysis methods for MA and SSE. Efficiency and reliability of the methods are assessed using simulation examples as well as by an application to data for Epo-induced JAK/STAT signaling. The application revealed that even if merely population-average data are available, MA and SSE improve parameter identifiability in comparison to RRE. Furthermore, the simulation examples revealed that the resulting estimates are more reliable for an intermediate volume regime. In this regime the estimation error is reduced and we propose methods to determine the regime boundaries. These results illustrate that inference using MA and SSE is feasible and possesses a high sensitivity.
Fabian Fröhlich, Philipp Thomas, Atefeh Kazeroonian, Fabian J. Theis, Ramon Grima, Jan Hasenauer
PLoS Comput. Biol.6
2015 Data2Dynamics: a modeling environment tailored to parameter estimation in dynamical systems
abstract
UNLABELLED: Modeling of dynamical systems using ordinary differential equations is a popular approach in the field of systems biology. Two of the most critical steps in this approach are to construct dynamical models of biochemical reaction networks for large datasets and complex experimental conditions and to perform efficient and reliable parameter estimation for model fitting. We present a modeling environment for MATLAB that pioneers these challenges. The numerically expensive parts of the calculations such as the solving of the differential equations and of the associated sensitivity system are parallelized and automatically compiled into efficient C code. A variety of parameter estimation algorithms as well as frequentist and Bayesian methods for uncertainty analysis have been implemented and used on a range of applications that lead to publications. AVAILABILITY AND IMPLEMENTATION: The Data2Dynamics modeling environment is MATLAB based, open source and freely available at http://www.data2dynamics.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andreas Raue, Bernhard Steiert, Max Schelker, Clemens Kreutz, Tim Maiwald, Helge Hass, Joep Vanlier, Christian Tönsing, Lorenz Adlung, Raphael Engesser, Wolfgang Mader 0002, Tim Heinemann, Jan Hasenauer, Marcel Schilling, Thomas Höfer, Edda Klipp, Fabian J. Theis, Ursula Klingmüller, Birgit Schoeberl, Jens Timmer
Bioinform.13
2014 ODE Constrained Mixture Modelling: A Method for Unraveling Subpopulation Structures and Dynamics
abstract
Functional cell-to-cell variability is ubiquitous in multicellular organisms as well as bacterial populations. Even genetically identical cells of the same cell type can respond differently to identical stimuli. Methods have been developed to analyse heterogeneous populations, e.g., mixture models and stochastic population models. The available methods are, however, either incapable of simultaneously analysing different experimental conditions or are computationally demanding and difficult to apply. Furthermore, they do not account for biological information available in the literature. To overcome disadvantages of existing methods, we combine mixture models and ordinary differential equation (ODE) models. The ODE models provide a mechanistic description of the underlying processes while mixture models provide an easy way to capture variability. In a simulation study, we show that the class of ODE constrained mixture models can unravel the subpopulation structure and determine the sources of cell-to-cell variability. In addition, the method provides reliable estimates for kinetic rates and subpopulation characteristics. We use ODE constrained mixture modelling to study NGF-induced Erk1/2 phosphorylation in primary sensory neurones, a process relevant in inflammatory and neuropathic pain. We propose a mechanistic pathway model for this process and reconstructed static and dynamical subpopulation characteristics across experimental conditions. We validate the model predictions experimentally, which verifies the capabilities of ODE constrained mixture models. These results illustrate that ODE constrained mixture models can reveal novel mechanistic insights and possess a high sensitivity.
Jan Hasenauer, Christine Hasenauer, Tim Hucho, Fabian J. Theis
PLoS Comput. Biol.1
2013 Visualizing edge-edge relations in graphs
abstract
Graphs are used to model relations between sets of objects. Objects are represented by vertices and relations by edges of the graph. Besides vertex-vertex relations, in some application domains also relations between edges exist. Our new visualization approach supports the investigation of both relation types in one diagram. Edge-edge relations are visualized as curves that are directly integrated into the node-link diagram that represents the object-relation structure. In contrast, vertex-vertex relations are illustrated distinguishably from edge-edge relations using straight links as representations. While the shape of links is used to differentiate between the relation types, the weights of the edge-edge relations are mapped to the width and color of the curves. To facilitate an extensive analysis of interrelations, our approach incorporates several interaction techniques that can be used for filtering and highlighting. The usability of our visualization is demonstrated with two case studies in the application domains of bioinformatics and financial services.
Corinna Vehlow, Jan Hasenauer, Fabian J. Theis, Daniel Weiskopf
PacificVis2
2013 Modeling of 2D diffusion processes based on microscopy data: parameter estimation and practical identifiability analysis
abstract
BACKGROUND: Diffusion is a key component of many biological processes such as chemotaxis, developmental differentiation and tissue morphogenesis. Since recently, the spatial gradients caused by diffusion can be assessed in-vitro and in-vivo using microscopy based imaging techniques. The resulting time-series of two dimensional, high-resolutions images in combination with mechanistic models enable the quantitative analysis of the underlying mechanisms. However, such a model-based analysis is still challenging due to measurement noise and sparse observations, which result in uncertainties of the model parameters. METHODS: We introduce a likelihood function for image-based measurements with log-normal distributed noise. Based upon this likelihood function we formulate the maximum likelihood estimation problem, which is solved using PDE-constrained optimization methods. To assess the uncertainty and practical identifiability of the parameters we introduce profile likelihoods for diffusion processes. RESULTS AND CONCLUSION: As proof of concept, we model certain aspects of the guidance of dendritic cells towards lymphatic vessels, an example for haptotaxis. Using a realistic set of artificial measurement data, we estimate the five kinetic parameters of this model and compute profile likelihoods. Our novel approach for the estimation of model parameters from image data as well as the proposed identifiability analysis approach is widely applicable to diffusion processes. The profile likelihood based method provides more rigorous uncertainty bounds in contrast to local approximation methods.
Sabrina Hock, Jan Hasenauer, Fabian J. Theis
BMC Bioinform.2
2013 iVUN: interactive Visualization of Uncertain biochemical reaction Networks
abstract
BACKGROUND: Mathematical models are nowadays widely used to describe biochemical reaction networks. One of the main reasons for this is that models facilitate the integration of a multitude of different data and data types using parameter estimation. Thereby, models allow for a holistic understanding of biological processes. However, due to measurement noise and the limited amount of data, uncertainties in the model parameters should be considered when conclusions are drawn from estimated model attributes, such as reaction fluxes or transient dynamics of biological species. METHODS AND RESULTS: We developed the visual analytics system iVUN that supports uncertainty-aware analysis of static and dynamic attributes of biochemical reaction networks modeled by ordinary differential equations. The multivariate graph of the network is visualized as a node-link diagram, and statistics of the attributes are mapped to the color of nodes and links of the graph. In addition, the graph view is linked with several views, such as line plots, scatter plots, and correlation matrices, to support locating uncertainties and the analysis of their time dependencies. As demonstration, we use iVUN to quantitatively analyze the dynamics of a model for Epo-induced JAK2/STAT5 signaling. CONCLUSION: Our case study showed that iVUN can be used to perform an in-depth study of biochemical reaction networks, including attribute uncertainties, correlations between these attributes and their uncertainties as well as the attribute dynamics. In particular, the linking of different visualization options turned out to be highly beneficial for the complex analysis tasks that come with the biological systems as presented here.
Corinna Vehlow, Jan Hasenauer, Andrei Kramer, Andreas Raue, Sabine Hug, Jens Timmer, Nicole Radde, Fabian J. Theis, Daniel Weiskopf
BMC Bioinform.2
2011 Identification of models of heterogeneous cell populations from population snapshot data
abstract
BACKGROUND: Most of the modeling performed in the area of systems biology aims at achieving a quantitative description of the intracellular pathways within a "typical cell". However, in many biologically important situations even clonal cell populations can show a heterogeneous response. These situations require study of cell-to-cell variability and the development of models for heterogeneous cell populations. RESULTS: In this paper we consider cell populations in which the dynamics of every single cell is captured by a parameter dependent differential equation. Differences among cells are modeled by differences in parameters which are subject to a probability density. A novel Bayesian approach is presented to infer this probability density from population snapshot data, such as flow cytometric analysis, which do not provide single cell time series data. The presented approach can deal with sparse and noisy measurement data. Furthermore, it is appealing from an application point of view as in contrast to other methods the uncertainty of the resulting parameter distribution can directly be assessed. CONCLUSIONS: The proposed method is evaluated using artificial experimental data from a model of the tumor necrosis factor signaling network. We demonstrate that the methods are computationally efficient and yield good estimation result even for sparse data sets.
Jan Hasenauer, Steffen Waldherr, Malgorzata Doszczak, Nicole Radde, Peter Scheurich, Frank Allgöwer
BMC Bioinform.1