VLDB 2026 Research / reviewers in the wild / expert
Clemens Kreutz
dblp:122/8986
· DBLP profile ↗
19ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-8796-5766ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mmContext: an open framework for multimodal contrastive learning of omics and text dataabstractSUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493. Jonatan Menger, Sonia Maria Krissmer, Clemens Kreutz, Harald Binder, Maren Hackenberg |
Bioinform. | 3 |
| 2024 | RTF: an R package for modelling time course dataabstractSUMMARY: The retarded transient function (RTF) approach serves as a complementary method to ordinary differential equations (ODEs) for modelling dynamics typically observed in cellular signalling processes. We introduce an R package that implements the RTF approach, originally implemented within the MATLAB-based Data2Dynamics modelling framework. This package facilitates the modelling of time and dose dependencies, and it includes the possibility of model reduction to minimize overfitting. It can be applied to experimental data or trajectories of ODE models to characterize their dynamics. Additionally, it can generate a low-dimensional representation based on the fitted RTF parameters of a set of time-resolved data, aiding in the identification of key targets of experimental perturbations. AVAILABILITY AND IMPLEMENTATION: The R package RTF is available at https://github.com/kreutz-lab/RTF. Eva Brombacher, Clemens Kreutz |
Bioinform. | 2 |
| 2024 | Dynamic modelling of signalling pathways when ordinary differential equations are not feasibleabstractMOTIVATION: Mathematical modelling plays a crucial role in understanding inter- and intracellular signalling processes. Currently, ordinary differential equations (ODEs) are the predominant approach in systems biology for modelling such pathways. While ODE models offer mechanistic interpretability, they also suffer from limitations, including the need to consider all relevant compounds, resulting in large models difficult to handle numerically and requiring extensive data. RESULTS: In previous work, we introduced the retarded transient function (RTF) as an alternative method for modelling temporal responses of signalling pathways. Here, we extend the RTF approach to integrate concentration or dose-dependencies into the modelling of dynamics. With this advancement, RTF modelling now fully encompasses the application range of ODE models, which comprises predictions in both time and concentration domains. Moreover, characterizing dose-dependencies provides an intuitive way to investigate and characterize signalling differences between biological conditions or cell types based on their response to stimulating inputs. To demonstrate the applicability of our extended approach, we employ data from time- and dose-dependent inflammasome activation in bone marrow-derived macrophages treated with nigericin sodium salt. Our results show the effectiveness of the extended RTF approach as a generic framework for modelling dose-dependent kinetics in cellular signalling. The approach results in intuitively interpretable parameters that describe signal dynamics and enables predictive modelling of time- and dose-dependencies even if only individual cellular components are quantified. AVAILABILITY AND IMPLEMENTATION: The presented approach is available within the MATLAB-based Data2Dynamics modelling toolbox at https://github.com/Data2Dynamics and https://zenodo.org/records/14008247 and as R code at https://github.com/kreutz-lab/RTF. Timo Rachel, Eva Brombacher, Svenja Wöhrle, Olaf Groß, Clemens Kreutz |
Bioinform. | 5 |
| 2023 | Likelihood-ratio test statistic for the finite-sample case in nonlinear ordinary differential equation modelsabstractLikelihood ratios are frequently utilized as basis for statistical tests, for model selection criteria and for assessing parameter and prediction uncertainties, e.g. using the profile likelihood. However, translating these likelihood ratios into p-values or confidence intervals requires the exact form of the test statistic's distribution. The lack of knowledge about this distribution for nonlinear ordinary differential equation (ODE) models requires an approximation which assumes the so-called asymptotic setting, i.e. a sufficiently large amount of data. Since the amount of data from quantitative molecular biology is typically limited in applications, this finite-sample case regularly occurs for mechanistic models of dynamical systems, e.g. biochemical reaction networks or infectious disease models. Thus, it is unclear whether the standard approach of using statistical thresholds derived for the asymptotic large-sample setting in realistic applications results in valid conclusions. In this study, empirical likelihood ratios for parameters from 19 published nonlinear ODE benchmark models are investigated using a resampling approach for the original data designs. Their distributions are compared to the asymptotic approximation and statistical thresholds are checked for conservativeness. It turns out, that corrections of the likelihood ratios in such finite-sample applications are required in order to avoid anti-conservative results. Christian Tönsing, Bernhard Steiert, Jens Timmer, Clemens Kreutz |
PLoS Comput. Biol. | 4 |
| 2021 | PEtab - Interoperable specification of parameter estimation problems in systems biologyabstractReproducibility and reusability of the results of data-based modeling studies are essential. Yet, there has been-so far-no broadly supported format for the specification of parameter estimation problems in systems biology. Here, we introduce PEtab, a format which facilitates the specification of parameter estimation problems using Systems Biology Markup Language (SBML) models and a set of tab-separated value files describing the observation model and experimental data as well as parameters to be estimated. We already implemented PEtab support into eight well-established model simulation and parameter estimation toolboxes with hundreds of users in total. We provide a Python library for validation and modification of a PEtab problem and currently 20 example parameter estimation problems based on recent studies. Leonard Schmiester, Yannik Schälte, Frank T. Bergmann, Tacio Camba, Erika Dudkin, Janine Egert, Fabian Fröhlich, Lara Fuhrmann, Adrian L. Hauber, Svenja Kemmer, Polina A. Lakrisenko, Carolin Loos, Simon Merkt, Wolfgang Müller 0001, Dilan Pathirana, Elba Raimúndez-Álvarez, Lukas Refisch, Marcus Rosenblatt, Paul Stapor, Philipp Städter, Dantong Wang, Franz-Georg Wieland, Julio R. Banga, Jens Timmer, Alejandro Fernández Villaverde, Sven Sahle, Clemens Kreutz, Jan Hasenauer, Daniel Weindl |
PLoS Comput. Biol. | 27 |
| 2020 | A blind and independent benchmark study for detecting differentially methylated regions in plantsabstractMOTIVATION: Bisulfite sequencing (BS-seq) is a state-of-the-art technique for investigating methylation of the DNA to gain insights into the epigenetic regulation. Several algorithms have been published for identification of differentially methylated regions (DMRs). However, the performances of the individual methods remain unclear and it is difficult to optimally select an algorithm in application settings. RESULTS: We analyzed BS-seq data from four plants covering three taxonomic groups. We first characterized the data using multiple summary statistics describing methylation levels, coverage, noise, as well as frequencies, magnitudes and lengths of methylated regions. Then, simulated datasets with most similar characteristics to real experimental data were created. Seven different algorithms (metilene, methylKit, MOABS, DMRcate, Defiant, BSmooth, MethylSig) for DMR identification were applied and their performances were assessed. A blind and independent study design was chosen to reduce bias and to derive practical method selection guidelines. Overall, metilene had superior performance in most settings. Data attributes, such as coverage and spread of the DMR lengths, were found to be useful for selecting the best method for DMR detection. A decision tree to select the optimal approach based on these data attributes is provided. The presented procedure might serve as a general strategy for deriving algorithm selection rules tailored to demands in specific application settings. AVAILABILITY AND IMPLEMENTATION: Scripts that were used for the analyses and that can be used for prediction of the optimal algorithm are provided at https://github.com/kreutz-lab/DMR-DecisionTree. Simulated and experimental data are available at https://doi.org/10.6084/m9.figshare.11619045. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Clemens Kreutz, Nilay S. Can, Ralf Schulze Bruening, Rabea Meyberg, Zsuzsanna Mérai, Noé Fernández-Pozo, Stefan A. Rensing |
Bioinform. | 1 |
| 2020 | A Blind and Independent Benchmark Study for Detecting Differentially Methylated Regions in PlantsabstractBioinformatics (2020) doi: 10.1093/bioinformatics/btaa191 The title of this article has been corrected to: A Blind and Independent Benchmark Study for Detecting Differentially Methylated Regions in Plants. Clemens Kreutz, Nilay S. Can, Ralf Schulze Bruening, Rabea Meyberg, Zsuzsanna Mérai, Noé Fernández-Pozo, Stefan A. Rensing |
Bioinform. | 1 |
| 2019 | Benchmark problems for dynamic modeling of intracellular processesabstractMOTIVATION: Dynamic models are used in systems biology to study and understand cellular processes like gene regulation or signal transduction. Frequently, ordinary differential equation (ODE) models are used to model the time and dose dependency of the abundances of molecular compounds as well as interactions and translocations. A multitude of computational approaches, e.g. for parameter estimation or uncertainty analysis have been developed within recent years. However, many of these approaches lack proper testing in application settings because a comprehensive set of benchmark problems is yet missing. RESULTS: We present a collection of 20 benchmark problems in order to evaluate new and existing methodologies, where an ODE model with corresponding experimental data is referred to as problem. In addition to the equations of the dynamical system, the benchmark collection provides observation functions as well as assumptions about measurement noise distributions and parameters. The presented benchmark models comprise problems of different size, complexity and numerical demands. Important characteristics of the models and methodological requirements are summarized, estimated parameters are provided, and some example studies were performed for illustrating the capabilities of the presented benchmark collection. AVAILABILITY AND IMPLEMENTATION: The models are provided in several standardized formats, including an easy-to-use human readable form and machine-readable SBML files. The data is provided as Excel sheets. All files are available at https://github.com/Benchmarking-Initiative/Benchmark-Models, including step-by-step explanations and MATLAB code to process and simulate the models. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Helge Hass, Carolin Loos, Elba Raimúndez-Álvarez, Jens Timmer, Jan Hasenauer, Clemens Kreutz |
Bioinform. | 6 |
| 2018 | An easy and efficient approach for testing identifiabilityabstractMotivation: The feasibility of uniquely estimating parameters of dynamical systems from observations is a widely discussed aspect of mathematical modelling. Several approaches have been published for analyzing this so-called identifiability of model parameters. However, they are typically computationally demanding, difficult to perform and/or not applicable in many application settings. Results: Here, an approach is presented which enables quickly testing of parameter identifiability. Numerical optimization with a penalty in radial direction enforcing displacement of the parameters is used to check whether estimated parameters are unique, or whether the parameters can be altered without loss of agreement with the data indicating non-identifiability. This Identifiability-Test by Radial Penalization (ITRP) can be employed for every model where optimization-based parameter estimation like least-squares or maximum likelihood is feasible and is therefore applicable for all typical systems biology models. The approach is illustrated and tested using 11 ordinary differential equation (ODE) models. Availability and implementation: The presented approach can be implemented without great efforts in any modelling framework. It is available within the free Matlab-based modelling toolbox Data2Dynamics. Source code is available at https://github.com/Data2Dynamics. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Clemens Kreutz |
Bioinform. | 1 |
| 2016 | Fast integration-based prediction bands for ordinary differential equation modelsabstractMOTIVATION: To gain a deeper understanding of biological processes and their relevance in disease, mathematical models are built upon experimental data. Uncertainty in the data leads to uncertainties of the model's parameters and in turn to uncertainties of predictions. Mechanistic dynamic models of biochemical networks are frequently based on nonlinear differential equation systems and feature a large number of parameters, sparse observations of the model components and lack of information in the available data. Due to the curse of dimensionality, classical and sampling approaches propagating parameter uncertainties to predictions are hardly feasible and insufficient. However, for experimental design and to discriminate between competing models, prediction and confidence bands are essential. To circumvent the hurdles of the former methods, an approach to calculate a profile likelihood on arbitrary observations for a specific time point has been introduced, which provides accurate confidence and prediction intervals for nonlinear models and is computationally feasible for high-dimensional models. RESULTS: In this article, reliable and smooth point-wise prediction and confidence bands to assess the model's uncertainty on the whole time-course are achieved via explicit integration with elaborate correction mechanisms. The corresponding system of ordinary differential equations is derived and tested on three established models for cellular signalling. An efficiency analysis is performed to illustrate the computational benefit compared with repeated profile likelihood calculations at multiple time points. AVAILABILITY AND IMPLEMENTATION: The integration framework and the examples used in this article are provided with the software package Data2Dynamics, which is based on MATLAB and freely available at http://www.data2dynamics.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Helge Hass, Clemens Kreutz, Jens Timmer, Daniel Kaschek |
Bioinform. | 2 |
| 2016 | L1 regularization facilitates detection of cell type-specific parameters in dynamical systemsabstractMOTIVATION: A major goal of drug development is to selectively target certain cell types. Cellular decisions influenced by drugs are often dependent on the dynamic processing of information. Selective responses can be achieved by differences between the involved cell types at levels of receptor, signaling, gene regulation or further downstream. Therefore, a systematic approach to detect and quantify cell type-specific parameters in dynamical systems becomes necessary. RESULTS: Here, we demonstrate that a combination of nonlinear modeling with L1 regularization is capable of detecting cell type-specific parameters. To adapt the least-squares numerical optimization routine to L1 regularization, sub-gradient strategies as well as truncation of proposed optimization steps were implemented. Likelihood-ratio tests were used to determine the optimal regularization strength resulting in a sparse solution in terms of a minimal number of cell type-specific parameters that is in agreement with the data. By applying our implementation to a realistic dynamical benchmark model of the DREAM6 challenge we were able to recover parameter differences with an accuracy of 78%. Within the subset of detected differences, 91% were in agreement with their true value. Furthermore, we found that the results could be improved using the profile likelihood. In conclusion, the approach constitutes a general method to infer an overarching model with a minimum number of individual parameters for the particular models. AVAILABILITY AND IMPLEMENTATION: A MATLAB implementation is provided within the freely available, open-source modeling environment Data2Dynamics. Source code for all examples is provided online at http://www.data2dynamics.org/ CONTACT: [email protected]. Bernhard Steiert, Jens Timmer, Clemens Kreutz |
Bioinform. | 3 |
| 2016 | Identification of Cell Type-Specific Differences in Erythropoietin Receptor Signaling in Primary Erythroid and Lung Cancer CellsabstractLung cancer, with its most prevalent form non-small-cell lung carcinoma (NSCLC), is one of the leading causes of cancer-related deaths worldwide, and is commonly treated with chemotherapeutic drugs such as cisplatin. Lung cancer patients frequently suffer from chemotherapy-induced anemia, which can be treated with erythropoietin (EPO). However, studies have indicated that EPO not only promotes erythropoiesis in hematopoietic cells, but may also enhance survival of NSCLC cells. Here, we verified that the NSCLC cell line H838 expresses functional erythropoietin receptors (EPOR) and that treatment with EPO reduces cisplatin-induced apoptosis. To pinpoint differences in EPO-induced survival signaling in erythroid progenitor cells (CFU-E, colony forming unit-erythroid) and H838 cells, we combined mathematical modeling with a method for feature selection, the L1 regularization. Utilizing an example model and simulated data, we demonstrated that this approach enables the accurate identification and quantification of cell type-specific parameters. We applied our strategy to quantitative time-resolved data of EPO-induced JAK/STAT signaling generated by quantitative immunoblotting, mass spectrometry and quantitative real-time PCR (qRT-PCR) in CFU-E and H838 cells as well as H838 cells overexpressing human EPOR (H838-HA-hEPOR). The established parsimonious mathematical model was able to simultaneously describe the data sets of CFU-E, H838 and H838-HA-hEPOR cells. Seven cell type-specific parameters were identified that included for example parameters for nuclear translocation of STAT5 and target gene induction. Cell type-specific differences in target gene induction were experimentally validated by qRT-PCR experiments. The systematic identification of pathway differences and sensitivities of EPOR signaling in CFU-E and H838 cells revealed potential targets for intervention to selectively inhibit EPO-induced signaling in the tumor cells but leave the responses in erythroid progenitor cells unaffected. Thus, the proposed modeling strategy can be employed as a general procedure to identify cell type-specific parameters and to recommend treatment strategies for the selective targeting of specific cell types. Ruth Merkle, Bernhard Steiert, Florian Salopiata, Sofia Depner, Andreas Raue, Nao Iwamoto, Max Schelker, Helge Hass, Marvin Wäsch, Martin E. Böhm, Oliver Mücke, Daniel B. Lipka, Christoph Plass, Wolf D. Lehmann, Clemens Kreutz, Jens Timmer, Marcel Schilling, Ursula Klingmüller |
PLoS Comput. Biol. | 15 |
| 2015 | Data2Dynamics: a modeling environment tailored to parameter estimation in dynamical systemsabstractUNLABELLED: Modeling of dynamical systems using ordinary differential equations is a popular approach in the field of systems biology. Two of the most critical steps in this approach are to construct dynamical models of biochemical reaction networks for large datasets and complex experimental conditions and to perform efficient and reliable parameter estimation for model fitting. We present a modeling environment for MATLAB that pioneers these challenges. The numerically expensive parts of the calculations such as the solving of the differential equations and of the associated sensitivity system are parallelized and automatically compiled into efficient C code. A variety of parameter estimation algorithms as well as frequentist and Bayesian methods for uncertainty analysis have been implemented and used on a range of applications that lead to publications. AVAILABILITY AND IMPLEMENTATION: The Data2Dynamics modeling environment is MATLAB based, open source and freely available at http://www.data2dynamics.org. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andreas Raue, Bernhard Steiert, Max Schelker, Clemens Kreutz, Tim Maiwald, Helge Hass, Joep Vanlier, Christian Tönsing, Lorenz Adlung, Raphael Engesser, Wolfgang Mader 0002, Tim Heinemann, Jan Hasenauer, Marcel Schilling, Thomas Höfer, Edda Klipp, Fabian J. Theis, Ursula Klingmüller, Birgit Schoeberl, Jens Timmer |
Bioinform. | 4 |
| 2015 | Summary of the DREAM8 Parameter Estimation Challenge: Toward Parameter Identification for Whole-Cell ModelsabstractWhole-cell models that explicitly represent all cellular components at the molecular level have the potential to predict phenotype from genotype. However, even for simple bacteria, whole-cell models will contain thousands of parameters, many of which are poorly characterized or unknown. New algorithms are needed to estimate these parameters and enable researchers to build increasingly comprehensive models. We organized the Dialogue for Reverse Engineering Assessments and Methods (DREAM) 8 Whole-Cell Parameter Estimation Challenge to develop new parameter estimation algorithms for whole-cell models. We asked participants to identify a subset of parameters of a whole-cell model given the model's structure and in silico "experimental" data. Here we describe the challenge, the best performing methods, and new insights into the identifiability of whole-cell models. We also describe several valuable lessons we learned toward improving future challenges. Going forward, we believe that collaborative efforts supported by inexpensive cloud computing have the potential to solve whole-cell model parameter estimation. Jonathan R. Karr, Alex H. Williams, Jeremy Zucker, Andreas Raue, Bernhard Steiert, Jens Timmer, Clemens Kreutz, Simon Wilkinson, Brandon A. Allgood, Brian M. Bot, Bruce R. Hoff, Michael R. Kellen, Markus W. Covert, Gustavo Stolovitzky, Pablo Meyer 0001 |
PLoS Comput. Biol. | 7 |
| 2012 | TSSi - an R package for transcription start site identification from 5′ mRNA tag dataabstractUNLABELLED: High-throughput sequencing has become an essential experimental approach for the investigation of transcriptional mechanisms. For some applications like ChIP-seq, several approaches for the prediction of peak locations exist. However, these methods are not designed for the identification of transcription start sites (TSSs) because such datasets contain qualitatively different noise. In this application note, the R package TSSi is presented which provides a heuristic framework for the identification of TSSs based on 5' mRNA tag data. Probabilistic assumptions for the distribution of the data, i.e. for the observed positions of the mapped reads, as well as for systematic errors, i.e. for reads which map closely but not exactly to a real TSS, are made and can be adapted by the user. The framework also comprises a regularization procedure which can be applied as a preprocessing step to decrease the noise and thereby reduce the number of false predictions. AVAILABILITY: The R package TSSi is available from the Bioconductor web site: www.bioconductor.org/packages/release/bioc/html/TSSi.html. Clemens Kreutz, J. S. Gehring, Daniel Lang 0001, Ralf Reski, Jens Timmer, Stefan A. Rensing |
Bioinform. | 1 |
| 2012 | Comprehensive estimation of input signals and dynamics in biochemical reaction networksabstractMOTIVATION: Cellular information processing can be described mathematically using differential equations. Often, external stimulation of cells by compounds such as drugs or hormones leading to activation has to be considered. Mathematically, the stimulus is represented by a time-dependent input function. Parameters such as rate constants of the molecular interactions are often unknown and need to be estimated from experimental data, e.g. by maximum likelihood estimation. For this purpose, the input function has to be defined for all times of the integration interval. This is usually achieved by approximating the input by interpolation or smoothing of the measured data. This procedure is suboptimal since the input uncertainties are not considered in the estimation process which often leads to overoptimistic confidence intervals of the inferred parameters and the model dynamics. RESULTS: This article presents a new approach which includes the input estimation into the estimation process of the dynamical model parameters by minimizing an objective function containing all parameters simultaneously. We applied this comprehensive approach to an illustrative model with simulated data and compared it to alternative methods. Statistical analyses revealed that our method improves the prediction of the model dynamics and the confidence intervals leading to a proper coverage of the confidence intervals of the dynamic parameters. The method was applied to the JAK-STAT signaling pathway. AVAILABILITY: MATLAB code is available on the authors' website http://www.fdmold.uni-freiburg.de/~schelker/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Additional information is available at Bioinformatics Online. Max Schelker, Andreas Raue, Jens Timmer, Clemens Kreutz |
Bioinform. | 4 |
| 2009 | Structural and practical identifiability analysis of partially observed dynamical models by exploiting the profile likelihoodabstractMOTIVATION: Mathematical description of biological reaction networks by differential equations leads to large models whose parameters are calibrated in order to optimally explain experimental data. Often only parts of the model can be observed directly. Given a model that sufficiently describes the measured data, it is important to infer how well model parameters are determined by the amount and quality of experimental data. This knowledge is essential for further investigation of model predictions. For this reason a major topic in modeling is identifiability analysis. RESULTS: We suggest an approach that exploits the profile likelihood. It enables to detect structural non-identifiabilities, which manifest in functionally related model parameters. Furthermore, practical non-identifiabilities are detected, that might arise due to limited amount and quality of experimental data. Last but not least confidence intervals can be derived. The results are easy to interpret and can be used for experimental planning and for model reduction. AVAILABILITY: An implementation is freely available for MATLAB and the PottersWheel modeling toolbox at http://web.me.com/andreas.raue/profile/software.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Andreas Raue, Clemens Kreutz, Thomas Maiwald, Julie Bachmann, Marcel Schilling, Ursula Klingmüller, Jens Timmer |
Bioinform. | 2 |
| 2007 | Data-based identifiability analysis of non-linear dynamical modelsabstractAbstract Motivation: Mathematical modelling of biological systems is becoming a standard approach to investigate complex dynamic, non-linear interaction mechanisms in cellular processes. However, models may comprise non-identifiable parameters which cannot be unambiguously determined. Non-identifiability manifests itself in functionally related parameters, which are difficult to detect. Results: We present the method of mean optimal transformations, a non-parametric bootstrap-based algorithm for identifiability testing, capable of identifying linear and non-linear relations of arbitrarily many parameters, regardless of model size or complexity. This is performed with use of optimal transformations, estimated using the alternating conditional expectation algorithm (ACE). An initial guess or prior knowledge concerning the underlying relation of the parameters is not required. Independent, and hence identifiable parameters are determined as well. The quality of data at disposal is included in our approach, i.e. the non-linear model is fitted to data and estimated parameter values are investigated with respect to functional relations. We exemplify our approach on a realistic dynamical model and demonstrate that the variability of estimated parameter values decreases from 81 to 1% after detection and fixation of structural non-identifiabilities. Availability: Our algorithm is written in Matlab and R. It is available from the authors on request. An implementation of ACE, written in Matlab as well as in C, is available online at www.stefanhengl.de Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Stefan Hengl, Clemens Kreutz, Jens Timmer, Thomas Maiwald |
Bioinform. | 2 |
| 2007 | An error model for protein quantificationabstractMOTIVATION: Quantitative experimental data is the critical bottleneck in the modeling of dynamic cellular processes in systems biology. Here, we present statistical approaches improving reproducibility of protein quantification by immunoprecipitation and immunoblotting. RESULTS: Based on a large data set with more than 3600 data points, we unravel that the main sources of biological variability and experimental noise are multiplicative and log-normally distributed. Therefore, we suggest a log-transformation of the data to obtain additive normally distributed noise. After this transformation, common statistical procedures can be applied to analyze the data. An error model is introduced to account for technical as well as biological variability. Elimination of these systematic errors decrease variability of measurements and allow for a more precise estimation of underlying dynamics of protein concentrations in cellular signaling. The proposed error model is relevant for simulation studies, parameter estimation and model selection, basic tools of systems biology. AVAILABILITY: Matlab and R code is available from the authors on request. The data can be downloaded from our website www.fdm.uni-freiburg.de/~ckreutz/data. Clemens Kreutz, M. M. Bartolome Rodriguez, Thomas Maiwald, Maxmilian Seidl, H. E. Blum, L. Mohr, Jens Timmer |
Bioinform. | 1 |