Olaf Wolkenhauer

dblp:98/3100 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-6105-2937ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multivariate Functional Linear Discriminant Analysis for Partially-Observed Time Series (Abstract Reprint)
abstract
The more extensive access to time-series data, especially for biomedical purposes, raises new methodological challenges, particularly regarding missing values. Functional linear discriminant analysis (FLDA) extends Linear Discriminant Analysis (LDA)-mediated multiclass classification and dimension reduction to data in the form of fragmented observations of a univariate function. For large multivariate and partially-observed data, there are two challenges: (i) statistical dependencies between different components of a multivariate function and (ii) heterogeneous sampling times with missing features. We here develop a multivariate version of FLDA, called MUDRA, to tackle these challenges and describe a computationally efficient expectation/conditional-maximisation (ECM) algorithm to infer its parameters without any tensor inversions. We assess its predictive power on the “Articulary Words” dataset and show its improvement over the state-of-the-art, especially in the case of missing data. This advancement in dimension reduction of multivariate functional data holds promise for enhancing classification accuracy in scenarios like partially observed short multivariate time series analysis.
Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer
AAAI5
2026 Anomaly detection via mean shift density enhancement
Pritam Kar, Rahul Bordoloi, Olaf Wolkenhauer, Saptarshi Bej
Data Min. Knowl. Discov.3
2026 Convex space learning for tabular synthetic data generation
Manjunath Mahendra, Chaithra Umesh, Kristian Schultz, Olaf Wolkenhauer, Saptarshi Bej
Neurocomputing4
2026 Correction: Multivariate functional linear discriminant analysis for partially‑observed time series
abstract
In the original publication of this article, the legend for Fig. 3 was inadvertently omitted in the published version.Specifically, the right-hand plot, depicting the performance of MUDRA on a synthetic dataset in terms of F1 score relative to the baseline ROCKET, was missing the legend identifying the respective models.For completeness and transparency, the incorrect and correct versions of Fig. 3 are presented with this correction article.The original article has been corrected.The original article can be found online at h t t p s : / / d o i . o r g / 1 0 . 1 0 0 7 / s 1 0 9 9 4 -0 2 5 -0 6 7 4 1 -0 .
Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer
Mach. Learn.5
2026 FUSE: Fast Semi-Supervised Node Embedding Learning via Structural and Label-Aware Optimization
Sujan Chakraborty, Rahul Bordoloi, Anindya Sengupta, Olaf Wolkenhauer, Saptarshi Bej
Mach. Learn.4
2026 Dependency-aware synthetic tabular data generation
abstract
Synthetic tabular data is increasingly used in privacy-sensitive domains such as healthcare, but existing generative models often fail to preserve inter-attribute relationships. In particular, functional dependencies (FDs) and logical dependencies (LDs), which capture deterministic and rule-based associations between features, are rarely or often poorly retained in synthetic datasets. To address this research gap, we propose the Hierarchical Feature Generation Framework (HFGF) for synthetic tabular data generation. We created benchmark datasets with known dependencies to evaluate our proposed HFGF. The framework first generates independent features using any standard generative model, and then reconstructs dependent features based on predefined FD and LD rules. Our experiments on four benchmark datasets and three publicly available real-world datasets with varying sizes, feature imbalance, and dependency complexity demonstrate that HFGF improves the preservation of FDs and LDs across six generative models, including CTGAN , TVAE , and GReaT . Utility analysis and qualitative dependency visualizations further show that HFGF significantly enhances the structural fidelity and utility of synthetic tabular data. 1
Chaithra Umesh, Kristian Schultz, Manjunath Mahendra, Saptarshi Bej, Olaf Wolkenhauer
Pattern Recognit.5
2025 Joint embedding-classifier learning for interpretable collaborative filtering
abstract
BACKGROUND: Interpretability is a topical question in recommender systems, especially in healthcare applications. An interpretable classifier quantifies the importance of each input feature for the predicted item-user association in a non-ambiguous fashion. RESULTS: We introduce the novel Joint Embedding Learning-classifier for improved Interpretability (JELI). By combining the training of a structured collaborative-filtering classifier and an embedding learning task, JELI predicts new user-item associations based on jointly learned item and user embeddings while providing feature-wise importance scores. Therefore, JELI flexibly allows the introduction of priors on the connections between users, items, and features. In particular, JELI simultaneously (a) learns feature, item, and user embeddings; (b) predicts new item-user associations; (c) provides importance scores for each feature. Moreover, JELI instantiates a generic approach to training recommender systems by encoding generic graph-regularization constraints. CONCLUSIONS: First, we show that the joint training approach yields a gain in the predictive power of the downstream classifier. Second, JELI can recover feature-association dependencies. Finally, JELI induces a restriction in the number of parameters compared to baselines in synthetic and drug-repurposing data sets.
Clémence Réda, Jill-Jênn Vie, Olaf Wolkenhauer
BMC Bioinform.3
2025 Multivariate functional linear discriminant analysis for partially-observed time series
abstract
Abstract The more extensive access to time-series data, especially for biomedical purposes, raises new methodological challenges, particularly regarding missing values. Functional linear discriminant analysis (FLDA) extends Linear Discriminant Analysis (LDA)-mediated multiclass classification and dimension reduction to data in the form of fragmented observations of a univariate function. For large multivariate and partially-observed data, there are two challenges: (i) statistical dependencies between different components of a multivariate function and (ii) heterogeneous sampling times with missing features. We here develop a multivariate version of FLDA, called MUDRA, to tackle these challenges and describe a computationally efficient expectation/conditional-maximisation (ECM) algorithm to infer its parameters without any tensor inversions. We assess its predictive power on the “Articulary Words” dataset and show its improvement over the state-of-the-art, especially in the case of missing data. This advancement in dimension reduction of multivariate functional data holds promise for enhancing classification accuracy in scenarios like partially observed short multivariate time series analysis.
Rahul Bordoloi, Clémence Réda, Orell Trautmann, Saptarshi Bej, Olaf Wolkenhauer
Mach. Learn.5
2025 Preserving logical and functional dependencies in synthetic tabular data
abstract
Dependencies among attributes are a common aspect of tabular data. However, whether existing tabular data generation algorithms preserve these dependencies while generating synthetic data is yet to be explored. In addition to the existing notion of functional dependencies, we introduce the notion of logical dependencies among the attributes in this article. Moreover, we provide a measure to quantify logical dependencies among attributes in tabular data. Utilizing this measure, we compare several state-of-the-art synthetic data generation algorithms and test their capability to preserve logical and functional dependencies on several publicly available datasets. We demonstrate that currently available synthetic tabular data generation algorithms do not fully preserve functional dependencies when they generate synthetic datasets. In addition, we also showed that some tabular synthetic data generation models can preserve inter-attribute logical dependencies. Our review and comparison of the state-of-the-art reveal research needs and opportunities to develop task-specific synthetic tabular data generation models. • We introduced the notion of logical dependencies in tabular data. • We introduce a novel Bayesian measure to identify logical dependencies. • State-of-the-art models are not able to preserve functional dependencies. • Our results show that state-of-the-art algorithms can preserve logical dependencies.
Chaithra Umesh, Kristian Schultz, Manjunath Mahendra, Saptarshi Bej, Olaf Wolkenhauer
Pattern Recognit.5
2024 ConvGeN: A convex space learning approach for deep-generative oversampling and imbalanced classification of small tabular datasets
abstract
Oversampling is commonly used to improve classifier performance for small tabular imbalanced datasets. State-of-the-art linear interpolation approaches can be used to generate synthetic samples from the convex space of the minority class. Generative networks are common deep learning approaches for synthetic sample generation. However, their scope on synthetic tabular data generation in the context of imbalanced classification is not adequately explored. In this article, we show that existing deep generative models perform poorly compared to linear interpolation-based approaches for imbalanced classification problems on small tabular datasets. To overcome this, we propose a deep generative model, ConvGeN that combines the idea of convex space learning with deep generative models . ConvGeN learns coefficients for the convex combinations of the minority class samples, such that the synthetic data is distinct enough from the majority class . Our benchmarking experiments demonstrate that our proposed model ConvGeN improves imbalanced classification on such small datasets, as compared to existing deep generative models, while being on par with the existing linear interpolation approaches. Moreover, we discuss how our model can be used for synthetic tabular data generation in general, even outside the scope of data imbalance, and thus improves the overall applicability of convex space learning.
Kristian Schultz, Saptarshi Bej, Waldemar Hahn, Markus Wolfien, Prashant Srivastava, Olaf Wolkenhauer
Pattern Recognit.6
2022 Melanoma 2.0. Skin cancer as a paradigm for emerging diagnostic technologies, computational modelling and artificial intelligence
abstract
We live in an unprecedented time in oncology. We have accumulated samples and cases in cohorts larger and more complex than ever before. New technologies are available for quantifying solid or liquid samples at the molecular level. At the same time, we are now equipped with the computational power necessary to handle this enormous amount of quantitative data. Computational models are widely used helping us to substantiate and interpret data. Under the label of systems and precision medicine, we are putting all these developments together to improve and personalize the therapy of cancer. In this review, we use melanoma as a paradigm to present the successful application of these technologies but also to discuss possible future developments in patient care linked to them. Melanoma is a paradigmatic case for disruptive improvements in therapies, with a considerable number of metastatic melanoma patients benefiting from novel therapies. Nevertheless, a large proportion of patients does not respond to therapy or suffers from adverse events. Melanoma is an ideal case study to deploy advanced technologies not only due to the medical need but also to some intrinsic features of melanoma as a disease and the skin as an organ. From the perspective of data acquisition, the skin is the ideal organ due to its accessibility and suitability for many kinds of advanced imaging techniques. We put special emphasis on the necessity of computational strategies to integrate multiple sources of quantitative data describing the tumour at different scales and levels.
Julio Vera, Xin Lai 0002, Andreas Baur, Michael Erdmann, Shailendra K. Gupta, Cristiano Guttà, Lucie Heinzerling, Markus V. Heppt, Philipp Maximilian Kazmierczak, Manfred Kunz, Christopher Lischer, Brigitte M. Pützer, Markus Rehm 0001, Christian Ostalecki, Jimmy Retzlaff, Stephan Witt, Olaf Wolkenhauer, Carola Berking
Briefings Bioinform.17
2021 Combining uniform manifold approximation with localized affine shadowsampling improves classification of imbalanced datasets
abstract
Oversampling approaches are a popular choice to improve classification on imbalanced datasets. The SMOTE algorithm is the pioneer for many algorithms, built as extensions of SMOTE, to solve its problem of over-generalization of the minority class. Some extensions adopt the approach of learning the minority class data distribution through clustering and manifold learning techniques. The Localised Random Affine Shadowsampling (LoRAS) algorithm, models the convex space, controlling the local variance of a synthetic sample by constructing them from convex combinations of multiple shadow samples generated by adding Gaussian noise to the original minority samples. LoRAS also uses t-SNE for a manifold learning step to identify minority class data neighbourhoods. The algorithm is known to outperform some early SMOTE extensions, improving F1-Score and Balanced accuracy for highly imbalanced classification problems. However, the state-of-the-art manifold learning algorithm UMAP is known to preserve the local and global structure of the latent data manifold better than t-SNE and is considerably faster. We have integrated the UMAP for manifold learning with localized affine shadowsampling, to build the LoRAS-UMAP algorithm. We have benchmarked the new algorithm LoRAS-UMAP against some state-of-the-art oversampling algorithms on 14 publicly available datasets characterized by high imbalance, high dimensionality, and high absolute imbalance. In summary, we incorporated UMAP for the manifold learning step yielding better F1-Score, Balanced accuracy and runtime for the LoRAS algorithm in comparison to t-SNE for manifold learning, particularly in the case of high-dimensional datasets.
Saptarshi Bej, Prashant Srivastava, Markus Wolfien, Olaf Wolkenhauer
IJCNN4
2021 Automated annotation of rare-cell types from single-cell RNA-sequencing data through synthetic oversampling
abstract
BACKGROUND: The research landscape of single-cell and single-nuclei RNA-sequencing is evolving rapidly. In particular, the area for the detection of rare cells was highly facilitated by this technology. However, an automated, unbiased, and accurate annotation of rare subpopulations is challenging. Once rare cells are identified in one dataset, it is usually necessary to generate further specific datasets to enrich the analysis (e.g., with samples from other tissues). From a machine learning perspective, the challenge arises from the fact that rare-cell subpopulations constitute an imbalanced classification problem. We here introduce a Machine Learning (ML)-based oversampling method that uses gene expression counts of already identified rare cells as an input to generate synthetic cells to then identify similar (rare) cells in other publicly available experiments. We utilize single-cell synthetic oversampling (sc-SynO), which is based on the Localized Random Affine Shadowsampling (LoRAS) algorithm. The algorithm corrects for the overall imbalance ratio of the minority and majority class. RESULTS: We demonstrate the effectiveness of our method for three independent use cases, each consisting of already published datasets. The first use case identifies cardiac glial cells in snRNA-Seq data (17 nuclei out of 8635). This use case was designed to take a larger imbalance ratio (~1 to 500) into account and only uses single-nuclei data. The second use case was designed to jointly use snRNA-Seq data and scRNA-Seq on a lower imbalance ratio (~1 to 26) for the training step to likewise investigate the potential of the algorithm to consider both single-cell capture procedures and the impact of "less" rare-cell types. The third dataset refers to the murine data of the Allen Brain Atlas, including more than 1 million cells. For validation purposes only, all datasets have also been analyzed traditionally using common data analysis approaches, such as the Seurat workflow. CONCLUSIONS: In comparison to baseline testing without oversampling, our approach identifies rare-cells with a robust precision-recall balance, including a high accuracy and low false positive detection rate. A practical benefit of our algorithm is that it can be readily implemented in other and existing workflows. The code basis in R and Python is publicly available at FairdomHub, as well as GitHub, and can easily be transferred to identify other rare-cell types.
Saptarshi Bej, Anne-Marie Galow, Robert David 0003, Markus Wolfien, Olaf Wolkenhauer
BMC Bioinform.5
2021 LoRAS: an oversampling approach for imbalanced datasets
abstract
Abstract The Synthetic Minority Oversampling TEchnique (SMOTE) is widely-used for the analysis of imbalanced datasets. It is known that SMOTE frequently over-generalizes the minority class, leading to misclassifications for the majority class, and effecting the overall balance of the model. In this article, we present an approach that overcomes this limitation of SMOTE, employing Localized Random Affine Shadowsampling (LoRAS) to oversample from an approximated data manifold of the minority class. We benchmarked our algorithm with 14 publicly available imbalanced datasets using three different Machine Learning (ML) algorithms and compared the performance of LoRAS, SMOTE and several SMOTE extensions that share the concept of using convex combinations of minority class data points for oversampling with LoRAS. We observed that LoRAS, on average generates better ML models in terms of F1-Score and Balanced accuracy. Another key observation is that while most of the extensions of SMOTE we have tested, improve the F1-Score with respect to SMOTE on an average, they compromise on the Balanced accuracy of a classification model. LoRAS on the contrary, improves both F1 Score and the Balanced accuracy thus produces better classification models. Moreover, to explain the success of the algorithm, we have constructed a mathematical framework to prove that LoRAS oversampling technique provides a better estimate for the mean of the underlying local data distribution of the minority class data space.
Saptarshi Bej, Narek Davtyan, Markus Wolfien, Mariam Nassar, Olaf Wolkenhauer
Mach. Learn.5
2020 GEMtractor: extracting views into genome-scale metabolic models
abstract
SUMMARY: Computational metabolic models typically encode for graphs of species, reactions and enzymes. Comparing genome-scale models through topological analysis of multipartite graphs is challenging. However, in many practical cases it is not necessary to compare the full networks. The GEMtractor is a web-based tool to trim models encoded in SBML. It can be used to extract subnetworks, for example focusing on reaction- and enzyme-centric views into the model. AVAILABILITY AND IMPLEMENTATION: The GEMtractor is licensed under the terms of GPLv3 and developed at github.com/binfalse/GEMtractor-a public version is available at sbi.uni-rostock.de/gemtractor.
Martin Scharm, Olaf Wolkenhauer, Mahdi Jalili, Ali Salehzadeh-Yazdi
Bioinform.2
2020 An integrative network-driven pipeline for systematic identification of lncRNA-associated regulatory network motifs in metastatic melanoma
abstract
BACKGROUND: Melanoma phenotype and the dynamics underlying its progression are determined by a complex interplay between different types of regulatory molecules. In particular, transcription factors (TFs), microRNAs (miRNAs), and long non-coding RNAs (lncRNAs) interact in layers that coalesce into large molecular interaction networks. Our goal here is to study molecules associated with the cross-talk between various network layers, and their impact on tumor progression. RESULTS: To elucidate their contribution to disease, we developed an integrative computational pipeline to construct and analyze a melanoma network focusing on lncRNAs, their miRNA and protein targets, miRNA target genes, and TFs regulating miRNAs. In the network, we identified three-node regulatory loops each composed of lncRNA, miRNA, and TF. To prioritize these motifs for their role in melanoma progression, we integrated patient-derived RNAseq dataset from TCGA (SKCM) melanoma cohort, using a weighted multi-objective function. We investigated the expression profile of the top-ranked motifs and used them to classify patients into metastatic and non-metastatic phenotypes. CONCLUSIONS: The results of this study showed that network motif UCA1/AKT1/hsa-miR-125b-1 has the highest prediction accuracy (ACC = 0.88) for discriminating metastatic and non-metastatic melanoma phenotypes. The observation is also confirmed by the progression-free survival analysis where the patient group characterized by the metastatic-type expression profile of the motif suffers a significant reduction in survival. The finding suggests a prognostic value of network motifs for the classification and treatment of melanoma.
Nivedita Singh, Martin Eberhardt, Olaf Wolkenhauer, Julio Vera, Shailendra K. Gupta
BMC Bioinform.3
2019 Harmonizing semantic annotations for computational models in biology
abstract
Life science researchers use computational models to articulate and test hypotheses about the behavior of biological systems. Semantic annotation is a critical component for enhancing the interoperability and reusability of such models as well as for the integration of the data needed for model parameterization and validation. Encoded as machine-readable links to knowledge resource terms, semantic annotations describe the computational or biological meaning of what models and data represent. These annotations help researchers find and repurpose models, accelerate model composition and enable knowledge integration across model repositories and experimental data stores. However, realizing the potential benefits of semantic annotation requires the development of model annotation standards that adhere to a community-based annotation protocol. Without such standards, tool developers must account for a variety of annotation formats and approaches, a situation that can become prohibitively cumbersome and which can defeat the purpose of linking model elements to controlled knowledge resource terms. Currently, no consensus protocol for semantic annotation exists among the larger biological modeling community. Here, we report on the landscape of current annotation practices among the COmputational Modeling in BIology NEtwork community and provide a set of recommendations for building a consensus approach to semantic annotation.
Maxwell Lewis Neal, Matthias König 0003, David P. Nickerson, Goksel Misirli, Reza Kalbasi, Andreas Dräger, Koray Atalag, Vijayalakshmi Chelliah, Mike T. Cooling, Daniel L. Cook, Sharon M. Crook, Miguel de Alba, Samuel H. Friedman, Alan Garny, John H. Gennari, Padraig Gleeson, Martin Golebiewski, Michael Hucka, Nick S. Juty, Chris J. Myers, Brett G. Olivier, Herbert M. Sauro, Martin Scharm, Jacky L. Snoep, Vasundra Touré, Anil Wipat, Olaf Wolkenhauer, Dagmar Waltemath
Briefings Bioinform.27
2018 Quick tips for creating effective and impactful biological pathways using the Systems Biology Graphical Notation
abstract
Quick tips for creating effective and impactful biological pathways using the Systems Biology Graphical Notation
Vasundra Touré, Nicolas Le Novère, Dagmar Waltemath, Olaf Wolkenhauer
PLoS Comput. Biol.4
2016 Personalized cancer immunotherapy using Systems Medicine approaches
abstract
The immune system is by definition multi-scale because it involves biochemical networks that regulate cell fates across cell boundaries, but also because immune cells communicate with each other by direct contact or through the secretion of local or systemic signals. Furthermore, tumor and immune cells communicate, and this interaction is affected by the tumor microenvironment. Altogether, the tumor-immunity interaction is a complex multi-scale biological system whose analysis requires a systemic view to succeed in developing efficient immunotherapies for cancer and immune-related diseases. In this review we discuss the necessity and the structure of a systems medicine approach for the design of anticancer immunotherapies. We support the idea that the approach must be a combination of algorithms and methods from bioinformatics and patient-data-driven mathematical models conceived to investigate the role of clinical interventions in the tumor-immunity interaction. For each step of the integrative approach proposed, we review the advancement with respect to the computational tools and methods available, but also successful case studies. We particularized our idea for the case of identifying novel tumor-associated antigens and therapeutic targets by integration of patient's immune and tumor profiling in case of aggressive melanoma.
Shailendra K. Gupta, Tanushree Jaitly, Ulf Schmitz, Gerold Schuler, Olaf Wolkenhauer, Julio Vera
Briefings Bioinform.5
2016 The RNA world in the 21st century - a systems approach to finding non-coding keys to clinical questions
abstract
There was evidence that RNAs are a functionally rich class of molecules not only since the arrival of the next-generation sequencing technology. Non-coding RNAs (ncRNA) could be the key to accelerated diagnosis and enhanced prediction of disease and therapy outcomes as well as the design of advanced therapeutic strategies to overcome yet unsatisfactory approaches.In this review, we discuss the state of the art in RNA systems biology with focus on the application in the systems biomedicine field. We propose guidelines for analysing the role of microRNAs and long non-coding RNAs in human pathologies. We introduce RNA expression profiling and network approaches for the identification of stable and effective RNomics-based biomarkers, providing insights into the role of ncRNAs in disease regulation. Towards this, we discuss ways to model the dynamics of gene regulatory networks and signalling pathways that involve ncRNAs. We also describe data resources and computational methods for finding putative mechanisms of action of ncRNAs. Finally, we discuss avenues for the computer-aided design of novel RNA-based therapeutics.
Ulf Schmitz, Hojjat Naderi-Meshkin, Shailendra K. Gupta, Olaf Wolkenhauer, Julio Vera
Briefings Bioinform.4
2016 An algorithm to detect and communicate the differences in computational models describing biological systems
abstract
MOTIVATION: Repositories support the reuse of models and ensure transparency about results in publications linked to those models. With thousands of models available in repositories, such as the BioModels database or the Physiome Model Repository, a framework to track the differences between models and their versions is essential to compare and combine models. Difference detection not only allows users to study the history of models but also helps in the detection of errors and inconsistencies. Existing repositories lack algorithms to track a model's development over time. RESULTS: Focusing on SBML and CellML, we present an algorithm to accurately detect and describe differences between coexisting versions of a model with respect to (i) the models' encoding, (ii) the structure of biological networks and (iii) mathematical expressions. This algorithm is implemented in a comprehensive and open source library called BiVeS. BiVeS helps to identify and characterize changes in computational models and thereby contributes to the documentation of a model's history. Our work facilitates the reuse and extension of existing models and supports collaborative modelling. Finally, it contributes to better reproducibility of modelling results and to the challenge of model provenance. AVAILABILITY AND IMPLEMENTATION: The workflow described in this article is implemented in BiVeS. BiVeS is freely available as source code and binary from sems.uni-rostock.de. The web interface BudHat demonstrates the capabilities of BiVeS at budhat.sems.uni-rostock.de.
Martin Scharm, Olaf Wolkenhauer, Dagmar Waltemath
Bioinform.2
2016 TRAPLINE: a standardized and automated pipeline for RNA sequencing data analysis, evaluation and annotation
abstract
BACKGROUND: Technical advances in Next Generation Sequencing (NGS) provide a means to acquire deeper insights into cellular functions. The lack of standardized and automated methodologies poses a challenge for the analysis and interpretation of RNA sequencing data. We critically compare and evaluate state-of-the-art bioinformatics approaches and present a workflow that integrates the best performing data analysis, data evaluation and annotation methods in a Transparent, Reproducible and Automated PipeLINE (TRAPLINE) for RNA sequencing data processing (suitable for Illumina, SOLiD and Solexa). RESULTS: Comparative transcriptomics analyses with TRAPLINE result in a set of differentially expressed genes, their corresponding protein-protein interactions, splice variants, promoter activity, predicted miRNA-target interactions and files for single nucleotide polymorphism (SNP) calling. The obtained results are combined into a single file for downstream analysis such as network construction. We demonstrate the value of the proposed pipeline by characterizing the transcriptome of our recently described stem cell derived antibiotic selected cardiac bodies ('aCaBs'). CONCLUSION: TRAPLINE supports NGS-based research by providing a workflow that requires no bioinformatics skills, decreases the processing time of the analysis and works in the cloud. The pipeline is implemented in the biomedical research platform Galaxy and is freely accessible via www.sbi.uni-rostock.de/RNAseqTRAPLINE or the specific Galaxy manual page (https://usegalaxy.org/u/mwolfien/p/trapline---manual).
Markus Wolfien, Christian Rimmbach, Ulf Schmitz, Julia Jeannine Jung, Stefan Krebs, Gustav Steinhoff, Robert David 0003, Olaf Wolkenhauer
BMC Bioinform.8
2013 Improving the reuse of computational models through version control
abstract
Abstract Motivation: Only models that are accessible to researchers can be reused. As computational models evolve over time, a number of different but related versions of a model exist. Consequently, tools are required to manage not only well-curated models but also their associated versions. Results: In this work, we discuss conceptual requirements for model version control. Focusing on XML formats such as Systems Biology Markup Language and CellML, we present methods for the identification and explanation of differences and for the justification of changes between model versions. In consequence, researchers can reflect on these changes, which in turn have considerable value for the development of new models. The implementation of model version control will therefore foster the exploration of published models and increase their reusability. Availability: We have implemented the proposed methods in a software library called Biochemical Model Version Control System. It is freely available at http://sems.uni-rostock.de/bives/. Biochemical Model Version Control System is also integrated in the online application BudHat, which is available for testing at http://sems.uni-rostock.de/budhat/ (The version described in this publication is available from http://budhat-demo.sems.uni-rostock.de/). Contact: [email protected]
Dagmar Waltemath, Ron Henkel, Robert Hälke, Martin Scharm, Olaf Wolkenhauer
Bioinform.5
2012 Parameter Identifiability and Sensitivity Analysis Predict Targets for Enhancement of STAT1 Activity in Pancreatic Cancer and Stellate Cells
abstract
The present work exemplifies how parameter identifiability analysis can be used to gain insights into differences in experimental systems and how uncertainty in parameter estimates can be handled. The case study, presented here, investigates interferon-gamma (IFNγ) induced STAT1 signalling in two cell types that play a key role in pancreatic cancer development: pancreatic stellate and cancer cells. IFNγ inhibits the growth for both types of cells and may be prototypic of agents that simultaneously hit cancer and stroma cells. We combined time-course experiments with mathematical modelling to focus on the common situation in which variations between profiles of experimental time series, from different cell types, are observed. To understand how biochemical reactions are causing the observed variations, we performed a parameter identifiability analysis. We successfully identified reactions that differ in pancreatic stellate cells and cancer cells, by comparing confidence intervals of parameter value estimates and the variability of model trajectories. Our analysis shows that useful information can also be obtained from nonidentifiable parameters. For the prediction of potential therapeutic targets we studied the consequences of uncertainty in the values of identifiable and nonidentifiable parameters. Interestingly, the sensitivity of model variables is robust against parameter variations and against differences between IFNγ induced STAT1 signalling in pancreatic stellate and cancer cells. This provides the basis for a prediction of therapeutic targets that are valid for both cell types.
Katja Rateitschak, Felix Winter, Falko Lange, Robert Jaster, Olaf Wolkenhauer
PLoS Comput. Biol.5
2011 Agent-based Simulation of Molecular Processes - An Application to Actin-polymerisation
Stefan Pauleweit, J. Barbara Nebe, Olaf Wolkenhauer
SIMULTECH3
2011 Minimum Information About a Simulation Experiment (MIASE)
abstract
This FAIRsharing record describes: The MIASE Guidelines, initiated by the BioModels.net effort, are a community effort to identify the Minimal Information About a Simulation Experiment, necessary to enable the reproducible simulation experiments. Consequently, the MIASE Guidelines list the information that a modeller needs to provide to enable the execution and reproduction of a numerical simulation experiment, derived from a given set of quantitative models. MIASE is a set of guidelines suitable for use with any structured format for simulation experiments. As such, MIASE is designed to help modelers and software tools to exchange their simulation settings and to foster collaboration.
Dagmar Waltemath, Richard R. Adams, Daniel A. Beard, Frank T. Bergmann, Upinder S. Bhalla, Randall Britten, Vijayalakshmi Chelliah, Mike T. Cooling, Jonathan Cooper, Edmund J. Crampin, Alan Garny, Stefan Hoops, Michael Hucka, Peter J. Hunter, Edda Klipp, Camille Laibe, Andrew K. Miller, Ion I. Moraru, David P. Nickerson, Poul M. F. Nielsen, Macha Nikolski, Sven Sahle, Herbert M. Sauro, Henning Schmidt, Jacky L. Snoep, Dominic P. Tolle, Olaf Wolkenhauer, Nicolas Le Novère
PLoS Comput. Biol.27
2010 Non-coding RNA detection methods combined to improve usability, reproducibility and precision
abstract
BACKGROUND: Non-coding RNAs gain more attention as their diverse roles in many cellular processes are discovered. At the same time, the need for efficient computational prediction of ncRNAs increases with the pace of sequencing technology. Existing tools are based on various approaches and techniques, but none of them provides a reliable ncRNA detector yet. Consequently, a natural approach is to combine existing tools. Due to a lack of standard input and output formats combination and comparison of existing tools is difficult. Also, for genomic scans they often need to be incorporated in detection workflows using custom scripts, which decreases transparency and reproducibility. RESULTS: We developed a Java-based framework to integrate existing tools and methods for ncRNA detection. This framework enables users to construct transparent detection workflows and to combine and compare different methods efficiently. We demonstrate the effectiveness of combining detection methods in case studies with the small genomes of Escherichia coli, Listeria monocytogenes and Streptococcus pyogenes. With the combined method, we gained 10% to 20% precision for sensitivities from 30% to 80%. Further, we investigated Streptococcus pyogenes for novel ncRNAs. Using multiple methods--integrated by our framework--we determined four highly probable candidates. We verified all four candidates experimentally using RT-PCR. CONCLUSIONS: We have created an extensible framework for practical, transparent and reproducible combination and comparison of ncRNA detection methods. We have proven the effectiveness of this approach in tests and by guiding experiments to find new ncRNAs. The software is freely available under the GNU General Public License (GPL), version 3 at http://www.sbi.uni-rostock.de/moses along with source code, screen shots, examples and tutorial material.
Peter Raasch, Ulf Schmitz, Nadja Patenge, Julio Vera, Bernd Kreikemeyer, Olaf Wolkenhauer
BMC Bioinform.6
2007 Interpreting Rosen
abstract
Some reflections on Robert Rosen, Chu and Ho, and Louie.
Olaf Wolkenhauer
Artif. Life1
2007 SBML export interface for the systems biology toolbox for MATLAB
abstract
UNLABELLED: In this application note, we present an Systems biology markup language (SBML) export interface for the Systems Biology Toolbox for MATLAB. This interface allows modelers to automatically convert models, represented in the toolbox's own format (SBmodels) to SBML files. Since SBmodels do not explicitly contain all the information that is required to generate SBML, the necessary information is gathered by parsing SBmodels. The export can be done in two different ways. First, it is possible to call the export from the command line, thereby directly converting a model to an SBML file. The second option is to inspect and edit the conversion results with the help of a graphical user interface and to subsequently export the model to SBML. AVAILABILITY: The SBML export interface has been integrated into the Systems Biology Toolbox for MATLAB, which is open source and freely available from http://www.sbtoolbox.org. The website also contains a tutorial, extensive documentation and examples.
Henning Schmidt, Gunnar Drews, Julio Vera, Olaf Wolkenhauer
Bioinform.4
2007 PLMaddon: a power-law module for the MatlabTM SBToolbox
abstract
Abstract Summary: PLMaddon is a General Public License (GPL) software module designed to expand the current version of the SBToolbox (a Matlab ™ toolbox for systems biology; www.sbtoolbox.org) with a set of functions for the analysis of power-law models, a specific class of kinetic models, set in ordinary differential equations (ODE) and in which the kinetic orders can have positive/negative non-integer values. The module includes functions to generate power-law Taylor expansions of other ODE models (e.g. Michaelis-Menten type models), as well as algorithms to estimate steady-states. The robustness and sensitivity of the models can also be analysed and visualized by computing the power-law's logarithmic gains and sensitivities. Availability: PLMaddon is an open source module for the analysis of power-law models based on the SBToolbox. The latest version of PLMaddon is freely available from: www.sbi.uni-rostock.de/plmaddon The website contains a tutorial with examples, as well as an interactive introductory course on power-law models in systems biology. Contact: [email protected]
Julio Vera, Yvonne Oertel, Olaf Wolkenhauer
Bioinform.4
2005 Clustering of unevenly sampled gene expression time-series data
Carla S. Möller-Levet, Frank Klawonn, Kwang-Hyun Cho, Hujun Yin, Olaf Wolkenhauer
Fuzzy Sets Syst.5
2004 Modelling gene expression time-series with radial basis function neural networks
abstract
Gene expression time-series are discrete, noisy, short and usually unevenly sampled. Most of the existing methods used to compare expression profiles, operate directly on the time points. While modelling, the profiles can lead to more generalised, smooth characterisation of gene expressions. In this paper, a radial basis function neural network is employed to model gene expression time-series. The orthogonal least square method, used for selection of centres, is further combined with a width optimisation scheme. The experiments on a number of expression datasets have shown the advantages of the approach in terms of generalisation and approximation. The results on known datasets have indeed coincided with biological interpretations.
Carla S. Möller-Levet, Kwang-Hyun Cho, Hujun Yin, Olaf Wolkenhauer
IJCNN4
2004 Advanced significance analysis of microarray data based on weighted resampling: a comparative study and application to gene deletions in Mycobacterium bovis
abstract
Abstract Motivation: When analyzing microarray data, non-biological variation introduces uncertainty in the analysis and interpretation. In this paper we focus on the validation of significant differences in gene expression levels, or normalized channel intensity levels with respect to different experimental conditions and with replicated measurements. A myriad of methods have been proposed to study differences in gene expression levels and to assign significance values as a measure of confidence. In this paper we compare several methods, including SAM, regularized t-test, mixture modeling, Wilk's lambda score and variance stabilization. From this comparison we developed a weighted resampling approach and applied it to gene deletions in Mycobacterium bovis. Results: We discuss the assumptions, model structure, computational complexity and applicability to microarray data. The results of our study justified the theoretical basis of the weighted resampling approach, which clearly outperforms the others. Availability: Algorithms were implemented using the statistical programming language R and available on the author's web-page. Supplementary information: For additional material see http://www.sbi.uni-rostock.de/
Zoltán Kutalik, Jacqueline Inwald, Steve V. Gordon, R. Glyn Hewinson, Philip D. Butcher, Jason Hinds, Kwang-Hyun Cho, Olaf Wolkenhauer
Bioinform.8
2003 Fuzzy Clustering of Short Time-Series and Unevenly Distributed Sampling Points
Carla S. Möller-Levet, Frank Klawonn, Kwang-Hyun Cho, Olaf Wolkenhauer
IDA4
2003 Level sets and minimum volume sets of probability density functions
Javier Nunez-Garcia, Zoltán Kutalik, Kwang-Hyun Cho, Olaf Wolkenhauer
Int. J. Approx. Reason.4
2002 Random set system identification
abstract
The paper gives a brief review of the basic mathematical aspects of random set theory. Concepts such as a random set mapping and its coverage function are introduced in a comprehensive way, avoiding too much detail. We adapt this theory to system identification and forecasting of time series. This is achieved by using the one-point coverage function of a random set as a possibility measure of the process which generates such a time series. The coverage function of a random set defines a fuzzy set, and we thereby establish the relationship between statistical objects and fuzzy systems. The possibility measure obtained in this way can be used for either prediction or to evaluate the quality of a model with respect to the training data. The technique is adapted to nonlinear time series analysis. A practical application of a nonlinear dynamic plant is presented.
Javier Nunez-Garcia, Olaf Wolkenhauer
IEEE Trans. Fuzzy Syst.2
2001 Random Sets and Histograms
abstract
One of the main reasons why histograms are the most used density estimators is that they are easier to implement and interpret than other density estimators. Some people have already exploited the connection between probability theory and possibility theory or fuzzy sets to set up membership functions and to create fuzzy sets models. Two different ways have been used: 1) transform the density function of a random variable into a possibility measure, which is an almost automatic operation; and 2) calculate the coverage function of a random set, which is a possibility measure. In this paper, we show that a histogram is the coverage function of a determined random set. This suggests other methods to create more accurate or different featured histograms by using the random set theory. One example of a histogram with overlapping classes is provided.
Javier Nunez-Garcia, Olaf Wolkenhauer
FUZZ-IEEE2
2001 Systems Biology: the Reincarnation of Systems Theory Applied in Biology?
abstract
With the availability of quantitative data on the transcriptome and proteome level, there is an increasing interest in formal mathematical models of gene expression and regulation. International conferences, research institutes and research groups concerned with systems biology have appeared in recent years and systems theory, the study of organisation and behaviour per se, is indeed a natural conceptual framework for such a task. This is, however, not the first time that systems theory has been applied in modelling cellular processes. Notably in the 1960s systems theory and biology enjoyed considerable interest among eminent scientists, mathematicians and engineers. Why did these early attempts vanish from research agendas? Here we shall review the domain of systems theory, its application to biology and the lessons that can be learned from the work of Robert Rosen. Rosen emerged from the early developments in the 1960s as a main critic but also developed a new alternative perspective to living systems, a concept that deserves a fresh look in the post-genome era of bioinformatics.
Olaf Wolkenhauer
Briefings Bioinform.1
1997 Qualitative Uncertainty Models from Random Set Theory
Olaf Wolkenhauer
IDA1
1997 Possibilistic Testing of Distribution Functions for Change Detection
abstract
Change detection algorithms are proposed that are based on the comparison of distribution functions. Estimated values of distributions are associated with a binomial distribution that is used to define fuzzy similarity classes. Fuzzy concepts are used to combine partial evaluations to a measure that indicates the departure of a signal from its reference.
Olaf Wolkenhauer, John M. Edmunds
Intell. Data Anal.1