Wolfram Weckwerth

dblp:40/23 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-9719-6358ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 since 2021
YearPublicationVenuePosition
2026 Kernel-DMD for multiome data integration and control
Iro Pierides, Hannes M. Kramml, Steffen Waldherr, Wolfram Weckwerth
PLoS Comput. Biol.4
2024 Polygenic Risk Score Simulation: A case study of sudden cardiovascular diseases derived from literature insights
abstract
Sudden cardiovascular diseases, including sudden cardiac arrest (SCA) and myocardial infarction (MI), remain among the leading causes of mortality worldwide. This study employs synthetic data simulations to develop polygenic risk scores (PRS) for improved identification of individuals at elevated risk. Leveraging genetic and phenotypic information from established literature, we modeled associations between multiple single nucleotide polymorphisms (SNPs) and the risk of sudden cardiovascular diseases. Our case study demonstrated that individuals with higher PRS values have a significantly increased risk of SCA and MI, with SNPs such as LTA (252A>G) identified as notable contributors. Additionally, combining PRS with traditional cardiovascular risk factors, such as smoking and diabetes, improved predictive accuracy, underscoring the value of integrating genetic data with clinical variables. These findings highlight the cumulative effect of genetic predispositions in determining cardiovascular risk and suggest potential applications of PRS models in personalized medicine to enhance preventive strategiesWhile promising, the study recognizes limitations related to the use of synthetic data and the need for validation in diverse, real-world populations. Future research should focus on refining these models and exploring additional genetic markers to further improve prediction capabilities.
Jana Schwarzerova, Lenka Piherova, Petra Polakovicova, Simona Guziurova, Eva Kutilova, Martina Adamova, Alice Krebsova, Valentine Provazník, Wolfram Weckwerth, Radka Sitkova
BIBM9
2024 A perspective on genetic and polygenic risk scores - advances and limitations and overview of associated tools
abstract
Polygenetic Risk Scores are used to evaluate an individual's vulnerability to developing specific diseases or conditions based on their genetic composition, by taking into account numerous genetic variations. This article provides an overview of the concept of Polygenic Risk Scores (PRS). We elucidate the historical advancements of PRS, their advantages and shortcomings in comparison with other predictive methods, and discuss their conceptual limitations in light of the complexity of biological systems. Furthermore, we provide a survey of published tools for computing PRS and associated resources. The various tools and software packages are categorized based on their technical utility for users or prospective developers. Understanding the array of available tools and their limitations is crucial for accurately assessing and predicting disease risks, facilitating early interventions, and guiding personalized healthcare decisions. Additionally, we also identify potential new avenues for future bioinformatic analyzes and advancements related to PRS.
Jana Schwarzerova, Martin Hurta, Vojtech Barton, Matej Lexa, Dirk Walther 0001, Valentine Provazník, Wolfram Weckwerth
Briefings Bioinform.7
2023 Utilizing Genetic Programming to Enhance Polygenic Risk Score Calculation
abstract
The polygenic risk score has proven to be a valuable tool for assessing an individual's genetic predisposition to phenotype (disease) within biomedicine in recent years. However, traditional regression-based methods for polygenic risk scores calculation have limitations that can impede their accuracy and predictive power. This study introduces an innovative approach to enhance polygenic risk scores calculation through the application of genetic programming. By harnessing the power of genetic programming, we aim to overcome the limitations of traditional regression techniques and improve the accuracy of polygenic risk scores predictions. Specifically, we showed that a polygenic risk score generated through Cartesian genetic programming yielded comparable or even more robust statistical distinctions between groups that we evaluated within three independent case studies.
Martin Hurta, Jana Schwarzerova, Thomas Nägele, Wolfram Weckwerth, Valentine Provazník, Lukás Sekanina
BIBM4
2023 COVRECON: automated integration of genome- and metabolome-scale network reconstruction and data-driven inverse modeling of metabolic interaction networks
abstract
MOTIVATION: One central goal of systems biology is to infer biochemical regulations from large-scale OMICS data. Many aspects of cellular physiology and organismal phenotypes can be understood as results of metabolic interaction network dynamics. Previously, we have proposed a convenient mathematical method, which addresses this problem using metabolomics data for the inverse calculation of biochemical Jacobian matrices revealing regulatory checkpoints of biochemical regulations. The proposed algorithms for this inference are limited by two issues: they rely on structural network information that needs to be assembled manually, and they are numerically unstable due to ill-conditioned regression problems for large-scale metabolic networks. RESULTS: To address these problems, we developed a novel regression loss-based inverse Jacobian algorithm, combining metabolomics COVariance and genome-scale metabolic RECONstruction, which allows for a fully automated, algorithmic implementation of the COVRECON workflow. It consists of two parts: (i) Sim-Network and (ii) inverse differential Jacobian evaluation. Sim-Network automatically generates an organism-specific enzyme and reaction dataset from Bigg and KEGG databases, which is then used to reconstruct the Jacobian's structure for a specific metabolomics dataset. Instead of directly solving a regression problem as in the previous workflow, the new inverse differential Jacobian is based on a substantially more robust approach and rates the biochemical interactions according to their relevance from large-scale metabolomics data. The approach is illustrated by in silico stochastic analysis with differently sized metabolic networks from the BioModels database and applied to a real-world example. The characteristics of the COVRECON implementation are that (i) it automatically reconstructs a data-driven superpathway model; (ii) more general network structures can be investigated, and (iii) the new inverse algorithm improves stability, decreases computation time, and extends to large-scale models. AVAILABILITY AND IMPLEMENTATION: The code is available in the website https://bitbucket.org/mosys-univie/covrecon.
Steffen Waldherr, Wolfram Weckwerth
Bioinform.3
2022 Systems biology approach for analysis of mobile genetic elements in chicken gut microbiome
abstract
Antibiotic use in farming for decades has led to the selection pressure on chicken commensal bacteria which adapted to those changes by the acquisition of antibiotic resistance genes via horizontal gene transfer. Extensive gene transfer between gut bacteria is mediated by mobile genetic elements such as plasmids, integrative conjugative elements and temperate phages. Therefore, chicken gut commensal bacteria are considered reservoirs of antibiotic resistance genes. Recently, we initiated a systemic cultivation of chicken gut anaerobes followed by subsequent whole genome sequencing and analysis. In this study, we implemented a novel in silico approach using systems biology methods to detect horizontal gene traits and mobile genetic elements associated with it in the genomes of chicken gut commensals. In total, we identified 1,427 genes envisaged to drive horizontal gene transfer between individual chicken gut microbiota members across different families. Except for genes known to be vertically transferred, we also revealed hypothetical genes which may represent important, yet unknown section of mobilome.
Jana Schwarzerova, Michal Zeman, Ivan Rychlik, Wolfram Weckwerth, Valentine Provazník, Monika Dolejská, Darina Cejková
BIBM4
2021 An Innovative Perspective on Metabolomics Data Analysis in Biomedical Research Using Concept Drift Detection
abstract
One of the most challenging scenarios of data analysis is prediction using time series data. As the underlying causal relationships of the data shift over time, a classification model trained on data at earlier points within the course starts to yield incorrect predictions on the current data. This phenomenon in machine learning is called concept drift. Within biomedical data, one of the molecular networks that changes significantly over a time is the metabolome. Using metabolomics analysis in biomedical applications produces an ideal tool in preventive healthcare, the pharmaceutical industry, and even ecology engineering. This study provides an innovative perspective on the analysis of metabolomics datasets using the concept of drift detection. The evaluation is based on two main objectives. The first objective is connected to the concept drift detection in available metabolomics datasets, and the second objective is to provide the assessment of commonly used machine learning tools for the best general detection approach in metabolomics datasets. The application of concept drift to metabolomics data has never been carried out before and is an original take on the analysis of highly dynamic molecular networks.
Jana Schwarzerova, Adam Bajger, Iro Pierdou, Lubos Popelínský, Karel Sedlár, Wolfram Weckwerth
BIBM6
2009 Detection and characterization of 3D-signature phosphorylation site motifs and their contribution towards improved phosphorylation site prediction in proteins
abstract
BACKGROUND: Phosphorylation of proteins plays a crucial role in the regulation and activation of metabolic and signaling pathways and constitutes an important target for pharmaceutical intervention. Central to the phosphorylation process is the recognition of specific target sites by protein kinases followed by the covalent attachment of phosphate groups to the amino acids serine, threonine, or tyrosine. The experimental identification as well as computational prediction of phosphorylation sites (P-sites) has proved to be a challenging problem. Computational methods have focused primarily on extracting predictive features from the local, one-dimensional sequence information surrounding phosphorylation sites. RESULTS: We characterized the spatial context of phosphorylation sites and assessed its usability for improved phosphorylation site predictions. We identified 750 non-redundant, experimentally verified sites with three-dimensional (3D) structural information available in the protein data bank (PDB) and grouped them according to their respective kinase family. We studied the spatial distribution of amino acids around phosphorserines, phosphothreonines, and phosphotyrosines to extract signature 3D-profiles. Characteristic spatial distributions of amino acid residue types around phosphorylation sites were indeed discernable, especially when kinase-family-specific target sites were analyzed. To test the added value of using spatial information for the computational prediction of phosphorylation sites, Support Vector Machines were applied using both sequence as well as structural information. When compared to sequence-only based prediction methods, a small but consistent performance improvement was obtained when the prediction was informed by 3D-context information. CONCLUSION: While local one-dimensional amino acid sequence information was observed to harbor most of the discriminatory power, spatial context information was identified as relevant for the recognition of kinases and their cognate target sites and can be used for an improved prediction of phosphorylation sites. A web-based service (Phos3D) implementing the developed structure-based P-site prediction method has been made available at (http://phos3d.mpimp-golm.mpg.de).
Pawel Durek, Christian Schudoma, Wolfram Weckwerth, Joachim Selbig, Dirk Walther 0001
BMC Bioinform.3
2007 ProMEX: a mass spectral reference database for proteins and protein phosphorylation sites
abstract
BACKGROUND: In the last decade, techniques were established for the large scale genome-wide analysis of proteins, RNA, and metabolites, and database solutions have been developed to manage the generated data sets. The Golm Metabolome Database for metabolite data (GMD) represents one such effort to make these data broadly available and to interconnect the different molecular levels of a biological system 1. As data interpretation in the light of already existing data becomes increasingly important, these initiatives are an essential part of current and future systems biology. RESULTS: A mass spectral library consisting of experimentally derived tryptic peptide product ion spectra was generated based on liquid chromatography coupled to ion trap mass spectrometry (LC-IT-MS). Protein samples derived from Arabidopsis thaliana, Chlamydomonas reinhardii, Medicago truncatula, and Sinorhizobium meliloti were analysed. With currently 4,557 manually validated spectra associated with 4,226 unique peptides from 1,367 proteins, the database serves as a continuously growing reference data set and can be used for protein identification and quantification in uncharacterized biological samples. For peptide identification, several algorithms were implemented based on a recently published study for peptide mass fingerprinting 2 and tested for false positive and negative rates. An algorithm which considers intensity distribution for match correlation scores was found to yield best results. For proof of concept, an LC-IT-MS analysis of a tryptic leaf protein digest was converted to mzData format and searched against the mass spectral library. The utility of the mass spectral library was also tested for the identification of phosphorylated tryptic peptides. We included in vivo phosphorylation sites of Arabidopsis thaliana proteins and the identification performance was found to be improved compared to genome-based search algorithms. Protein identification by ProMEX is linked to other levels of biological organization such as metabolite, pathway, and transcript data. The database is further connected to annotation and classification services via BioMoby. CONCLUSION: The ProMEX protein/peptide database represents a mass spectral reference library with the capability of matching unknown samples for protein identification. The database allows text searches based on metadata such as experimental information of the samples, mass spectrometric instrument parameters or unique protein identifier like AGI codes. ProMEX integrates proteomics data with other levels of molecular organization including metabolite, pathway, and transcript information and may thus become a useful resource for plant systems biology studies. The ProMEX mass spectral library is available at http://promex.mpimp-golm.mpg.de/.
Jan Hummel, Michaela Niemann, Stefanie Wienkoop, Waltraud X. Schulze, Dirk Steinhauser, Joachim Selbig, Dirk Walther 0001, Wolfram Weckwerth
BMC Bioinform.8
2005 [email protected]: the Golm Metabolome Database
abstract
UNLABELLED: Metabolomics, in particular gas chromatography-mass spectrometry (GC-MS) based metabolite profiling of biological extracts, is rapidly becoming one of the cornerstones of functional genomics and systems biology. Metabolite profiling has profound applications in discovering the mode of action of drugs or herbicides, and in unravelling the effect of altered gene expression on metabolism and organism performance in biotechnological applications. As such the technology needs to be available to many laboratories. For this, an open exchange of information is required, like that already achieved for transcript and protein data. One of the key-steps in metabolite profiling is the unambiguous identification of metabolites in highly complex metabolite preparations from biological samples. Collections of mass spectra, which comprise frequently observed metabolites of either known or unknown exact chemical structure, represent the most effective means to pool the identification efforts currently performed in many laboratories around the world. Here we present GMD, The Golm Metabolome Database, an open access metabolome database, which should enable these processes. GMD provides public access to custom mass spectral libraries, metabolite profiling experiments as well as additional information and tools, e.g. with regard to methods, spectral information or compounds. The main goal will be the representation of an exchange platform for experimental research activities and bioinformatics to develop and improve metabolomics by multidisciplinary cooperation. AVAILABILITY: http://csbdb.mpimp-golm.mpg.de/gmd.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://csbdb.mpimp-golm.mpg.de/
Joachim Kopka, Nicolas Schauer, Stephan Krueger, Claudia Birkemeyer, Björn Usadel, Eveline Bergmüller, Peter Dörmann, Wolfram Weckwerth, Yves Gibon, Mark Stitt, Lothar Willmitzer, Alisdair R. Fernie, Dirk Steinhauser
Bioinform.8
2005 Species-specific analysis of protein sequence motifs using mutual information
abstract
BACKGROUND: Protein sequence motifs are by definition short fragments of conserved amino acids, often associated with a specific function. Accordingly protein sequence profiles derived from multiple sequence alignments provide an alternative description of functional motifs characterizing families of related sequences. Such profiles conveniently reflect functional necessities by pointing out proximity at conserved sequence positions as well as depicting distances at variable positions. Discovering significant conservation characteristics within the variable positions of profiles mirrors group-specific and, in particular, evolutionary features of the underlying sequences. RESULTS: We describe the tool PROfile analysis based on Mutual Information (PROMI) that enables comparative analysis of user-classified protein sequences. PROMI is implemented as a web service using Perl and R as well as other publicly available packages and tools on the server-side. On the client-side platform-independence is achieved by generally applied internet delivery standards. As one possible application analysis of the zinc finger C2H2-type protein domain is introduced to illustrate the functionality of the tool. CONCLUSION: The web service PROMI should assist researchers to detect evolutionary correlations in protein profiles of defined biological sequences. It is available at http://promi.mpimp-golm.mpg.de where additional documentation can be found.
Jan Hummel, Nima Keshvari, Wolfram Weckwerth, Joachim Selbig
BMC Bioinform.3
2003 Observing and Interpreting Correlations in Metabolic Networks
abstract
MOTIVATION: Metabolite profiling aims at an unbiased identification and quantification of all the metabolites present in a biological sample. Based on their pair-wise correlations, the data obtained from metabolomic experiments are organized into metabolic correlation networks and the key challenge is to deduce unknown pathways based on the observed correlations. However, the data generated is fundamentally different from traditional biological measurements and thus the analysis is often restricted to rather pragmatic approaches, such as data mining tools, to discriminate between different metabolic phenotypes. METHODS AND RESULTS: We investigate to what extent the data generated networks reflect the structure of the underlying biochemical pathways. The purpose of this work is 2-fold: Based on the theory of stochastic systems, we first introduce a framework which shows that the emergent correlations can be interpreted as a 'fingerprint' of the underlying biophysical system. This result leads to a systematic relationship between observed correlation networks and the underlying biochemical pathways. In a second step, we investigate to what extent our result is applicable to the problem of reverse engineering, i.e. to recover the underlying enzymatic reaction network from data. The implications of our findings for other bioinformatics approaches are discussed.
Ralph E. Steuer, Jürgen Kurths, Oliver Fiehn, Wolfram Weckwerth
Bioinform.4
2001 Visualizing plant metabolomic correlation networks using clique-metabolite matrices
abstract
MOTIVATION: Today, metabolite levels in biological samples can be determined using multiparallel, fast, and precise metabolomic approaches. Correlations between the levels of various metabolites can be searched to gain information about metabolic links. Such correlations are the net result of direct enzymatic conversions and of indirect cellular regulation over transcriptional or biochemical processes. In order to visualize metabolic networks derived from correlation lists graphically, each metabolite pair may be represented as vertices connected by an edge. However, graph complexity rapidly increases with the number of edges and vertices. To gain structural information from metabolite correlation networks, improvements in clarity are needed. RESULTS: To achieve this clarity, three algorithms are combined. First, a list of linear metabolite correlations is generated that can be regarded as a set of pairs of edges (or as 2-cliques). Next, a branch-and-bound algorithm was developed to find all maximal cliques by combining submaximal cliques. Due to a clique assignment procedure, the generation of unnecessary submaximal cliques is avoided in order to maintain high efficiency. Differences and similarities to the Bron-Kerbosch algorithm are pointed out. Lastly, metabolite correlation networks are visualized by clique-metabolite matrices that are sorted to minimize the length of lines that connect different cliques and metabolites. Examples of biochemical hypotheses are given that can be built from interpretation of such clique matrices. AVAILABILITY: The algorithms are implemented in Visual Basic and can be downloaded from our web site along with a test data set (http://www.mpimp-golm.mpg.de/fiehn/projekte/data-mining-e.html). CONTACT: [email protected]
Frank Kose, Wolfram Weckwerth, Thomas Linke, Oliver Fiehn
Bioinform.2