Martin Hofmann-Apitius

dblp:45/5227 · also Martin Hofmann 0009 · DBLP profile ↗
← Back
35ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0001-9012-6720ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 12 since 2021Systems, architecture and hardware · 4Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 KGG: a fully automated workflow for creating disease-specific knowledge graphs
abstract
MOTIVATION: Knowledge graphs (KGs) in life sciences have become an important application of systems biology as they delineate complex biological and pathophysiological phenomena. They are composed of biological and chemical entities represented with standard ontologies to comply with Findable, Accessible, Interoperable and Reusable (FAIR) principles. Alongside serving as a graph database, KGs hold the potential to address complex scientific queries and facilitate downstream analyses. However, the process of constructing KGs is expensive and time consuming as it primarily relies on manual curation from published literature and experimental data. The existing text-mining workflows are still in their infancy and fail to achieve the accuracy and reliability of manual curation. RESULTS: Knowledge graph generator (KGG) is an automated workflow for representing chemotype and phenotype of diseases and medical conditions. It embeds the underlying schema of curated databases such as OpenTargets, Uniprot, ChEMBL, Integrated Interactions Database and GWAS Central resembling a clockwork-esque mechanism. The resultant KG is a comprehensive and rational assembly of disease-associated entities such as proteins, protein-related pathways, biological processes and functions, genetic variants, chemicals, mechanism of actions, assays and adverse effects. As use cases, we have used KGs to identify shared entities for possible link of comorbidity and compared them with KGs from other sources. We have also demonstrated a use case of identifying putative new targets and repurposing drug candidates in Parkinson's Disease. Lastly, we have developed reusable workflows to explore drug-likeness of chemicals and identify structures of proteins. AVAILABILITY AND IMPLEMENTATION: The resources and codes for KGG are publicly available at: https://github.com/Fraunhofer-ITMP/kgg.
Reagon Karki, Yojana Gadiya, Andrea Zaliani, Bishab Pokharel, Negin Sadat Babaiha, Marek Ostaszewski, Martin Hofmann-Apitius, Philip Gribbon
Bioinform.7
2023 PEMT: a patent enrichment tool for drug discovery
abstract
MOTIVATION: Drug discovery practitioners in industry and academia use semantic tools to extract information from online scientific literature to generate new insights into targets, therapeutics and diseases. However, due to complexities in access and analysis, patent-based literature is often overlooked as a source of information. As drug discovery is a highly competitive field, naturally, tools that tap into patent literature can provide any actor in the field an advantage in terms of better informed decision-making. Hence, we aim to facilitate access to patent literature through the creation of an automatic tool for extracting information from patents described in existing public resources. RESULTS: Here, we present PEMT, a novel patent enrichment tool, that takes advantage of public databases like ChEMBL and SureChEMBL to extract relevant patent information linked to chemical structures and/or gene names described through FAIR principles and metadata annotations. PEMT aims at supporting drug discovery and research by establishing a patent landscape around genes of interest. The pharmaceutical focus of the tool is mainly due to the subselection of International Patent Classification codes, but in principle, it can be used for other patent fields, provided that a link between a concept and chemical structure is investigated. Finally, we demonstrate a use-case in rare diseases by generating a gene-patent list based on the epidemiological prevalence of these diseases and exploring their underlying patent landscapes. AVAILABILITY AND IMPLEMENTATION: PEMT is an open-source Python tool and its source code and PyPi package are available at https://github.com/Fraunhofer-ITMP/PEMT and https://pypi.org/project/PEMT/, respectively. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yojana Gadiya, Andrea Zaliani, Philip Gribbon, Martin Hofmann-Apitius
Bioinform.4
2022 On the influence of several factors on pathway enrichment analysis
abstract
Pathway enrichment analysis has become a widely used knowledge-based approach for the interpretation of biomedical data. Its popularity has led to an explosion of both enrichment methods and pathway databases. While the elegance of pathway enrichment lies in its simplicity, multiple factors can impact the results of such an analysis, which may not be accounted for. Researchers may fail to give influential aspects their due, resorting instead to popular methods and gene set collections, or default settings. Despite ongoing efforts to establish set guidelines, meaningful results are still hampered by a lack of consensus or gold standards around how enrichment analysis should be conducted. Nonetheless, such concerns have prompted a series of benchmark studies specifically focused on evaluating the influence of various factors on pathway enrichment results. In this review, we organize and summarize the findings of these benchmarks to provide a comprehensive overview on the influence of these factors. Our work covers a broad spectrum of factors, spanning from methodological assumptions to those related to prior biological knowledge, such as pathway definitions and database choice. In doing so, we aim to shed light on how these aspects can lead to insignificant, uninteresting or even contradictory results. Finally, we conclude the review by proposing future benchmarks as well as solutions to overcome some of the challenges, which originate from the outlined factors.
Sarah Mubeen, Alpha Tom Kodamullil, Martin Hofmann-Apitius, Daniel Domingo-Fernández
Briefings Bioinform.3
2022 STonKGs: a sophisticated transformer trained on biomedical text and knowledge graphs
abstract
MOTIVATION: The majority of biomedical knowledge is stored in structured databases or as unstructured text in scientific publications. This vast amount of information has led to numerous machine learning-based biological applications using either text through natural language processing (NLP) or structured data through knowledge graph embedding models. However, representations based on a single modality are inherently limited. RESULTS: To generate better representations of biological knowledge, we propose STonKGs, a Sophisticated Transformer trained on biomedical text and Knowledge Graphs (KGs). This multimodal Transformer uses combined input sequences of structured information from KGs and unstructured text data from biomedical literature to learn joint representations in a shared embedding space. First, we pre-trained STonKGs on a knowledge base assembled by the Integrated Network and Dynamical Reasoning Assembler consisting of millions of text-triple pairs extracted from biomedical literature by multiple NLP systems. Then, we benchmarked STonKGs against three baseline models trained on either one of the modalities (i.e. text or KG) across eight different classification tasks, each corresponding to a different biological application. Our results demonstrate that STonKGs outperforms both baselines, especially on the more challenging tasks with respect to the number of classes, improving upon the F1-score of the best baseline by up to 0.084 (i.e. from 0.881 to 0.965). Finally, our pre-trained model as well as the model architecture can be adapted to various other transfer learning applications. AVAILABILITY AND IMPLEMENTATION: We make the source code and the Python package of STonKGs available at GitHub (https://github.com/stonkgs/stonkgs) and PyPI (https://pypi.org/project/stonkgs/). The pre-trained STonKGs models and the task-specific classification models are respectively available at https://huggingface.co/stonkgs/stonkgs-150k and https://zenodo.org/communities/stonkgs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Helena Balabin, Charles Tapley Hoyt, Colin Birkenbihl, Benjamin M. Gyori, John A. Bachman, Alpha Tom Kodamullil, Paul-Gerhard Plöger, Martin Hofmann-Apitius, Daniel Domingo-Fernández
Bioinform.8
2022 Common data model for COVID-19 datasets
abstract
MOTIVATION: A global medical crisis like the coronavirus disease 2019 (COVID-19) pandemic requires interdisciplinary and highly collaborative research from all over the world. One of the key challenges for collaborative research is a lack of interoperability among various heterogeneous data sources. Interoperability, standardization and mapping of datasets are necessary for data analysis and applications in advanced algorithms such as developing personalized risk prediction modeling. RESULTS: To ensure the interoperability and compatibility among COVID-19 datasets, we present here a common data model (CDM) which has been built from 11 different COVID-19 datasets from various geographical locations. The current version of the CDM holds 4639 data variables related to COVID-19 such as basic patient information (age, biological sex and diagnosis) as well as disease-specific data variables, for example, Anosmia and Dyspnea. Each of the data variables in the data model is associated with specific data types, variable mappings, value ranges, data units and data encodings that could be used for standardizing any dataset. Moreover, the compatibility with established data standards like OMOP and FHIR makes the CDM a well-designed CDM for COVID-19 data interoperability. AVAILABILITY AND IMPLEMENTATION: The CDM is available in a public repo here: https://github.com/Fraunhofer-SCAI-Applied-Semantics/COVID-19-Global-Model. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Philipp Wegner, Geena Mariya Jose, Vanessa Lage-Rupprecht, Sepehr Golriz Khatami, Bide Zhang, Stephan Springstubbe, Marc Jacobs 0001, Thomas Linden, Cindy Ku, Bruce Schultz, Martin Hofmann-Apitius, Alpha Tom Kodamullil
Bioinform.11
2022 Integrative data semantics through a model-enabled data stewardship
abstract
MOTIVATION: The importance of clinical data in understanding the pathophysiology of complex disorders has prompted the launch of multiple initiatives designed to generate patient-level data from various modalities. While these studies can reveal important findings relevant to the disease, each study captures different yet complementary aspects and modalities which, when combined, generate a more comprehensive picture of disease etiology. However, achieving this requires a global integration of data across studies, which proves to be challenging given the lack of interoperability of cohort datasets. RESULTS: Here, we present the Data Steward Tool (DST), an application that allows for semi-automatic semantic integration of clinical data into ontologies and global data models and data standards. We demonstrate the applicability of the tool in the field of dementia research by establishing a Clinical Data Model (CDM) in this domain. The CDM currently consists of 277 common variables covering demographics (e.g. age and gender), diagnostics, neuropsychological tests and biomarker measurements. The DST combined with this disease-specific data model shows how interoperability between multiple, heterogeneous dementia datasets can be achieved. AVAILABILITY AND IMPLEMENTATION: The DST source code and Docker images are respectively available at https://github.com/SCAI-BIO/data-steward and https://hub.docker.com/r/phwegner/data-steward. Furthermore, the DST is hosted at https://data-steward.bio.scai.fraunhofer.de/data-steward. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Philipp Wegner, Sebastian Schaaf, Mischa Uebachs, Daniel Domingo-Fernández, Yasamin Salimi, Stephan Gebel, Astghik Sargsyan, Colin Birkenbihl, Stephan Springstubbe, Thomas Klockgether, Juliane Fluck, Martin Hofmann-Apitius, Alpha Tom Kodamullil
Bioinform.12
2022 Elucidating gene expression patterns across multiple biological contexts through a large-scale investigation of transcriptomic datasets
abstract
Distinct gene expression patterns within cells are foundational for the diversity of functions and unique characteristics observed in specific contexts, such as human tissues and cell types. Though some biological processes commonly occur across contexts, by harnessing the vast amounts of available gene expression data, we can decipher the processes that are unique to a specific context. Therefore, with the goal of developing a portrait of context-specific patterns to better elucidate how they govern distinct biological processes, this work presents a large-scale exploration of transcriptomic signatures across three different contexts (i.e., tissues, cell types, and cell lines) by leveraging over 600 gene expression datasets categorized into 98 subcontexts. The strongest pairwise correlations between genes from these subcontexts are used for the construction of co-expression networks. Using a network-based approach, we then pinpoint patterns that are unique and common across these subcontexts. First, we focused on patterns at the level of individual nodes and evaluated their functional roles using a human protein-protein interactome as a referential network. Next, within each context, we systematically overlaid the co-expression networks to identify specific and shared correlations as well as relations already described in scientific literature. Additionally, in a pathway-level analysis, we overlaid node and edge sets from co-expression networks against pathway knowledge to identify biological processes that are related to specific subcontexts or groups of them. Finally, we have released our data and scripts at https://zenodo.org/record/5831786 and https://github.com/ContNeXt/ , respectively and developed ContNeXt ( https://contnext.scai.fraunhofer.de/ ), a web application to explore the networks generated in this work.
Rebeca Queiroz Figueiredo, Sara Díaz del Ser, Tamara Raschka, Martin Hofmann-Apitius, Alpha Tom Kodamullil, Sarah Mubeen, Daniel Domingo-Fernández
BMC Bioinform.4
2022 GuiltyTargets: Prioritization of Novel Therapeutic Targets With Network Representation Learning
abstract
The majority of clinical trials fail due to low efficacy of investigated drugs, often resulting from a poor choice of target protein. Existing computational approaches aim to support target selection either via genetic evidence or by putting potential targets into the context of a disease specific network reconstruction. The purpose of this work was to investigate whether network representation learning techniques could be used to allow for a machine learning based prioritization of putative targets. We propose a novel target prioritization approach, GuiltyTargets, which relies on attributed network representation learning of a genome-wide protein-protein interaction network annotated with disease-specific differential gene expression and uses positive-unlabeled (PU) machine learning for candidate ranking. We evaluated our approach on 12 datasets from six diseases of different type (cancer, metabolic, neurodegenerative) within a 10 times repeated 5-fold stratified cross-validation and achieved AUROC values between 0.92 - 0.97, significantly outperforming previous approaches that relied on manually engineered topological features. Moreover, we showed that GuiltyTargets allows for target repositioning across related disease areas. An application of GuiltyTargets to Alzheimer's disease resulted in a number of highly ranked candidates that are currently discussed as targets in the literature. Interestingly, one (COMT) is also the target of an approved drug (Tolcapone) for Parkinson's disease, highlighting the potential for target repositioning with our method. The GuiltyTargets Python package is available on PyPI and all code used for analysis can be found under the MIT License at https://github.com/GuiltyTargets. Attributed network representation learning techniques provide an interesting approach to effectively leverage the existing knowledge about the molecular mechanisms in different diseases. In this work, the combination with positive-unlabeled learning for target prioritization demonstrated a clear superiority compared to classical feature engineering approaches. Our work highlights the potential of attributed network representation learning for target prioritization. Given the overarching relevance of networks in computational biology we believe that attributed network representation learning techniques could have a broader impact in the future.
Ozlem Muslu, Charles Tapley Hoyt, Mauricio Lacerda, Martin Hofmann-Apitius, Holger Fröhlich
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 CLEP: a hybrid data- and knowledge-driven framework for generating patient representations
abstract
SUMMARY: As machine learning and artificial intelligence increasingly attain a larger number of applications in the biomedical domain, at their core, their utility depends on the data used to train them. Due to the complexity and high dimensionality of biomedical data, there is a need for approaches that combine prior knowledge around known biological interactions with patient data. Here, we present CLinical Embedding of Patients (CLEP), a novel approach that generates new patient representations by leveraging both prior knowledge and patient-level data. First, given a patient-level dataset and a knowledge graph containing relations across features that can be mapped to the dataset, CLEP incorporates patients into the knowledge graph as new nodes connected to their most characteristic features. Next, CLEP employs knowledge graph embedding models to generate new patient representations that can ultimately be used for a variety of downstream tasks, ranging from clustering to classification. We demonstrate how using new patient representations generated by CLEP significantly improves performance in classifying between patients and healthy controls for a variety of machine learning models, as compared to the use of the original transcriptomics data. Furthermore, we also show how incorporating patients into a knowledge graph can foster the interpretation and identification of biological features characteristic of a specific disease or patient subgroup. Finally, we released CLEP as an open source Python package together with examples and documentation. AVAILABILITY AND IMPLEMENTATION: CLEP is available to the bioinformatics community as an open source Python package at https://github.com/hybrid-kg/clep under the Apache 2.0 License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vinay Srinivas Bharadhwaj, Mehdi Ali, Colin Birkenbihl, Sarah Mubeen, Jens Lehmann 0001, Martin Hofmann-Apitius, Charles Tapley Hoyt, Daniel Domingo-Fernández
Bioinform.6
2021 COVID-19 Knowledge Graph: a computable, multi-modal, cause-and-effect knowledge model of COVID-19 pathophysiology
abstract
SUMMARY: The COVID-19 crisis has elicited a global response by the scientific community that has led to a burst of publications on the pathophysiology of the virus. However, without coordinated efforts to organize this knowledge, it can remain hidden away from individual research groups. By extracting and formalizing this knowledge in a structured and computable form, as in the form of a knowledge graph, researchers can readily reason and analyze this information on a much larger scale. Here, we present the COVID-19 Knowledge Graph, an expansive cause-and-effect network constructed from scientific literature on the new coronavirus that aims to provide a comprehensive view of its pathophysiology. To make this resource available to the research community and facilitate its exploration and analysis, we also implemented a web application and released the KG in multiple standard formats. AVAILABILITY AND IMPLEMENTATION: The COVID-19 Knowledge Graph is publicly available under CC-0 license at https://github.com/covid19kg and https://bikmi.covid19-knowledgespace.de. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel Domingo-Fernández, Shounak Baksi, Bruce Schultz, Yojana Gadiya, Reagon Karki, Tamara Raschka, Christian Ebeling, Martin Hofmann-Apitius, Alpha Tom Kodamullil
Bioinform.8
2021 MultiPaths: a Python framework for analyzing multi-layer biological networks using diffusion algorithms
abstract
SUMMARY: High-throughput screening yields vast amounts of biological data which can be highly challenging to interpret. In response, knowledge-driven approaches emerged as possible solutions to analyze large datasets by leveraging prior knowledge of biomolecular interactions represented in the form of biological networks. Nonetheless, given their size and complexity, their manual investigation quickly becomes impractical. Thus, computational approaches, such as diffusion algorithms, are often employed to interpret and contextualize the results of high-throughput experiments. Here, we present MultiPaths, a framework consisting of two independent Python packages for network analysis. While the first package, DiffuPy, comprises numerous commonly used diffusion algorithms applicable to any generic network, the second, DiffuPath, enables the application of these algorithms on multi-layer biological networks. To facilitate its usability, the framework includes a command line interface, reproducible examples and documentation. To demonstrate the framework, we conducted several diffusion experiments on three independent multi-omics datasets over disparate networks generated from pathway databases, thus, highlighting the ability of multi-layer networks to integrate multiple modalities. Finally, the results of these experiments demonstrate how the generation of harmonized networks from disparate databases can improve predictive performance with respect to individual resources. AVAILABILITY AND IMPLEMENTATION: DiffuPy and DiffuPath are publicly available under the Apache License 2.0 at https://github.com/multipaths. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Josep Marín-Llaó, Sarah Mubeen, Alexandre Perera-Lluna, Martin Hofmann-Apitius, Sergio Picart-Armada, Daniel Domingo-Fernández
Bioinform.4
2021 The COVID-19 Ontology
abstract
MOTIVATION: The COVID-19 pandemic has prompted an impressive, worldwide response by the academic community. In order to support text mining approaches as well as data description, linking and harmonization in the context of COVID-19, we have developed an ontology representing major novel coronavirus (SARS-CoV-2) entities. The ontology has a strong scope on chemical entities suited for drug repurposing, as this is a major target of ongoing COVID-19 therapeutic development. RESULTS: The ontology comprises 2270 classes of concepts and 38 987 axioms (2622 logical axioms and 2434 declaration axioms). It depicts the roles of molecular and cellular entities in virus-host interactions and in the virus life cycle, as well as a wide spectrum of medical and epidemiological concepts linked to COVID-19. The performance of the ontology has been tested on Medline and the COVID-19 corpus provided by the Allen Institute. AVAILABILITYAND IMPLEMENTATION: COVID-19 Ontology is released under a Creative Commons 4.0 License and shared via https://github.com/covid-19-ontology/covid-19. The ontology is also deposited in BioPortal at https://bioportal.bioontology.org/ontologies/COVID-19. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Astghik Sargsyan, Alpha Tom Kodamullil, Shounak Baksi, Johannes Darms, Sumit Madan, Stephan Gebel, Oliver Keminer, Geena Mariya Jose, Helena Balabin, Lauren Nicole Delong, Manfred Kohler, Marc Jacobs 0001, Martin Hofmann-Apitius
Bioinform.13
2020 PS4DR: a multimodal workflow for identification and prioritization of drugs based on pathway signatures
abstract
BACKGROUND: During the last decade, there has been a surge towards computational drug repositioning owing to constantly increasing -omics data in the biomedical research field. While numerous existing methods focus on the integration of heterogeneous data to propose candidate drugs, it is still challenging to substantiate their results with mechanistic insights of these candidate drugs. Therefore, there is a need for more innovative and efficient methods which can enable better integration of data and knowledge for drug repositioning. RESULTS: Here, we present a customizable workflow (PS4DR) which not only integrates high-throughput data such as genome-wide association study (GWAS) data and gene expression signatures from disease and drug perturbations but also takes pathway knowledge into consideration to predict drug candidates for repositioning. We have collected and integrated publicly available GWAS data and gene expression signatures for several diseases and hundreds of FDA-approved drugs or those under clinical trial in this study. Additionally, different pathway databases were used for mechanistic knowledge integration in the workflow. Using this systematic consolidation of data and knowledge, the workflow computes pathway signatures that assist in the prediction of new indications for approved and investigational drugs. CONCLUSION: We showcase PS4DR with applications demonstrating how this tool can be used for repositioning and identifying new drugs as well as proposing drugs that can simulate disease dysregulations. We were able to validate our workflow by demonstrating its capability to predict FDA-approved drugs for their known indications for several diseases. Further, PS4DR returned many potential drug candidates for repositioning that were backed up by epidemiological evidence extracted from scientific literature. Source code is freely available at https://github.com/ps4dr/ps4dr.
Mohammad Asif Emon, Daniel Domingo-Fernández, Charles Tapley Hoyt, Martin Hofmann-Apitius
BMC Bioinform.4
2020 Drug2ways: Reasoning over causal paths in biological networks for drug discovery
abstract
Elucidating the causal mechanisms responsible for disease can reveal potential therapeutic targets for pharmacological intervention and, accordingly, guide drug repositioning and discovery. In essence, the topology of a network can reveal the impact a drug candidate may have on a given biological state, leading the way for enhanced disease characterization and the design of advanced therapies. Network-based approaches, in particular, are highly suited for these purposes as they hold the capacity to identify the molecular mechanisms underlying disease. Here, we present drug2ways, a novel methodology that leverages multimodal causal networks for predicting drug candidates. Drug2ways implements an efficient algorithm which reasons over causal paths in large-scale biological networks to propose drug candidates for a given disease. We validate our approach using clinical trial information and demonstrate how drug2ways can be used for multiple applications to identify: i) single-target drug candidates, ii) candidates with polypharmacological properties that can optimize multiple targets, and iii) candidates for combination therapy. Finally, we make drug2ways available to the scientific community as a Python package that enables conducting these applications on multiple standard network formats.
Daniel Rivas-Barragan, Sarah Mubeen, Francesc Guim 0001, Martin Hofmann-Apitius, Daniel Domingo-Fernández
PLoS Comput. Biol.4
2019 PathMe: merging and exploring mechanistic pathway knowledge
abstract
BACKGROUND: The complexity of representing biological systems is compounded by an ever-expanding body of knowledge emerging from multi-omics experiments. A number of pathway databases have facilitated pathway-centric approaches that assist in the interpretation of molecular signatures yielded by these experiments. However, the lack of interoperability between pathway databases has hindered the ability to harmonize these resources and to exploit their consolidated knowledge. Such a unification of pathway knowledge is imperative in enhancing the comprehension and modeling of biological abstractions. RESULTS: Here, we present PathMe, a Python package that transforms pathway knowledge from three major pathway databases into a unified abstraction using Biological Expression Language as the pivotal, integrative schema. PathMe is complemented by a novel web application (freely available at https://pathme.scai.fraunhofer.de/ ) which allows users to comprehensively explore pathway crosstalk and compare areas of consensus and discrepancies. CONCLUSIONS: This work has harmonized three major pathway databases and transformed them into a unified schema in order to gain a holistic picture of pathway knowledge. We demonstrate the utility of the PathMe framework in: i) integrating pathway landscapes at the database level, ii) comparing the degree of consensus at the pathway level, and iii) exploring pathway crosstalk and investigating consensus at the molecular level.
Daniel Domingo-Fernández, Sarah Mubeen, Josep Marín-Llaó, Charles Tapley Hoyt, Martin Hofmann-Apitius
BMC Bioinform.5
2019 Quantifying mechanisms in neurodegenerative diseases (NDDs) using candidate mechanism perturbation amplitude (CMPA) algorithm
abstract
BACKGROUND: Literature derived knowledge assemblies have been used as an effective way of representing biological phenomenon and understanding disease etiology in systems biology. These include canonical pathway databases such as KEGG, Reactome and WikiPathways and disease specific network inventories such as causal biological networks database, PD map and NeuroMMSig. The represented knowledge in these resources delineates qualitative information focusing mainly on the causal relationships between biological entities. Genes, the major constituents of knowledge representations, tend to express differentially in different conditions such as cell types, brain regions and disease stages. A classical approach of interpreting a knowledge assembly is to explore gene expression patterns of the individual genes. However, an approach that enables quantification of the overall impact of differentially expressed genes in the corresponding network is still lacking. RESULTS: Using the concept of heat diffusion, we have devised an algorithm that is able to calculate the magnitude of regulation of a biological network using expression datasets. We have demonstrated that molecular mechanisms specific to Alzheimer (AD) and Parkinson Disease (PD) regulate with different intensities across spatial and temporal resolutions. Our approach depicts that the mitochondrial dysfunction in PD is severe in cortex and advanced stages of PD patients. Similarly, we have shown that the intensity of aggregation of neurofibrillary tangles (NFTs) in AD increases as the disease progresses. This finding is in concordance with previous studies that explain the burden of NFTs in stages of AD. CONCLUSIONS: This study is one of the first attempts that enable quantification of mechanisms represented as biological networks. We have been able to quantify the magnitude of regulation of a biological network and illustrate that the magnitudes are different across spatial and temporal resolution.
Reagon Karki, Alpha Tom Kodamullil, Charles Tapley Hoyt, Martin Hofmann-Apitius
BMC Bioinform.4
2018 BEL2ABM: agent-based simulation of static models in Biological Expression Language
abstract
Summary: While cause-and-effect knowledge assembly models encoded in Biological Expression Language are able to support generation of mechanistic hypotheses, they are static and limited in their ability to encode temporality. Here, we present BEL2ABM, a software for producing continuous, dynamic, executable agent-based models from BEL templates. Availability and implementation: The tool has been developed in Java and NetLogo. Code, data and documentation are available under the Apache 2.0 License at https://github.com/pybel/bel2abm. Supplementary information: Supplementary data are available at Bioinformatics online.
Michaela Gündel, Charles Tapley Hoyt, Martin Hofmann-Apitius
Bioinform.3
2017 Multimodal mechanistic signatures for neurodegenerative diseases (NeuroMMSig): a web server for mechanism enrichment
abstract
MOTIVATION: The concept of a 'mechanism-based taxonomy of human disease' is currently replacing the outdated paradigm of diseases classified by clinical appearance. We have tackled the paradigm of mechanism-based patient subgroup identification in the challenging area of research on neurodegenerative diseases. RESULTS: We have developed a knowledge base representing essential pathophysiology mechanisms of neurodegenerative diseases. Together with dedicated algorithms, this knowledge base forms the basis for a 'mechanism-enrichment server' that supports the mechanistic interpretation of multiscale, multimodal clinical data. AVAILABILITY AND IMPLEMENTATION: NeuroMMSig is available at http://neurommsig.scai.fraunhofer.de/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daniel Domingo-Fernández, Alpha Tom Kodamullil, Anandhi Iyappan, Mufassra Naz, Mohammad Asif Emon, Tamara Raschka, Reagon Karki, Stephan Springstubbe, Christian Ebeling, Martin Hofmann-Apitius
Bioinform.10
2016 Reasoning over genetic variance information in cause-and-effect models of neurodegenerative diseases
abstract
The work we present here is based on the recent extension of the syntax of the Biological Expression Language (BEL), which now allows for the representation of genetic variation information in cause-and-effect models. In our article, we describe, how genetic variation information can be used to identify candidate disease mechanisms in diseases with complex aetiology such as Alzheimer's disease and Parkinson's disease. In those diseases, we have to assume that many genetic variants contribute moderately to the overall dysregulation that in the case of neurodegenerative diseases has such a long incubation time until the first clinical symptoms are detectable. Owing to the multilevel nature of dysregulation events, systems biomedicine modelling approaches need to combine mechanistic information from various levels, including gene expression, microRNA (miRNA) expression, protein-protein interaction, genetic variation and pathway. OpenBEL, the open source version of BEL, has recently been extended to match this requirement, and we demonstrate in our article, how candidate mechanisms for early dysregulation events in Alzheimer's disease can be identified based on an integrative mining approach that identifies 'chains of causation' that include single nucleotide polymorphism information in BEL models.
Mufassra Naz, Alpha Tom Kodamullil, Martin Hofmann-Apitius
Briefings Bioinform.3
2014 A new optimization phase for scientific workflow management systems
Sonja Holl, Olav Zimmermann, Magnus Palmblad, Yassene Mohammed, Martin Hofmann-Apitius
Future Gener. Comput. Syst.5
2013 'HypothesisFinder: ' A Strategy for the Detection of Speculative Statements in Scientific Text
abstract
Speculative statements communicating experimental findings are frequently found in scientific articles, and their purpose is to provide an impetus for further investigations into the given topic. Automated recognition of speculative statements in scientific text has gained interest in recent years as systematic analysis of such statements could transform speculative thoughts into testable hypotheses. We describe here a pattern matching approach for the detection of speculative statements in scientific text that uses a dictionary of speculative patterns to classify sentences as hypothetical. To demonstrate the practical utility of our approach, we applied it to the domain of Alzheimer's disease and showed that our automated approach captures a wide spectrum of scientific speculations on Alzheimer's disease. Subsequent exploration of derived hypothetical knowledge leads to generation of a coherent overview on emerging knowledge niches, and can thus provide added value to ongoing research activities.
Ashutosh Malhotra, Erfan Younesi, Harsha Gurulingappa, Martin Hofmann-Apitius
PLoS Comput. Biol.4
2012 A new optimization phase for scientific workflow management systems
abstract
Scientific workflows have emerged as an important tool for combining computational power with data analysis for all scientific domains in e-science. They help scientists to design and execute complex in silico experiments. However, with increasing complexity it becomes more and more infeasible to optimize scientific workflows by trial and error. To address this issue, this paper describes the design of a new optimization phase integrated in the established scientific workflow life cycle. We have also developed a flexible optimization application programming interface (API) and have integrated it into a scientific workflow management system. A sample plugin for parameter optimization based on genetic algorithms illustrates, how the API enables rapid implementation of concrete workflow optimization methods. Finally, a use case taken from the area of structural bioinformatics validates how the optimization approach facilitates setup, execution and monitoring of workflow parameter optimization in high performance computing e-science environments.
Sonja Holl, Olav Zimmermann, Martin Hofmann-Apitius
eScience3
2012 Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports
Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, Luca Toldo
J. Biomed. Informatics5
2011 A UNICORE Plugin for HPC-Enabled Scientific Workflows in Taverna 2.2
abstract
As scientific workflows are becoming more complex and apply compute-intensive methods to increasingly large data volumes, access to HPC resources is becoming mandatory. We describe the development of a novel plug in for the Tavern a workflow system, which provides transparent and secure access to HPC/grid resources via the UNICORE grid middleware, while maintaining the ease of use that has been the main reason for the success of scientific workflow systems. A use case from the bioinformatics domain demonstrates the potential of the UNICORE plug in for Tavern a by creating a scientific workflow that executes the central parts in parallel on a cluster resource.
Sonja Holl, Olav Zimmermann, Martin Hofmann-Apitius
SERVICES3
2011 PLIO: an ontology for formal description of protein-ligand interactions
abstract
MOTIVATION: Biomedical ontologies have proved to be valuable tools for data analysis and data interoperability. Protein-ligand interactions are key players in drug discovery and development; however, existing public ontologies that describe the knowledge space of biomolecular interactions do not cover all aspects relevant to pharmaceutical modelling and simulation. RESULTS: The protein--ligand interaction ontology (PLIO) was developed around three main concepts, namely target, ligand and interaction, and was enriched by adding synonyms, useful annotations and references. The quality of the ontology was assessed based on structural, functional and usability features. Validation of the lexicalized ontology by means of natural language processing (NLP)-based methods showed a satisfactory performance (F-score = 81%). Through integration into our information retrieval environment we can demonstrate that PLIO supports lexical search in PubMed abstracts. The usefulness of PLIO is demonstrated by two use-case scenarios and it is shown that PLIO is able to capture both confirmatory and new knowledge from simulation and empirical studies. AVAILABILITY: The PLIO ontology is made freely available to the public at http://www.scai.fraunhofer.de/bioinformatics/downloads.html.
Olga Ivchenko, Erfan Younesi, Antje Wolf, Bernd Müller, Martin Hofmann-Apitius
Bioinform.6
2011 Challenges in the association of human single nucleotide polymorphism mentions with unique database identifiers
abstract
BACKGROUND: Most information on genomic variations and their associations with phenotypes are covered exclusively in scientific publications rather than in structured databases. These texts commonly describe variations using natural language; database identifiers are seldom mentioned. This complicates the retrieval of variations, associated articles, as well as information extraction, e. g. the search for biological implications. To overcome these challenges, procedures to map textual mentions of variations to database identifiers need to be developed. RESULTS: This article describes a workflow for normalization of variation mentions, i.e. the association of them to unique database identifiers. Common pitfalls in the interpretation of single nucleotide polymorphism (SNP) mentions are highlighted and discussed. The developed normalization procedure achieves a precision of 98.1 % and a recall of 67.5% for unambiguous association of variation mentions with dbSNP identifiers on a text corpus based on 296 MEDLINE abstracts containing 527 mentions of SNPs. The annotated corpus is freely available at http://www.scai.fraunhofer.de/snp-normalization-corpus.html. CONCLUSIONS: Comparable approaches usually focus on variations mentioned on the protein sequence and neglect problems for other SNP mentions. The results presented here indicate that normalizing SNPs described on DNA level is more difficult than the normalization of SNPs described on protein level. The challenges associated with normalization are exemplified with ambiguities and errors, which occur in this corpus.
Philippe Thomas 0002, Roman Klinger, Laura Inés Furlong, Martin Hofmann-Apitius, Christoph M. Friedrich
BMC Bioinform.4
2009 Identification of histone modifications in biomedical text for supporting epigenomic research
abstract
BACKGROUND: Posttranslational modifications of histones influence the structure of chromatine and in such a way take part in the regulation of gene expression. Certain histone modification patterns, distributed over the genome, are connected to cell as well as tissue differentiation and to the adaption of organisms to their environment. Abnormal changes instead influence the development of disease states like cancer. The regulation mechanisms for modifying histones and its functionalities are the subject of epigenomics investigation and are still not completely understood. Text provides a rich resource of knowledge on epigenomics and modifications of histones in particular. It contains information about experimental studies, the conditions used, and results. To our knowledge, no approach has been published so far for identifying histone modifications in text. RESULTS: We have developed an approach for identifying histone modifications in biomedical literature with Conditional Random Fields (CRF) and for resolving the recognized histone modification term variants by term standardization. For the term identification F1 measures of 0.84 by 10-fold cross-validation on the training corpus and 0.81 on an independent test corpus have been obtained. The standardization enabled the correct transformation of 96% of the terms from training and 98% from test the corpus. Due to the lack of terminologies exhaustively covering specific histone modification types, we developed a histone modification term hierarchy for use in a semantic text retrieval system. CONCLUSION: The developed approach highly improves the retrieval of articles describing histone modifications. Since text contains context information about performed studies and experiments, the identification of histone modifications is the basis for supporting literature-based knowledge discovery and hypothesis generation to accelerate epigenomic research.
Corinna Kolárik, Roman Klinger, Martin Hofmann-Apitius
BMC Bioinform.3
2008 @neurIST - Towards a System Architecture for Advanced Disease Management through Integration of Heterogeneous Data, Computing, and Complex Processing Services
abstract
This paper presents the system architecture of the @neurIST project, which aims at supporting the research and treatment of cerebral aneurysms by bringing together heterogeneous data, computing and complex processing services. The architecture is generic enough to adapt it to the treatment of other diseases beyond cerebral aneurysms. The paper describes the generic requirements of the system and presents the architecture, applications and middleware technologies used to realise the system and highlights the innovations in @neurIST.
Hariharan Rajasekaran, Luigi Lo Iacono, Peer Hasselmeyer, Jochen Fingberg, Paul E. Summers, Siegfried Benkner, Gerhard Engelbrecht, Antonio Arbona, Alessandro Chiarini, Christoph M. Friedrich, Martin Hofmann-Apitius, Kai Kumpf, Bob Moore, Philippe Bijlenga, Jimison Iavindrasana, Henning Müller, Rod D. Hose, Robert Dunlop, Alejandro F. Frangi
CBMS11
2008 Detection of IUPAC and IUPAC-like chemical names
abstract
MOTIVATION: Chemical compounds like small signal molecules or other biological active chemical substances are an important entity class in life science publications and patents. Several representations and nomenclatures for chemicals like SMILES, InChI, IUPAC or trivial names exist. Only SMILES and InChI names allow a direct structure search, but in biomedical texts trivial names and Iupac like names are used more frequent. While trivial names can be found with a dictionary-based approach and in such a way mapped to their corresponding structures, it is not possible to enumerate all IUPAC names. In this work, we present a new machine learning approach based on conditional random fields (CRF) to find mentions of IUPAC and IUPAC-like names in scientific text as well as its evaluation and the conversion rate with available name-to-structure tools. RESULTS: We present an IUPAC name recognizer with an F(1) measure of 85.6% on a MEDLINE corpus. The evaluation of different CRF orders and offset conjunction orders demonstrates the importance of these parameters. An evaluation of hand-selected patent sections containing large enumerations and terms with mixed nomenclature shows a good performance on these cases (F(1) measure 81.5%). Remaining recognition problems are to detect correct borders of the typically long terms, especially when occurring in parentheses or enumerations. We demonstrate the scalability of our implementation by providing results from a full MEDLINE run. AVAILABILITY: We plan to publish the corpora, annotation guideline as well as the conditional random field model as a UIMA component.
Roman Klinger, Corinna Kolárik, Juliane Fluck, Martin Hofmann-Apitius, Christoph M. Friedrich
ISMB4
2008 OSIRISv1.2: A named entity recognition system for sequence variants of genes in biomedical literature
abstract
BACKGROUND: Single Nucleotide Polymorphisms, among other type of sequence variants, constitute key elements in genetic epidemiology and pharmacogenomics. While sequence data about genetic variation is found at databases such as dbSNP, clues about the functional and phenotypic consequences of the variations are generally found in biomedical literature. The identification of the relevant documents and the extraction of the information from them are hampered by the large size of literature databases and the lack of widely accepted standard notation for biomedical entities. Thus, automatic systems for the identification of citations of allelic variants of genes in biomedical texts are required. RESULTS: Our group has previously reported the development of OSIRIS, a system aimed at the retrieval of literature about allelic variants of genes http://ibi.imim.es/osirisform.html. Here we describe the development of a new version of OSIRIS (OSIRISv1.2, http://ibi.imim.es/OSIRISv1.2.html) which incorporates a new entity recognition module and is built on top of a local mirror of the MEDLINE collection and HgenetInfoDB: a database that collects data on human gene sequence variations. The new entity recognition module is based on a pattern-based search algorithm for the identification of variation terms in the texts and their mapping to dbSNP identifiers. The performance of OSIRISv1.2 was evaluated on a manually annotated corpus, resulting in 99% precision, 82% recall, and an F-score of 0.89. As an example, the application of the system for collecting literature citations for the allelic variants of genes related to the diseases intracranial aneurysm and breast cancer is presented. CONCLUSION: OSIRISv1.2 can be used to link literature references to dbSNP database entries with high accuracy, and therefore is suitable for collecting current knowledge on gene sequence variations and supporting the functional annotation of variation databases. The application of OSIRISv1.2 in combination with controlled vocabularies like MeSH provides a way to identify associations of biomedical interest, such as those that relate SNPs with diseases.
Laura Inés Furlong, Holger Dach, Martin Hofmann-Apitius, Ferran Sanz
BMC Bioinform.3
2008 Public microarray repository semantic annotation with ontologies employing text mining and expression profile correlation
David Ruau, Corinna Kolárik, Heinz-Theodor Mevissen, Emmanuel Müller, Ira Assent, Ralph Krieger, Thomas Seidl 0001, Martin Hofmann-Apitius, Martin Zenke
BMC Bioinform.8
2008 Grid-enabled Virtual Screening Against Malaria
Nicolas Jacq, Jean Salzemann, Florence Jacq, Yannick Legré, Emmanuel Medernach, Johan Montagnat, Astrid Maaß, Matthieu Reichstadt, Horst Schwichtenberg, Mahendrakar Sridhar, Vinod Kasam, Marc Jacobs 0001, Martin Hofmann-Apitius, Vincent Breton
J. Grid Comput.13
2008 Grid-Added Value to Address Malaria
abstract
Through this paper, we call for a distributed, Internet-based collaboration to address one of the worst plagues of our present world, malaria. The spirit is a nonproprietary peer-production of information-embedding goods. And we propose to use the grid technology to enable such a worldwide "open-source" like collaboration. The first step toward this vision has been achieved during the summer 2005 on the enabling grids for E-scienceE (EGEE) grid infrastructure where 42 million ligands were docked for a total amount of 80 CPU years in 6 weeks in the quest for new drugs. The impact of this first deployment has significantly raised the interest of the research community so that several laboratories all around the world expressed interest to propose targets for a second large-scale deployment against malaria.
Vincent Breton, Nicolas Jacq, Vinod Kasam, Martin Hofmann-Apitius
IEEE Trans. Inf. Technol. Biomed.4
2007 Virtual screening on large scale grids
Nicolas Jacq, Vincent Breton, Hsin-Yen Chen, Li-Yung Ho, Martin Hofmann-Apitius, Vinod Kasam, Hurng-Chun Lee, Yannick Legré, Simon C. Lin, Astrid Maaß
Parallel Comput.5
2006 Grid Added Value to Address Malaria
Vincent Breton, Nicolas Jacq, Martin Hofmann-Apitius
CCGRID3