Michel Dumontier

dblp:d/MichelDumontier · also Michel Justin Dumontier · DBLP profile ↗
← Back
61ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0003-4727-9435ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 41 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 AudSemThinker: Enhancing Audio-Language Models Through Reasoning over Semantics of Sound
abstract
Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound. In this paper, we present AudSemThinker, a model whose reasoning is structured around a framework of auditory semantics inspired by human cognition. To support this, we introduce AudSem, a novel dataset specifically curated for semantic descriptor reasoning in audio-language models. AudSem addresses the persistent challenge of data contamination in zero-shot evaluations by providing a carefully filtered collection of audio samples paired with captions generated through a robust multi-stage pipeline. Our experiments demonstrate that AudSemThinker outperforms state-of-the-art models across multiple training settings, highlighting its strength in semantic audio reasoning. Both AudSemThinker and the AudSem dataset are released publicly.
Gijs Wijngaard, Elia Formisano, Michele Esposito, Michel Dumontier
NeurIPS4
2025 Privacy preservation in blockchain-based healthcare data sharing: A systematic review
abstract
Blockchain technology promises enhanced data ownership, control, and interoperability in healthcare, yet security and privacy concerns continue to hinder its adoption. Existing surveys examine blockchain-based privacy challenges, but they lack a systematic analysis and maturity evaluation of privacy-preserving techniques tailored to healthcare data sharing. This paper presents a systematic review of blockchain-based privacy-preserving solutions, analyzing blockchain details, applied privacy methods, regulatory compliance, and maturity levels using Technology Readiness Levels (TRLs). Our findings reveal that authentication and authorization is the most explored stage, dominated by smart contracts and ciphertext-policy attribute-based encryption. Proxy re-encryption is frequently used for data transfer, while privacy-preserving search and verification remain underexplored. On/off-chain mechanisms are commonly applied to balance privacy and storage efficiency. TRL assessment shows that most solutions remain at the proof-of-concept stage (TRL3), with only limited progress to prototype validation (TRL4-TRL5), highlighting the gap between experimental designs and real-world deployment. To guide developers and researchers, we identify two primary patterns of blockchain integration and propose a framework for system design. We also compare methods across data-sharing stages, outlining their strengths and limitations to support informed selection. In conclusion, while research interest is growing, the field remains at an early stage of maturity. Addressing this gap requires stronger implementation capacity, access to clinical data, and robust regulatory alignment. We emphasize the importance of clinical validation and real-world testing to advance privacy-preserving blockchain solutions toward practical adoption in healthcare.
Ankur Lohachab, Michel Dumontier, Visara Urovi
Peer Peer Netw. Appl.3
2023 Graph-based Categorical Embedding
abstract
Categorical features are a challenge for most machine learning algorithms that only accept numerical vectors in input.However, the emergence of graph neural networks is revolutionizing the application of machine learning models to traditional data sets.This is thanks to the possibility of introducing graph relationships amongst features and samples.In this contribution, we describe an algorithm leveraging the assignment matrix of a DiffPool graph neural network to calculate embeddings for categorical features, using as an adjacency matrix the co-occurrence matrix between the categorical values and as nodes feature the one hot encoded categorical values.We show that the algorithm proposed is scalable and presents a competitive performance in three publicly available data sets presenting both numerical and categorical values. 41
Stefano Bromuri, Michel Dumontier
ESANN3
2023 Semantically-Informed Deep Neural Networks For Sound Recognition
abstract
Deep neural networks (DNNs) for sound recognition learn to categorize a barking sound as a "dog" and a meowing sound as a "cat" but do not exploit information inherent to the semantic relations between classes (e.g., both are animal vocalisations). Cognitive neuroscience research, however, suggests that human listeners automatically exploit higher-level semantic information on the sources besides acoustic information. Inspired by this notion, we introduce here a DNN that learns to recognize sounds and simultaneously learns the semantic relation between the sources (semDNN). Comparison of semDNN with a homologous network trained with categorical labels (catDNN) revealed that semDNN produces semantically more accurate labelling than catDNN in sound recognition tasks and that semDNN-embeddings preserve higherlevel semantic relations between sound sources. Importantly, through a model-based analysis of human dissimilarity ratings of natural sounds, we show that semDNN approximates the behaviour of human listeners better than catDNN and several other DNN and NLP comparison models.
Michele Esposito, Giancarlo Valente, Yenisel Plasencia, Michel Dumontier, Bruno L. Giordano, Elia Formisano
ICASSP4
2023 Generating synthetic personal health data using conditional generative adversarial networks combining with differential privacy
abstract
A large amount of personal health data that is highly valuable to the scientific community is still not accessible or requires a lengthy request process due to privacy concerns and legal restrictions. As a solution, synthetic data has been studied and proposed to be a promising alternative to this issue. However, generating realistic and privacy-preserving synthetic personal health data retains challenges such as simulating the characteristics of the patients' data that are in the minority classes, capturing the relations among variables in imbalanced data and transferring them to the synthetic data, and preserving individual patients' privacy. In this paper, we propose a differentially private conditional Generative Adversarial Network model (DP-CGANS) consisting of data transformation, sampling, conditioning, and network training to generate realistic and privacy-preserving personal data. Our model distinguishes categorical and continuous variables and transforms them into latent space separately for better training performance. We tackle the unique challenges of generating synthetic patient data due to the special data characteristics of personal health data. For example, patients with a certain disease are typically the minority in the dataset and the relations among variables are crucial to be observed. Our model is structured with a conditional vector as an additional input to present the minority class in the imbalanced data and maximally capture the dependency between variables. Moreover, we inject statistical noise into the gradients in the networking training process of DP-CGANS to provide a differential privacy guarantee. We extensively evaluate our model with state-of-the-art generative models on personal socio-economic datasets and real-world personal health datasets in terms of statistical similarity, machine learning performance, and privacy measurement. We demonstrate that our model outperforms other comparable models, especially in capturing the dependence between variables. Finally, we present the balance between data utility and privacy in synthetic data generation considering the different data structures and characteristics of real-world personal health data such as imbalanced classes, abnormal distributions, and data sparsity.
Chang Sun 0001, Johan van Soest, Michel Dumontier
J. Biomed. Informatics3
2022 Semantic Correlation Graph Embedding
abstract
Many data sets include categorical features in the form of nominal and ordinal features. However, most machine learning algorithms cannot deal with categorical features directly because they require numerical input features. Categorical embeddings are an effective approach to converting categorical features into numerical vectors. This work proposes a novel embedding approach, called Semantic Correlation Graph Embedding, to create embeddings from knowledge graphs. The approach constructs a semantic correlation graph of triplets among the categorical features to learn numerical embeddings. Our approach aims to uncover relationships taking place in categorical data in terms of low-level knowledge and semantics that may help group the features of the data sets under semantic entities. Three distinct embedding models are proposed according to how the graph is constructed. The results are evaluated with two public data sets. They show that the learned embeddings produce a statistically significant improvement in the performance of the classification tasks in terms of AUC, F1 score, precision, and recall.
Stefano Bromuri, Michel Dumontier
FUZZ-IEEE4
2022 LUCE: A blockchain-based data sharing platform for monitoring data License accoUntability and CompliancE
abstract
Easy access to data is one of the main avenues to accelerate scientific research. As a key element of scientific innovations, data sharing allows the reproduction of results and helps prevent data fabrication, falsification, and misuse. Although the research benefits from data reuse are widely acknowledged, the data collections existing today are still kept in silos. Indeed, monitoring what happens to data once they have been handed to a third party is currently not feasible within the current data sharing practices. We propose a blockchain-based system to trace data collections and potentially create a more trustworthy data sharing process. In this paper, we present the LUCE (License accoUntability and CompliancE) architecture as a decentralized blockchain-based platform supporting data sharing and reuse. LUCE is designed to provide full transparency on what happens to the data after they are shared with third parties. The contributions of this work consist of i) the design of a decentralized data sharing solution with accountability and compliance by design and ii) the inclusion of a dynamic consent model for personalized data sharing preferences and for enabling legal compliance mechanisms. We test the scalability of the platform in a real-time environment where a growing number of users access and reuse different datasets. Compared to existing data sharing solutions, LUCE provides transparency over data sharing practices, enables data reuse, and supports regulatory requirements. The experimentation shows that the platform can be scaled for a large number of users.
Visara Urovi, Vikas Jaiman, Arno Angerer, Michel Dumontier
Blockchain Res. Appl.4
2022 Studying the association of diabetes and healthcare cost on distributed data from the Maastricht Study and Statistics Netherlands using a privacy-preserving federated learning infrastructure
abstract
The mining of personal data collected by multiple organizations remains challenging in the presence of technical barriers, privacy concerns, and legal and/or organizational restrictions. While a number of privacy-preserving and data mining frameworks have recently emerged, much remains to show their practical utility. In this study, we implement and utilize a secure infrastructure using data from Statistics Netherlands and the Maastricht Study to learn the association between Type 2 Diabetes Mellitus (T2DM) and healthcare expenses considering the impact of lifestyle, physical activities, and complications of T2DM. Through experiments using real-world distributed personal data, we present the feasibility and effectiveness of the secure infrastructure for practical use cases of linking and analyzing vertically partitioned data across multiple organizations. We discovered that individuals diagnosed with T2DM had significantly higher expenses than those with prediabetes, while participants with prediabetes spent more than those without T2DM in all the included healthcare categories to different degrees. We further discuss a joint effort from technical, ethical-legal, and domain-specific experts that is highly valued for applying such a secure infrastructure to real-life use cases to protect data privacy.
Chang Sun 0001, Johan van Soest, Annemarie Koster, Simone J. P. M. Eussen, Miranda T. Schram, Coen D. A. Stehouwer, Pieter C. Dagnelie, Michel Dumontier
J. Biomed. Informatics8
2021 User-friendly Composition of FAIR Workflows in a Notebook Environment
abstract
There has been a large focus in recent years on making assets in scientific research findable, accessible, interoperable and reusable, collectively known as the FAIR principles. A particular area of focus lies in applying these principles to scientific computational workflows. Jupyter notebooks are a very popular medium by which to program and communicate computational scientific analyses. However, they present unique challenges when it comes to reuse of only particular steps of an analysis without disrupting the usual flow and benefits of the notebook approach, making it difficult to fully comply with the FAIR principles. Here we present an approach and toolset for adding the power of semantic technologies to Python-encoded scientific workflows in a simple, automated and minimally intrusive manner. The semantic descriptions are published as a series of nanopublications that can be searched and used in other notebooks by means of a Jupyter Lab plugin. We describe the implementation of the proposed approach and toolset, and provide the results of a user study with 15 participants, designed around image processing workflows, to evaluate the usability of the system and its perceived effect on FAIRness. Our results show that our approach is feasible and perceived as user-friendly. Our system received an overall score of 78.75 on the System Usability Scale, which is above the average score reported in the literature.
Robin A. Richardson, Remzi Çelebi, Sven van der Burg, Djura Smits, Lars Ridder, Michel Dumontier, Tobias Kuhn
K-CAP6
2021 Relation extraction from DailyMed structured product labels by optimally combining crowd, experts and machines
abstract
The effectiveness of machine learning models to provide accurate and consistent results in drug discovery and clinical decision support is strongly dependent on the quality of the data used. However, substantive amounts of open data that drive drug discovery suffer from a number of issues including inconsistent representation, inaccurate reporting, and incomplete context. For example, databases of FDA-approved drug indications used in computational drug repositioning studies do not distinguish between treatments that simply offer symptomatic relief from those that target the underlying pathology. Moreover, drug indication sources often lack proper provenance and have little overlap. Consequently, new predictions can be of poor quality as they offer little in the way of new insights. Hence, work remains to be done to establish higher quality databases of drug indications that are suitable for use in drug discovery and repositioning studies. Here, we report on the combination of weak supervision (i.e., programmatic labeling and crowdsourcing) and deep learning methods for relation extraction from DailyMed text to create a higher quality drug-disease relation dataset. The generated drug-disease relation data shows a high overlap with DrugCentral, a manually curated dataset. Using this dataset, we constructed a machine learning model to classify relations between drugs and diseases from text into four categories; treatment, symptomatic relief, contradiction, and effect, exhibiting an improvement of 15.5% with Bi-LSTM (F1 score of 71.8%) over the best performing discrete method. Access to high quality data is crucial to building accurate and reliable drug repurposing prediction models. Our work suggests how the combination of crowds, experts, and machine learning methods can go hand-in-hand to improve datasets and predictive models.
Krist Shingjergji, Remzi Çelebi, Jan C. Scholtes, Michel Dumontier
J. Biomed. Informatics4
2020 Accelerating Discovery Science with an Internet of FAIR Data and Services
abstract
Biomedicine has always been a fertile and challenging domain for computational discovery science. Indeed, the existence of millions of scientific articles, thousands of databases, and hundreds of ontologies offer exciting opportunities to mine our collective knowledge, were we not stymied by incompatible formats, incomplete and overlapping vocabularies, confusing licensing policies, and heterogeneous data access points.
Michel Dumontier
CIKM1
2020 Sleeping Beauties in Case Law
abstract
A challenge in computational legal research is the quantitative assessment of “relevance” in a network of court decisions. The term “sleeping beauty” (SB) was coined to denote an article that received almost no attention immediately after publication, but suddenly received multiple citations many years later. These publications can be identified by calculating their Beauty coefficient (B-coefficient). In this contribution, we apply approaches used for identifying SBs to decisions arising from the Court of Justice of the European Union (CJEU). We compared B-coefficients of CJEU cases with their centrality scores from classical algorithms from network analysis, finding that these measures tend to correlate. We discuss the implications of this that are interesting for legal scholars, acknowledging that future work is required to calibrate the scale of the time variable in the B-coefficient formula for finer-grained application to case law. Our study’s setup provides a foundation for new case law analytics methodologies that extends the power of traditional network analysis techniques for answering questions about the behavior of European courts.
Pedro Hernandez Serrano, Kodylan Moodley, Gijs van Dijck, Michel Dumontier
JURIX4
2020 Ten simple rules for making training materials FAIR
abstract
Everything we do today is becoming more and more reliant on the use of computers. The field of biology is no exception; but most biologists receive little or no formal preparation for the increasingly computational aspects of their discipline. In consequence, informal training courses are often needed to plug the gaps; and the demand for such training is growing worldwide. To meet this demand, some training programs are being expanded, and new ones are being developed. Key to both scenarios is the creation of new course materials. Rather than starting from scratch, however, it's sometimes possible to repurpose materials that already exist. Yet finding suitable materials online can be difficult: They're often widely scattered across the internet or hidden in their home institutions, with no systematic way to find them. This is a common problem for all digital objects. The scientific community has attempted to address this issue by developing a set of rules (which have been called the Findable, Accessible, Interoperable and Reusable [FAIR] principles) to make such objects more findable and reusable. Here, we show how to apply these rules to help make training materials easier to find, (re)use, and adapt, for the benefit of all.
Leyla Jael Castro, Bérénice Batut, Melissa L. Burke, Mateusz Kuzak, Fotis E. Psomopoulos, Ricardo Arcila, Terri K. Attwood, Niall Beard, Denise Carvalho-Silva, Alexandros C. Dimopoulos, Victoria Dominguez Del Angel, Michel Dumontier, Kim T. Gurwitz, Roland Krause, Peter McQuilton, Loredana Le Pera, Sarah L. Morgan, Päivi Rauste, Allegra Via, Pascal Kahlem, Gabriella Rustici, Celia W. G. van Gelder, Patricia M. Palagi
PLoS Comput. Biol.12
2019 Finding the evidence base using citation networks: Do 300 to 400 U.S. physicians die by suicide annually?
Tiffany I. Leung, Sima Pendharkar, Chwen-Yuen Angie Chen, Michel Dumontier
AMIA4
2019 Similarity and Relevance of Court Decisions: A Computational Study on CJEU Cases
abstract
Identification of relevant or similar court decisions is a core activity in legal decision making for case law researchers and practitioners. With an ever increasing body of case law, a manual analysis of court decisions can become practically impossible. As a result, some decisions are inevitably overlooked. Alternatively, network analysis may be applied to detect relevant precedents and landmark cases. Previous research suggests that citation networks of court decisions frequently provide relevant precedents and landmark cases. The advent of text similarity measures (both syntactic and semantic) has meant that potentially relevant cases can be identified without the need to manually read them. However, how close do these measures come to approximating the notion of relevance captured in the citation network? In this contribution, we explore this question by measuring the level of agreement of state-of-the-art text similarity algorithms with the citation behavior in the case citation network. For this paper, we focus on judgements by the Court of Justice of the European Union (CJEU) as published in the EUR-Lex database. Our results show that similarity of the full texts of CJEU court decisions does not closely mirror citation behaviour, there is a substantial overlap. In particular, we found syntactic measures surprisingly outperform semantic ones in approximating the citation network.
Kodylan Moodley, Pedro Hernandez Serrano, Gijs van Dijck, Michel Dumontier
JURIX4
2019 Evaluation of knowledge graph embedding approaches for drug-drug interaction prediction in realistic settings
abstract
BACKGROUND: Current approaches to identifying drug-drug interactions (DDIs), include safety studies during drug development and post-marketing surveillance after approval, offer important opportunities to identify potential safety issues, but are unable to provide complete set of all possible DDIs. Thus, the drug discovery researchers and healthcare professionals might not be fully aware of potentially dangerous DDIs. Predicting potential drug-drug interaction helps reduce unanticipated drug interactions and drug development costs and optimizes the drug design process. Methods for prediction of DDIs have the tendency to report high accuracy but still have little impact on translational research due to systematic biases induced by networked/paired data. In this work, we aimed to present realistic evaluation settings to predict DDIs using knowledge graph embeddings. We propose a simple disjoint cross-validation scheme to evaluate drug-drug interaction predictions for the scenarios where the drugs have no known DDIs. RESULTS: We designed different evaluation settings to accurately assess the performance for predicting DDIs. The settings for disjoint cross-validation produced lower performance scores, as expected, but still were good at predicting the drug interactions. We have applied Logistic Regression, Naive Bayes and Random Forest on DrugBank knowledge graph with the 10-fold traditional cross validation using RDF2Vec, TransE and TransD. RDF2Vec with Skip-Gram generally surpasses other embedding methods. We also tested RDF2Vec on various drug knowledge graphs such as DrugBank, PharmGKB and KEGG to predict unknown drug-drug interactions. The performance was not enhanced significantly when an integrated knowledge graph including these three datasets was used. CONCLUSION: We showed that the knowledge embeddings are powerful predictors and comparable to current state-of-the-art methods for inferring new DDIs. We addressed the evaluation biases by introducing drug-wise and pairwise disjoint test classes. Although the performance scores for drug-wise and pairwise disjoint seem to be low, the results can be considered to be realistic in predicting the interactions for drugs with limited interaction information.
Remzi Çelebi, Hüseyin Uyar, Erkan Yasar, Özgür Gümüs, Oguz Dikenelli, Michel Dumontier
BMC Bioinform.6
2018 Nanopublications: A Growing Resource of Provenance-Centric Scientific Linked Data
abstract
Nanopublications are a Linked Data format for scholarly data publishing that has received considerable uptake in the last few years. In contrast to the common Linked Data publishing practice, nanopublications work at the granular level of atomic information snippets and provide a consistent container format to attach provenance and metadata at this atomic level. While the nanopublications format is domain-independent, the datasets that have become available in this format are mostly from Life Science domains, including data about diseases, genes, proteins, drugs, biological pathways, and biotic interactions. More than 10 million such nanopublications have been published, which now form a valuable resource for studies on the domain level of the given Life Science domains as well as on the more technical levels of provenance modeling and heterogeneous Linked Data. We provide here an overview of this combined nanopublication dataset, show the results of some overarching analyses, and describe how it can be accessed and queried.
Tobias Kuhn, Albert Meroño-Peñuela, Alexander Malic, Jorrit H. Poelen, Allen H. Hurlbert, Emilio Centeno, Laura Inés Furlong, Núria Queralt-Rosinach, Christine Chichester, Juan M. Banda, Egon L. Willighagen, Friederike Ehrhart, Chris T. A. Evelo, Tareq B. Malas, Michel Dumontier
eScience15
2018 Drug-drug interaction extraction via hierarchical RNNs on sequence and shortest dependency paths
abstract
Motivation: Adverse events resulting from drug-drug interactions (DDI) pose a serious health issue. The ability to automatically extract DDIs described in the biomedical literature could further efforts for ongoing pharmacovigilance. Most of neural networks-based methods typically focus on sentence sequence to identify these DDIs, however the shortest dependency path (SDP) between the two entities contains valuable syntactic and semantic information. Effectively exploiting such information may improve DDI extraction. Results: In this article, we present a hierarchical recurrent neural networks (RNNs)-based method to integrate the SDP and sentence sequence for DDI extraction task. Firstly, the sentence sequence is divided into three subsequences. Then, the bottom RNNs model is employed to learn the feature representation of the subsequences and SDP, and the top RNNs model is employed to learn the feature representation of both sentence sequence and SDP. Furthermore, we introduce the embedding attention mechanism to identify and enhance keywords for the DDI extraction task. We evaluate our approach using the DDI extraction 2013 corpus. Our method is competitive or superior in performance as compared with other state-of-the-art methods. Experimental results show that the sentence sequence and SDP are complementary to each other. Integrating the sentence sequence with SDP can effectively improve the DDI extraction performance. Availability and implementation: The experimental data is available at https://github.com/zhangyijia1979/hierarchical-RNNs-model-for-DDI-extraction. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Yi-Jia Zhang 0001, Wei Zheng 0003, Hongfei Lin, Jian Wang 0021, Michel Dumontier
Bioinform.6
2017 Is Crowdsourcing Patient-Reported Outcomes the Future of Evidence-Based Medicine? A Case Study of Back Pain
Mor Peleg, Tiffany I. Leung, Manisha Desai, Michel Dumontier
AIME4
2017 smartAPI: Towards a More Intelligent Network of Web APIs
Amrapali Zaveri, Shima Dastgheib, Chunlei Wu, Trish Whetzel, Ruben Verborgh, Paul Avillach, Gabor Korodi, Raymond Terryn, Kathleen M. Jagodnik, Pedro Assis, Michel Dumontier
ESWC (2)11
2017 Cleaning by clustering: methodology for addressing data quality issues in biomedical metadata
abstract
BACKGROUND: The ability to efficiently search and filter datasets depends on access to high quality metadata. While most biomedical repositories require data submitters to provide a minimal set of metadata, some such as the Gene Expression Omnibus (GEO) allows users to specify additional metadata in the form of textual key-value pairs (e.g. sex: female). However, since there is no structured vocabulary to guide the submitter regarding the metadata terms to use, consequently, the 44,000,000+ key-value pairs in GEO suffer from numerous quality issues including redundancy, heterogeneity, inconsistency, and incompleteness. Such issues hinder the ability of scientists to hone in on datasets that meet their requirements and point to a need for accurate, structured and complete description of the data. METHODS: In this study, we propose a clustering-based approach to address data quality issues in biomedical, specifically gene expression, metadata. First, we present three different kinds of similarity measures to compare metadata keys. Second, we design a scalable agglomerative clustering algorithm to cluster similar keys together. RESULTS: Our agglomerative cluster algorithm identified metadata keys that were similar, based on (i) name, (ii) core concept and (iii) value similarities, to each other and grouped them together. We evaluated our method using a manually created gold standard in which 359 keys were grouped into 27 clusters based on six types of characteristics: (i) age, (ii) cell line, (iii) disease, (iv) strain, (v) tissue and (vi) treatment. As a result, the algorithm generated 18 clusters containing 355 keys (four clusters with only one key were excluded). In the 18 clusters, there were keys that were identified correctly to be related to that cluster, but there were 13 keys which were not related to that cluster. We compared our approach with four other published methods. Our approach significantly outperformed them for most metadata keys and achieved the best average F-Score (0.63). CONCLUSION: Our algorithm identified keys that were similar to each other and grouped them together. Our intuition that underpins cleaning by clustering is that, dividing keys into different clusters resolves the scalability issues for data observation and cleaning, and keys in the same cluster with duplicates and errors can easily be found. Our algorithm can also be applied to other biomedical data types.
Wei Hu 0007, Amrapali Zaveri, Honglei Qiu, Michel Dumontier
BMC Bioinform.4
2017 Formalizing drug indications on the road to therapeutic intent
abstract
Therapeutic intent, the reason behind the choice of a therapy and the context in which a given approach should be used, is an important aspect of medical practice. There are unmet needs with respect to current electronic mapping of drug indications. For example, the active ingredient sildenafil has 2 distinct indications, which differ solely on dosage strength. In progressing toward a practice of precision medicine, there is a need to capture and structure therapeutic intent for computational reuse, thus enabling more sophisticated decision-support tools and a possible mechanism for computer-aided drug repurposing. The indications for drugs, such as those expressed in the Structured Product Labels approved by the US Food and Drug Administration, appears to be a tractable area for developing an application ontology of therapeutic intent.
Stuart J. Nelson, Tudor I. Oprea, Oleg Ursu, Cristian Bologa, Amrapali Zaveri, Jayme Holmes, Jeremy J. Yang, Stephen L. Mathias, Subramani Mani, Mark S. Tuttle, Michel Dumontier
J. Am. Medical Informatics Assoc.11
2017 Drug-drug interaction discovery and demystification using Semantic Web technologies
abstract
OBJECTIVE: To develop a novel pharmacovigilance inferential framework to infer mechanistic explanations for asserted drug-drug interactions (DDIs) and deduce potential DDIs. MATERIALS AND METHODS: A mechanism-based DDI knowledge base was constructed by integrating knowledge from several existing sources at the pharmacokinetic, pharmacodynamic, pharmacogenetic, and multipathway interaction levels. A query-based framework was then created to utilize this integrated knowledge base in conjunction with 9 inference rules to infer mechanistic explanations for asserted DDIs and deduce potential DDIs. RESULTS: The drug-drug interactions discovery and demystification (D3) system achieved an overall 85% recall rate in terms of inferring mechanistic explanations for the DDIs integrated into its knowledge base, while demonstrating a 61% precision rate in terms of the inference or lack of inference of mechanistic explanations for a balanced, randomly selected collection of interacting and noninteracting drug pairs. DISCUSSION: The successful demonstration of the D3 system's ability to confirm interactions involving well-studied drugs enhances confidence in its ability to deduce interactions involving less-studied drugs. In its demonstration, the D3 system infers putative explanations for most of its integrated DDIs. Further enhancements to this work in the future might include ranking interaction mechanisms based on likelihood of applicability, determining the likelihood of deduced DDIs, and making the framework publicly available. CONCLUSION: The D3 system provides an early-warning framework for augmenting knowledge of known DDIs and deducing unknown DDIs. It shows promise in suggesting interaction pathways of research and evaluation interest and aiding clinicians in evaluating and adjusting courses of drug therapy.
Adeeb Noor, Abdullah Assiri, Serkan Ayvaz, Connor Clark, Michel Dumontier
J. Am. Medical Informatics Assoc.5
2017 Developing a framework for digital objects in the Big Data to Knowledge (BD2K) commons: Report from the Commons Framework Pilots workshop
Kathleen M. Jagodnik, Simon Koplev, Sherry L. Jenkins, Lucila Ohno-Machado, Benedict Paten, Stephan C. Schürer, Michel Dumontier, Ruben Verborgh, Alex Bui, Peipei Ping, Neil J. McKenna, Ravi K. Madduri, Ajay Pillai, Avi Ma'ayan
J. Biomed. Informatics7
2017 Predicting biomedical metadata in CEDAR: A study of Gene Expression Omnibus (GEO)
abstract
A crucial and limiting factor in data reuse is the lack of accurate, structured, and complete descriptions of data, known as metadata. Towards improving the quantity and quality of metadata, we propose a novel metadata prediction framework to learn associations from existing metadata that can be used to predict metadata values. We evaluate our framework in the context of experimental metadata from the Gene Expression Omnibus (GEO). We applied four rule mining algorithms to the most common structured metadata elements (sample type, molecular type, platform, label type and organism) from over 1.3million GEO records. We examined the quality of well supported rules from each algorithm and visualized the dependencies among metadata elements. Finally, we evaluated the performance of the algorithms in terms of accuracy, precision, recall, and F-measure. We found that PART is the best algorithm outperforming Apriori, Predictive Apriori, and Decision Table. All algorithms perform significantly better in predicting class values than the majority vote classifier. We found that the performance of the algorithms is related to the dimensionality of the GEO elements. The average performance of all algorithm increases due of the decreasing of dimensionality of the unique values of these elements (2697 platforms, 537 organisms, 454 labels, 9 molecules, and 5 types). Our work suggests that experimental metadata such as present in GEO can be accurately predicted using rule mining algorithms. Our work has implications for both prospective and retrospective augmentation of metadata quality, which are geared towards making data easier to find and reuse.
Maryam Panahiazar, Michel Dumontier, Olivier Gevaert
J. Biomed. Informatics2
2016 The digital revolution in phenotyping
abstract
Phenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data.
Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier
Briefings Bioinform.19
2016 Is the crowd better as an assistant or a replacement in ontology engineering? An exploration through the lens of the Gene Ontology
Jonathan Mortensen, Natalie Telis, Jacob J. Hughey, Hua Fan-Minogue, Kimberly Van Auken, Michel Dumontier, Mark A. Musen
J. Biomed. Informatics6
2015 Drug-Disease Associations in Guidelines, Drug Labels, and Practice
Tiffany I. Leung, Michel Dumontier
AMIA2
2015 Provenance-Centered Dataset of Drug-Drug Interactions
Juan M. Banda, Tobias Kuhn, Nigam H. Shah, Michel Dumontier
ISWC (2)4
2015 Link Analysis of Life Science Linked Data
Wei Hu 0007, Honglei Qiu, Michel Dumontier
ISWC (2)3
2015 Publishing Without Publishers: A Decentralized Approach to Dissemination, Retrieval, and Archiving of Data
abstract
Making available and archiving scientific results is for the most part still considered the task of classical publishing companies, despite the fact that classical forms of publishing centered around printed narrative articles no longer seem well-suited in the digital age. In particular, there exist currently no efficient, reliable, and agreed-upon methods for publishing scientific datasets, which have become increasingly important for science. Here we propose to design scientific data publishing as a Web-based bottom-up process, without top-down control of central authorities such as publishing companies. Based on a novel combination of existing concepts and technologies, we present a server network to decentrally store and archive data in the form of nanopublications, an RDF-based format to represent scientific data. We show how this approach allows researchers to publish, retrieve, verify, and recombine datasets of nanopublications in a reliable and trustworthy manner, and we argue that this architecture could be used for the Semantic Web in general. Evaluation of the current small network shows that this system is efficient and reliable. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Tobias Kuhn, Christine Chichester, Michael Krauthammer, Michel Dumontier
ISWC (1)4
2015 SPARQL-enabled identifier conversion with Identifiers.org
abstract
MOTIVATION: On the semantic web, in life sciences in particular, data is often distributed via multiple resources. Each of these sources is likely to use their own International Resource Identifier for conceptually the same resource or database record. The lack of correspondence between identifiers introduces a barrier when executing federated SPARQL queries across life science data. RESULTS: We introduce a novel SPARQL-based service to enable on-the-fly integration of life science data. This service uses the identifier patterns defined in the Identifiers.org Registry to generate a plurality of identifier variants, which can then be used to match source identifiers with target identifiers. We demonstrate the utility of this identifier integration approach by answering queries across major producers of life science Linked Data. AVAILABILITY AND IMPLEMENTATION: The SPARQL-based identifier conversion service is available without restriction at http://identifiers.org/services/sparql.
Sarala M. Wimalaratne, Jerven T. Bolleman, Nick S. Juty, Toshiaki Katayama, Michel Dumontier, Nicole Redaschi, Nicolas Le Novère, Henning Hermjakob, Camille Laibe
Bioinform.5
2015 An evidence-based approach to identify aging-related genes in Caenorhabditis elegans
abstract
BACKGROUND: Extensive studies have been carried out on Caenorhabditis elegans as a model organism to elucidate mechanisms of aging and the effects of perturbing known aging-related genes on lifespan and behavior. This research has generated large amounts of experimental data that is increasingly difficult to integrate and analyze with existing databases and domain knowledge. To address this challenge, we demonstrate a scalable and effective approach for automatic evidence gathering and evaluation that leverages existing experimental data and literature-curated facts to identify genes involved in aging and lifespan regulation in C. elegans. RESULTS: We developed a semantic knowledge base for aging by integrating data about C. elegans genes from WormBase with data about 2005 human and model organism genes from GenAge and 149 genes from GenDR, and with the Bio2RDF network of linked data for the life sciences. Using HyQue (a Semantic Web tool for hypothesis-based querying and evaluation) to interrogate this knowledge base, we examined 48,231 C. elegans genes for their role in modulating lifespan and aging. HyQue identified 24 novel but well-supported candidate aging-related genes for further experimental validation. CONCLUSIONS: We use semantic technologies to discover candidate aging genes whose effects on lifespan are not yet well understood. Our customized HyQue system, the aging research knowledge base it operates over, and HyQue evaluations of all C. elegans genes are freely available at http://hyque.semanticscience.org .
Alison Callahan, Juan Cifuentes, Michel Dumontier
BMC Bioinform.3
2015 The center for expanded data annotation and retrieval
abstract
The Center for Expanded Data Annotation and Retrieval is studying the creation of comprehensive and expressive metadata for biomedical datasets to facilitate data discovery, data interpretation, and data reuse. We take advantage of emerging community-based standard templates for describing different kinds of biomedical datasets, and we investigate the use of computational techniques to help investigators to assemble templates and to fill in their values. We are creating a repository of metadata from which we plan to identify metadata patterns that will drive predictive data entry when filling in metadata templates. The metadata repository not only will capture annotations specified when experimental datasets are initially created, but also will incorporate links to the published literature, including secondary analyses and possible refinements or retractions of experimental interpretations. By working initially with the Human Immunology Project Consortium and the developers of the ImmPort data repository, we are developing and evaluating an end-to-end solution to the problems of metadata authoring and management that will generalize to other data-management environments.
Mark A. Musen, Carol A. Bean, Kei-Hoi Cheung, Michel Dumontier, Kim A. Durante, Olivier Gevaert, Alejandra N. González-Beltrán, Purvesh Khatri, Steven H. Kleinstein, Martin J. O'Connor, Yannick Pouliot, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jeffrey A. Wiser
J. Am. Medical Informatics Assoc.4
2015 Toward a complete dataset of drug-drug interaction information from publicly available sources
abstract
Although potential drug-drug interactions (PDDIs) are a significant source of preventable drug-related harm, there is currently no single complete source of PDDI information. In the current study, all publically available sources of PDDI information that could be identified using a comprehensive and broad search were combined into a single dataset. The combined dataset merged fourteen different sources including 5 clinically-oriented information sources, 4 Natural Language Processing (NLP) Corpora, and 5 Bioinformatics/Pharmacovigilance information sources. As a comprehensive PDDI source, the merged dataset might benefit the pharmacovigilance text mining community by making it possible to compare the representativeness of NLP corpora for PDDI text extraction tasks, and specifying elements that can be useful for future PDDI extraction purposes. An analysis of the overlap between and across the data sources showed that there was little overlap. Even comprehensive PDDI lists such as DrugBank, KEGG, and the NDF-RT had less than 50% overlap with each other. Moreover, all of the comprehensive lists had incomplete coverage of two data sources that focus on PDDIs of interest in most clinical settings. Based on this information, we think that systems that provide access to the comprehensive lists, such as APIs into RxNorm, should be careful to inform users that the lists may be incomplete with respect to PDDIs that drug experts suggest clinicians be aware of. In spite of the low degree of overlap, several dozen cases were identified where PDDI information provided in drug product labeling might be augmented by the merged dataset. Moreover, the combined dataset was also shown to improve the performance of an existing PDDI NLP pipeline and a recently published PDDI pharmacovigilance protocol. Future work will focus on improvement of the methods for mapping between PDDI information sources, identifying methods to improve the use of the merged dataset in PDDI NLP algorithms, integrating high-quality PDDI information from the merged dataset into Wikidata, and making the combined dataset accessible as Semantic Web Linked Data.
Serkan Ayvaz, John R. Horn, Oktie Hassanzadeh, Qian Zhu 0003, Johann Stan, Nicholas P. Tatonetti, Santiago Vilar, Mathias Brochhausen, Matthias Samwald, Majid Rastegar-Mojarad, Michel Dumontier, Richard D. Boyce
J. Biomed. Informatics11
2015 Making Digital Artifacts on the Web Verifiable and Reliable
abstract
The current Web has no general mechanisms to make digital artifacts-such as datasets, code, texts, and images-verifiable and permanent. For digital artifacts that are supposed to be immutable, there is moreover no commonly accepted method to enforce this immutability. These shortcomings have a serious negative impact on the ability to reproduce the results of processes that rely on Web resources, which in turn heavily impacts areas such as science where reproducibility is important. To solve this problem, we propose trusty URIs containing cryptographic hash values. We show how trusty URIs can be used for the verification of digital artifacts, in a manner that is independent of the serialization format in the case of structured data files such as nanopublications. We demonstrate how the contents of these files become immutable, including dependencies to external digital artifacts and thereby extending the range of verifiability to the entire reference tree. Our approach sticks to the core principles of the Web, namely openness and decentralized architecture, and is fully compatible with existing standards and protocols. Evaluation of our reference implementations shows that these design goals are indeed accomplished by our approach, and that it remains practical even for very large files.
Tobias Kuhn, Michel Dumontier
IEEE Trans. Knowl. Data Eng.2
2014 Automating Identification of Multiple Chronic Conditions in Clinical Practice Guidelines
Tiffany I. Leung, Hawre Jalal, Donna M. Zulman, Douglas K. Owens, Mark A. Musen, Michel Dumontier, Mary K. Goldstein
AMIA6
2014 Trusty URIs: Verifiable, Immutable, and Permanent Digital Artifacts for Linked Data
Tobias Kuhn, Michel Dumontier
ESWC2
2014 Erratum: Trusty URIs: Verifiable, Immutable, and Permanent Digital Artifacts for Linked Data
Tobias Kuhn, Michel Dumontier
ESWC2
2014 Mouse model phenotypes provide information about human drug targets
abstract
MOTIVATION: Methods for computational drug target identification use information from diverse information sources to predict or prioritize drug targets for known drugs. One set of resources that has been relatively neglected for drug repurposing is animal model phenotype. RESULTS: We investigate the use of mouse model phenotypes for drug target identification. To achieve this goal, we first integrate mouse model phenotypes and drug effects, and then systematically compare the phenotypic similarity between mouse models and drug effect profiles. We find a high similarity between phenotypes resulting from loss-of-function mutations and drug effects resulting from the inhibition of a protein through a drug action, and demonstrate how this approach can be used to suggest candidate drug targets. AVAILABILITY AND IMPLEMENTATION: Analysis code and supplementary data files are available on the project Web site at https://drugeffects.googlecode.com.
Robert Hoehndorf, Tanya Hiebert, Nigel W. Hardy, Paul N. Schofield, Georgios V. Gkoutos, Michel Dumontier
Bioinform.6
2013 Bio2RDF Release 2: Improved Coverage, Interoperability and Provenance of Life Science Linked Data
Alison Callahan, Jose Cruz-Toledo, Peter Ansell, Michel Dumontier
ESWC4
2013 Evaluation of research in biomedical ontologies
abstract
Ontologies are now pervasive in biomedicine, where they serve as a means to standardize terminology, to enable access to domain knowledge, to verify data consistency and to facilitate integrative analyses over heterogeneous biomedical data. For this purpose, research on biomedical ontologies applies theories and methods from diverse disciplines such as information management, knowledge representation, cognitive science, linguistics and philosophy. Depending on the desired applications in which ontologies are being applied, the evaluation of research in biomedical ontologies must follow different strategies. Here, we provide a classification of research problems in which ontologies are being applied, focusing on the use of ontologies in basic and translational research, and we demonstrate how research results in biomedical ontologies can be evaluated. The evaluation strategies depend on the desired application and measure the success of using an ontology for a particular biomedical problem. For many applications, the success can be quantified, thereby facilitating the objective evaluation and comparison of research in biomedical ontology. The objective, quantifiable comparison of research results based on scientific applications opens up the possibility for systematically improving the utility of ontologies in biomedical research.
Robert Hoehndorf, Michel Dumontier, Georgios V. Gkoutos
Briefings Bioinform.2
2013 Evaluation of the OQuaRE framework for ontology quality
Astrid Duque-Ramos, Jesualdo Tomás Fernández-Breis, Miguela Iniesta-Moreno, Michel Dumontier, Mikel Egaña Aranguren, Stefan Schulz 0001, Nathalie Aussenac-Gilles, Robert Stevens 0001
Expert Syst. Appl.4
2013 State of the art and open challenges in community-driven knowledge curation
Tudor Groza, Tania Tudorache, Michel Dumontier
J. Biomed. Informatics3
2012 Evaluating Scientific Hypotheses Using the SPARQL Inferencing Notation
Alison Callahan, Michel Dumontier
ESWC2
2012 Building an HIV data mashup using Bio2RDF
abstract
We present an update to the Bio2RDF Linked Data Network, which now comprises ∼30 billion statements across 30 data sets. Significant changes to the framework include the accommodation of global mirrors, offline data processing and new search and integration services. The utility of this new network of knowledge is illustrated through a Bio2RDF-based mashup with microarray gene expression results and interaction data obtained from the HIV-1, Human Protein Interaction Database (HHPID) with respect to the infection of human macrophages with the human immunodeficiency virus type 1 (HIV-1).
Marc-Alexandre Nolin, Michel Dumontier, François Belleau, Jacques Corbeil
Briefings Bioinform.2
2012 Identifying aberrant pathways through integrated analysis of knowledge in pharmacogenomics
abstract
MOTIVATION: Many complex diseases are the result of abnormal pathway functions instead of single abnormalities. Disease diagnosis and intervention strategies must target these pathways while minimizing the interference with normal physiological processes. Large-scale identification of disease pathways and chemicals that may be used to perturb them requires the integration of information about drugs, genes, diseases and pathways. This information is currently distributed over several pharmacogenomics databases. An integrated analysis of the information in these databases can reveal disease pathways and facilitate novel biomedical analyses. RESULTS: We demonstrate how to integrate pharmacogenomics databases through integration of the biomedical ontologies that are used as meta-data in these databases. The additional background knowledge in these ontologies can then be used to enable novel analyses. We identify disease pathways using a novel multi-ontology enrichment analysis over the Human Disease Ontology, and we identify significant associations between chemicals and pathways using an enrichment analysis over a chemical ontology. The drug-pathway and disease-pathway associations are a valuable resource for research in disease and drug mechanisms and can be used to improve computational drug repurposing. AVAILABILITY: http://pharmgkb-owl.googlecode.com
Robert Hoehndorf, Michel Dumontier, Georgios V. Gkoutos
Bioinform.2
2012 Self-organizing ontology of biochemically relevant small molecules
abstract
BACKGROUND: The advent of high-throughput experimentation in biochemistry has led to the generation of vast amounts of chemical data, necessitating the development of novel analysis, characterization, and cataloguing techniques and tools. Recently, a movement to publically release such data has advanced biochemical structure-activity relationship research, while providing new challenges, the biggest being the curation, annotation, and classification of this information to facilitate useful biochemical pattern analysis. Unfortunately, the human resources currently employed by the organizations supporting these efforts (e.g. ChEBI) are expanding linearly, while new useful scientific information is being released in a seemingly exponential fashion. Compounding this, currently existing chemical classification and annotation systems are not amenable to automated classification, formal and transparent chemical class definition axiomatization, facile class redefinition, or novel class integration, thus further limiting chemical ontology growth by necessitating human involvement in curation. Clearly, there is a need for the automation of this process, especially for novel chemical entities of biological interest. RESULTS: To address this, we present a formal framework based on Semantic Web technologies for the automatic design of chemical ontology which can be used for automated classification of novel entities. We demonstrate the automatic self-assembly of a structure-based chemical ontology based on 60 MeSH and 40 ChEBI chemical classes. This ontology is then used to classify 200 compounds with an accuracy of 92.7%. We extend these structure-based classes with molecular feature information and demonstrate the utility of our framework for classification of functionally relevant chemicals. Finally, we discuss an iterative approach that we envision for future biochemical ontology development. CONCLUSIONS: We conclude that the proposed methodology can ease the burden of chemical data annotators and dramatically increase their productivity. We anticipate that the use of formal logic in our proposed framework will make chemical classification criteria more transparent to humans and machines alike and will thus facilitate predictive and integrative bioactivity model development.
Leonid L. Chepelev, Janna Hastings, Marcus Ennis, Christoph Steinbeck, Michel Dumontier
BMC Bioinform.5
2011 A common layer of interoperability for biomedical ontologies based on OWL EL
abstract
MOTIVATION: Ontologies are essential in biomedical research due to their ability to semantically integrate content from different scientific databases and resources. Their application improves capabilities for querying and mining biological knowledge. An increasing number of ontologies is being developed for this purpose, and considerable effort is invested into formally defining them in order to represent their semantics explicitly. However, current biomedical ontologies do not facilitate data integration and interoperability yet, since reasoning over these ontologies is very complex and cannot be performed efficiently or is even impossible. We propose the use of less expressive subsets of ontology representation languages to enable efficient reasoning and achieve the goal of genuine interoperability between ontologies. RESULTS: We present and evaluate EL Vira, a framework that transforms OWL ontologies into the OWL EL subset, thereby enabling the use of tractable reasoning. We illustrate which OWL constructs and inferences are kept and lost following the conversion and demonstrate the performance gain of reasoning indicated by the significant reduction of processing time. We applied EL Vira to the open biomedical ontologies and provide a repository of ontologies resulting from this conversion. EL Vira creates a common layer of ontological interoperability that, for the first time, enables the creation of software solutions that can employ biomedical ontologies to perform inferences and answer complex queries to support scientific analyses. AVAILABILITY AND IMPLEMENTATION: The EL Vira software is available from http://el-vira.googlecode.com and converted OBO ontologies and their mappings are available from http://bioonto.gen.cam.ac.uk/el-ont.
Robert Hoehndorf, Michel Dumontier, Anika Oellrich, Sarala M. Wimalaratne, Dietrich Rebholz-Schuhmann, Paul N. Schofield, Georgios V. Gkoutos
Bioinform.2
2011 Prototype Semantic Infrastructure for Automated Small Molecule Classification and Annotation in Lipidomics
abstract
BACKGROUND: The development of high-throughput experimentation has led to astronomical growth in biologically relevant lipids and lipid derivatives identified, screened, and deposited in numerous online databases. Unfortunately, efforts to annotate, classify, and analyze these chemical entities have largely remained in the hands of human curators using manual or semi-automated protocols, leaving many novel entities unclassified. Since chemical function is often closely linked to structure, accurate structure-based classification and annotation of chemical entities is imperative to understanding their functionality. RESULTS: As part of an exploratory study, we have investigated the utility of semantic web technologies in automated chemical classification and annotation of lipids. Our prototype framework consists of two components: an ontology and a set of federated web services that operate upon it. The formal lipid ontology we use here extends a part of the LiPrO ontology and draws on the lipid hierarchy in the LIPID MAPS database, as well as literature-derived knowledge. The federated semantic web services that operate upon this ontology are deployed within the Semantic Annotation, Discovery, and Integration (SADI) framework. Structure-based lipid classification is enacted by two core services. Firstly, a structural annotation service detects and enumerates relevant functional groups for a specified chemical structure. A second service reasons over lipid ontology class descriptions using the attributes obtained from the annotation service and identifies the appropriate lipid classification. We extend the utility of these core services by combining them with additional SADI services that retrieve associations between lipids and proteins and identify publications related to specified lipid types. We analyze the performance of SADI-enabled eicosanoid classification relative to the LIPID MAPS classification and reflect on the contribution of our integrative methodology in the context of high-throughput lipidomics. CONCLUSIONS: Our prototype framework is capable of accurate automated classification of lipids and facile integration of lipid class information with additional data obtained with SADI web services. The potential of programming-free integration of external web services through the SADI framework offers an opportunity for development of powerful novel applications in lipidomics. We conclude that semantic web technologies can provide an accurate and versatile means of classification and annotation of lipids.
Leonid L. Chepelev, Alexandre Riazanov, Alexandre Kouznetsov, Hong Sang Low, Michel Dumontier, Christopher J. O. Baker
BMC Bioinform.5
2011 MoSuMo: A Semantic Web service to generate electrostatic potentials across solvent excluded protein surfaces and binding pockets
Alexander Gawronski, Michel Dumontier
Comput. Graph.2
2010 Realism for scientific ontologies
abstract
Science aims to develop an accurate understanding of reality through a variety of rigorously empirical and formal methods. Ontologies are used to formalize the meaning of terms within a domain of discourse. The Basic Formal Ontology (BFO) is an ontology of particular importance in the biomedical domains, where it provides the top-level for numerous ontologies, including those admitted as part of the OBO Foundry collection. The BFO requires that all classes in an ontology are actually instantiated in reality. Despite the fact that it is hard to show whether entities of some kind exist or do not exist in reality (especially for unobservable entities like elementary particles), this criterion fails to satisfy the need of scientists to communicate their findings and theories unambiguously. We discuss the problems that arise due to the BFO's realism criterion and suggest viable alternatives.
Michel Dumontier, Robert Hoehndorf
FOIS1
2010 Relations as patterns: bridging the gap between OBO and OWL
abstract
BACKGROUND: most biomedical ontologies are represented in the OBO Flatfile Format, which is an easy-to-use graph-based ontology language. The semantics of the OBO Flatfile Format 1.2 enforces a strict predetermined interpretation of relationship statements between classes. It does not allow flexible specifications that provide better approximations of the intuitive understanding of the considered relations. If relations cannot be accurately expressed then ontologies built upon them may contain false assertions and hence lead to false inferences. Ontologies in the OBO Foundry must formalize the semantics of relations according to the OBO Relationship Ontology (RO). Therefore, being able to accurately express the intended meaning of relations is of crucial importance. Since the Web Ontology Language (OWL) is an expressive language with a formal semantics, it is suitable to de ne the meaning of relations accurately. RESULTS: we developed a method to provide definition patterns for relations between classes using OWL and describe a novel implementation of the RO based on this method. We implemented our extension in software that converts ontologies in the OBO Flatfile Format to OWL, and also provide a prototype to extract relational patterns from OWL ontologies using automated reasoning. The conversion software is freely available at http://bioonto.de/obo2owl, and can be accessed via a web interface. CONCLUSIONS: explicitly defining relations permits their use in reasoning software and leads to a more flexible and powerful way of representing biomedical ontologies. Using the extended langua0067e and semantics avoids several mistakes commonly made in formalizing biomedical ontologies, and can be used to automatically detect inconsistencies. The use of our method enables the use of graph-based ontologies in OWL, and makes complex OWL ontologies accessible in a graph-based form. Thereby, our method provides the means to gradually move the representation of biomedical ontologies into formal knowledge representation languages that incorporates an explicit semantics. Our method facilitates the use of OWL-based software in the back-end while ontology curators may continue to develop ontologies with an OBO-style front-end.
Robert Hoehndorf, Anika Oellrich, Michel Dumontier, Janet Kelso, Dietrich Rebholz-Schuhmann, Heinrich Herre
BMC Bioinform.3
2010 Modeling and querying graphical representations of statistical data
Michel Dumontier, Leo Ferres, Natalia Villanueva-Rosales
J. Web Semant.1
2009 Towards pharmacogenomics knowledge discovery with the semantic web
abstract
Pharmacogenomics aims to understand pharmacological response with respect to genetic variation. Essential to the delivery of better health care is the use of pharmacogenomics knowledge to answer questions about therapeutic, pharmacological or genetic aspects. Several XML markup languages have been developed to capture pharmacogenomic and related information so as to facilitate data sharing. However, recent advances in semantic web technologies have presented exciting new opportunities for pharmacogenomics knowledge discovery by representing the information with machine understandable semantics. Progress in this area is illustrated with reference to the personalized medicine project that aims to facilitate pharmacogenomics knowledge discovery through intuitive knowledge capture and sophisticated question answering using automated reasoning over expressive ontologies.
Michel Dumontier, Natalia Villanueva-Rosales
Briefings Bioinform.1
2008 Report on semantic web for health care and life sciences workshop
abstract
The Semantic Web for Health Care and Life Sciences Workshop will be held in Beijing, China, on April 22, 2008. The goal of the workshop is to foster the development and advancement in the use of Semantic Web technologies to facilitate collaboration, research and development, and innovation adoption in the domains of Health Care and Life Sciences, We also encourage the participation of all research communities in this event, with enhanced participation from Asia due to the location of the event. The workshop consists of two invited keynote talks, eight peer-reviewed presentations, and one panel discussion.
Huajun Chen, Kei-Hoi Cheung, Michel Dumontier, Eric Prud'hommeaux, Alan Ruttenberg, Susie Stephens
WWW3
2008 yOWL: An ontology-driven knowledge base for yeast biologists
Natalia Villanueva-Rosales, Michel Dumontier
J. Biomed. Informatics2
2006 Domain-based small molecule binding site annotation
abstract
BACKGROUND: Accurate small molecule binding site information for a protein can facilitate studies in drug docking, drug discovery and function prediction, but small molecule binding site protein sequence annotation is sparse. The Small Molecule Interaction Database (SMID), a database of protein domain-small molecule interactions, was created using structural data from the Protein Data Bank (PDB). More importantly it provides a means to predict small molecule binding sites on proteins with a known or unknown structure and unlike prior approaches, removes large numbers of false positive hits arising from transitive alignment errors, non-biologically significant small molecules and crystallographic conditions that overpredict ion binding sites. DESCRIPTION: Using a set of co-crystallized protein-small molecule structures as a starting point, SMID interactions were generated by identifying protein domains that bind to small molecules, using NCBI's Reverse Position Specific BLAST (RPS-BLAST) algorithm. SMID records are available for viewing at http://smid.blueprint.org. The SMID-BLAST tool provides accurate transitive annotation of small-molecule binding sites for proteins not found in the PDB. Given a protein sequence, SMID-BLAST identifies domains using RPS-BLAST and then lists potential small molecule ligands based on SMID records, as well as their aligned binding sites. A heuristic ligand score is calculated based on E-value, ligand residue identity and domain entropy to assign a level of confidence to hits found. SMID-BLAST predictions were validated against a set of 793 experimental small molecule interactions from the PDB, of which 472 (60%) of predicted interactions identically matched the experimental small molecule and of these, 344 had greater than 80% of the binding site residues correctly identified. Further, we estimate that 45% of predictions which were not observed in the PDB validation set may be true positives. CONCLUSION: By focusing on protein domain-small molecule interactions, SMID is able to cluster similar interactions and detect subtle binding patterns that would not otherwise be obvious. Using SMID-BLAST, small molecule targets can be predicted for any protein sequence, with the only limitation being that the small molecule must exist in the PDB. Validation results and specific examples within illustrate that SMID-BLAST has a high degree of accuracy in terms of predicting both the small molecule ligand and binding site residue positions for a query protein.
Kevin A. Snyder, Howard J. Feldman, Michel Dumontier, John J. Salama, Christopher W. V. Hogue
BMC Bioinform.3
2002 NBLAST: a cluster variant of BLAST for NxN comparisons
abstract
BACKGROUND: The BLAST algorithm compares biological sequences to one another in order to determine shared motifs and common ancestry. However, the comparison of all non-redundant (NR) sequences against all other NR sequences is a computationally intensive task. We developed NBLAST as a cluster computer implementation of the BLAST family of sequence comparison programs for the purpose of generating pre-computed BLAST alignments and neighbour lists of NR sequences. RESULTS: NBLAST performs the heuristic BLAST algorithm and generates an exhaustive database of alignments, but it only computes alignments (i.e. the upper triangle) of a possible N2 alignments, where N is the set of all sequences to be compared. A task-partitioning algorithm allows for cluster computing across all cluster nodes and the NBLAST master process produces a BLAST sequence alignment database and a list of sequence neighbours for each sequence record. The resulting sequence alignment and neighbour databases are used to serve the SeqHound query system through a C/C++ and PERL Application Programming Interface (API). CONCLUSIONS: NBLAST offers a local alternative to the NCBI's remote Entrez system for pre-computed BLAST alignments and neighbour queries. On our 216-processor 450 MHz PIII cluster, NBLAST requires ~24 hrs to compute neighbours for 850000 proteins currently in the non-redundant protein database.
Michel Dumontier, Christopher W. V. Hogue
BMC Bioinform.1
2002 Species-specific protein sequence and fold optimizations
abstract
BACKGROUND: An organism's ability to adapt to its particular environmental niche is of fundamental importance to its survival and proliferation. In the largest study of its kind, we sought to identify and exploit the amino-acid signatures that make species-specific protein adaptation possible across 100 complete genomes. RESULTS: Environmental niche was determined to be a significant factor in variability from correspondence analysis using the amino acid composition of over 360,000 predicted open reading frames (ORFs) from 17 archaea, 76 bacteria and 7 eukaryote complete genomes. Additionally, we found clusters of phylogenetically unrelated archaea and bacteria that share similar environments by amino acid composition clustering. Composition analyses of conservative, domain-based homology modeling suggested an enrichment of small hydrophobic residues Ala, Gly, Val and charged residues Asp, Glu, His and Arg across all genomes. However, larger aromatic residues Phe, Trp and Tyr are reduced in folds, and these results were not affected by low complexity biases. We derived two simple log-odds scoring functions from ORFs (CG) and folds (CF) for each of the complete genomes. CF achieved an average cross-validation success rate of 85 +/- 8% whereas the CG detected 73 +/- 9% species-specific sequences when competing against all other non-redundant CG. Continuously updated results are available at http://genome.mshri.on.ca. CONCLUSION: Our analysis of amino acid compositions from the complete genomes provides stronger evidence for species-specific and environmental residue preferences in genomic sequences as well as in folds. Scoring functions derived from this work will be useful in future protein engineering experiments and possibly in identifying horizontal transfer events.
Michel Dumontier, Katerina Michalickova, Christopher W. V. Hogue
BMC Bioinform.1
2002 SeqHound: biological sequence and structure database as a platform for bioinformatics research
abstract
BACKGROUND: SeqHound has been developed as an integrated biological sequence, taxonomy, annotation and 3-D structure database system. It provides a high-performance server platform for bioinformatics research in a locally-hosted environment. RESULTS: SeqHound is based on the National Center for Biotechnology Information data model and programming tools. It offers daily updated contents of all Entrez sequence databases in addition to 3-D structural data and information about sequence redundancies, sequence neighbours, taxonomy, complete genomes, functional annotation including Gene Ontology terms and literature links to PubMed. SeqHound is accessible via a web server through a Perl, C or C++ remote API or an optimized local API. It provides functionality necessary to retrieve specialized subsets of sequences, structures and structural domains. Sequences may be retrieved in FASTA, GenBank, ASN.1 and XML formats. Structures are available in ASN.1, XML and PDB formats. Emphasis has been placed on complete genomes, taxonomy, domain and functional annotation as well as 3-D structural functionality in the API, while fielded text indexing functionality remains under development. SeqHound also offers a streamlined WWW interface for simple web-user queries. CONCLUSIONS: The system has proven useful in several published bioinformatics projects such as the BIND database and offers a cost-effective infrastructure for research. SeqHound will continue to develop and be provided as a service of the Blueprint Initiative at the Samuel Lunenfeld Research Institute. The source code and examples are available under the terms of the GNU public license at the Sourceforge site http://sourceforge.net/projects/slritools/ in the SLRI Toolkit.
Katerina Michalickova, Gary D. Bader, Michel Dumontier, Hao Lieu, Doron Betel, Ruth Isserlin, Christopher W. V. Hogue
BMC Bioinform.3