Olivier Bodenreider

dblp:29/3051 · DBLP profile ↗
← Back
118ranked-venue papers
24as first author
8since 2021 · last 2023
0000-0003-4769-4217ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 110 · 23 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2023 A deep learning approach to identify missing is-a relations in SNOMED CT
abstract
OBJECTIVE: SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. MATERIALS AND METHODS: Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. RESULTS: We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. CONCLUSIONS: The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT.
Rashmie Abeysinghe, Fengbo Zheng, Elmer V. Bernstam, Jay Shi, Olivier Bodenreider, Licong Cui
J. Am. Medical Informatics Assoc.5
2023 A practical strategy to use the ICD-11 for morbidity coding in the United States without a clinical modification
abstract
OBJECTIVE: The aim of this study was to derive and evaluate a practical strategy of replacing ICD-10-CM codes by ICD-11 for morbidity coding in the United States, without the creation of a Clinical Modification. MATERIALS AND METHODS: A stepwise strategy is described, using first the ICD-11 stem codes from the Mortality and Morbidity Statistics (MMS) linearization, followed by exposing Foundation entities, then adding postcoordination (with existing codes and adding new stem codes if necessary), with creating new stem codes as the last resort. The strategy was evaluated by recoding 2 samples of ICD-10-CM codes comprised of frequently used codes and all codes from the digestive diseases chapter. RESULTS: Among the 1725 ICD-10-CM codes examined, the cumulative coverage at the stem code, Foundation, and postcoordination levels are 35.2%, 46.5% and 89.4% respectively. 7.1% of codes require new extension codes and 3.5% require new stem codes. Among the new extension codes, severity scale values and anatomy are the most common categories. 5.5% of codes are not one-to-one matches (1 ICD-10-CM code matched to 1 ICD-11 stem code or Foundation entity) which could be potentially challenging. CONCLUSION: Existing ICD-11 content can achieve full representation of almost 90% of ICD-10-CM codes, provided that postcoordination can be used and the coding guidelines and hierarchical structures of ICD-10-CM and ICD-11 can be harmonized. The various options examined in this study should be carefully considered before embarking on the traditional approach of a full-fledged ICD-11-CM.
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
J. Am. Medical Informatics Assoc.5
2023 Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learning
abstract
OBJECTIVE: Use heuristic, deep learning (DL), and hybrid AI methods to predict semantic group (SG) assignments for new UMLS Metathesaurus atoms, with target accuracy ≥95%. MATERIALS AND METHODS: We used train-test datasets from successive 2020AA-2022AB UMLS Metathesaurus releases. Our heuristic "waterfall" approach employed a sequence of 7 different SG prediction methods. Atoms not qualifying for a method were passed on to the next method. The DL approach generated BioWordVec and SapBERT embeddings for atom names, BioWordVec embeddings for source vocabulary names, and BioWordVec embeddings for atom names of the second-to-top nodes of an atom's source hierarchy. We fed a concatenation of the 4 embeddings into a fully connected multilayer neural network with an output layer of 15 nodes (one for each SG). For both approaches, we developed methods to estimate the probability that their predicted SG for an atom would be correct. Based on these estimations, we developed 2 hybrid SG prediction methods combining the strengths of heuristic and DL methods. RESULTS: The heuristic waterfall approach accurately predicted 94.3% of SGs for 1 563 692 new unseen atoms. The DL accuracy on the same dataset was also 94.3%. The hybrid approaches achieved an average accuracy of 96.5%. CONCLUSION: Our study demonstrated that AI methods can predict SG assignments for new UMLS atoms with sufficient accuracy to be potentially useful as an intermediate step in the time-consuming task of assigning new atoms to UMLS concepts. We showed that for SG prediction, combining heuristic methods and DL methods can produce better results than either alone.
Randolph A. Miller, Olivier Bodenreider, Vinh Nguyen 0002, Kin Wah Fung
J. Am. Medical Informatics Assoc.3
2022 Context-Enriched Learning Models for Aligning Biomedical Vocabularies at Scale in the UMLS Metathesaurus
abstract
The Unified Medical Language System (UMLS) Metathesaurus construction process mainly relies on lexical algorithms and manual expert curation for integrating over 200 biomedical vocabularies. A lexical-based learning model (LexLM) was developed to predict synonymy among Metathesaurus terms and largely outperforms a rule-based approach (RBA) that approximates the current construction process. However, the LexLM has the potential for being improved further because it only uses lexical information from the source vocabularies, while the RBA also takes advantage of contextual information. We investigate the role of multiple types of contextual information available to the UMLS editors, namely source synonymy (SS), source semantic group (SG), and source hierarchical relations (HR), for the UMLS vocabulary alignment (UVA) problem. In this paper, we develop multiple variants of context-enriched learning models (ConLMs) by adding to the LexLM the types of contextual information listed above. We represent these context types in context-enriched knowledge graphs (ConKGs) with four variants ConSS, ConSG, ConHR, and ConAll. We train these ConKG embeddings using seven KG embedding techniques. We create the ConLMs by concatenating the ConKG embedding vectors with the word embedding vectors from the LexLM. We evaluate the performance of the ConLMs using the UVA generalization test datasets with hundreds of millions of pairs. Our extensive experiments show a significant performance improvement from the ConLMs over the LexLM, namely +5.0% in precision (93.75%), +0.69% in recall (93.23%), +2.88% in F1 (93.49%) for the best ConLM. Our experiments also show that the ConAll variant including the three context types takes more time, but does not always perform better than other variants with a single context type. Finally, our experiments show that the pairs of terms with high lexical similarity benefit most from adding contextual information, namely +6.56% in precision (94.97%), +2.13% in recall (93.23%), +4.35% in F1 (94.09%) for the best ConLM. The pairs with lower degrees of lexical similarity also show performance improvement with +0.85% in F1 (96%) for low similarity and +1.31% in F1 (96.34%) for no similarity. These results demonstrate the importance of using contextual information in the UVA problem.
Vinh Nguyen 0002, Hong Yung Yip, Goonmeet Bajaj, Thilini Wijesiriwardene, Vishesh Javangula, Srinivasan Parthasarathy 0001, Amit P. Sheth, Olivier Bodenreider
WWW8
2021 Reproducibility in biomedical natural language processing: A FAIR approach to what we need to know
Kevin Cohen 0001, Anna Ripple, Asma Ben Abacha, Olivier Bodenreider, Orin Hargraves, Karin Verspoor, Pierre Zweigenbaum, Dina Demner-Fushman
AMIA4
2021 Can ICD-11 Replace ICD-10-CM for Morbidity Coding in the U.S.?
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
AMIA5
2021 Biomedical Vocabulary Alignment at Scale in the UMLS Metathesaurus
abstract
With 214 source vocabularies, the construction and maintenance process of the UMLS (Unified Medical Language System) Metathesaurus terminology integration system is costly, time-consuming, and error-prone as it primarily relies on (1) lexical and semantic processing for suggesting groupings of synonymous terms, and (2) the expertise of UMLS editors for curating these synonymy predictions. This paper aims to improve the UMLS Metathesaurus construction process by developing a novel supervised learning approach for improving the task of suggesting synonymous pairs that can scale to the size and diversity of the UMLS source vocabularies. We evaluate this deep learning (DL) approach against a rule-based approach (RBA) that approximates the current UMLS Metathesaurus construction process. The key to the generalizability of our approach is the use of various degrees of lexical similarity in negative pairs during the training process. Our initial experiments demonstrate the strong performance across multiple datasets of our DL approach in terms of recall (91-92%), precision (88-99%), and F1 score (89-95%). Our DL approach largely outperforms the RBA method in recall (+23%), precision (+2.4%), and F1 score (+14.1%). This novel approach has great potential for improving the UMLS Metathesaurus construction process by providing better synonymy suggestions to the UMLS editors.
Vinh Nguyen 0002, Hong Yung Yip, Olivier Bodenreider
WWW3
2021 Feasibility of replacing the ICD-10-CM with the ICD-11 for morbidity coding: A content analysis
abstract
OBJECTIVE: The study sought to assess the feasibility of replacing the International Classification of Diseases-Tenth Revision-Clinical Modification (ICD-10-CM) with the International Classification of Diseases-11th Revision (ICD-11) for morbidity coding based on content analysis. MATERIALS AND METHODS: The most frequently used ICD-10-CM codes from each chapter covering 60% of patients were identified from Medicare claims and hospital data. Each ICD-10-CM code was recoded in the ICD-11, using postcoordination (combination of codes) if necessary. Recoding was performed by 2 terminologists independently. Failure analysis was done for cases where full representation was not achieved even with postcoordination. After recoding, the coding guidance (inclusions, exclusions, and index) of the ICD-10-CM and ICD-11 codes were reviewed for conflict. RESULTS: Overall, 23.5% of 943 codes could be fully represented by the ICD-11 without postcoordination. Postcoordination is the potential game changer. It supports the full representation of 8.6% of 943 codes. Moreover, with the addition of only 9 extension codes, postcoordination supports the full representation of 35.2% of 943 codes. Coding guidance review identified potential conflicts in 10% of codes, but mostly not affecting recoding. The majority of the conflicts resulted from differences in granularity and default coding assumptions between the ICD-11 and ICD-10-CM. CONCLUSIONS: With some minor enhancements to postcoordination, the ICD-11 can fully represent almost 60% of the most frequently used ICD-10-CM codes. Even without postcoordination, 23.5% full representation is comparable to the 24.3% of ICD-9-CM codes with exact match in the ICD-10-CM, so migrating from the ICD-10-CM to the ICD-11 is not necessarily more disruptive than from the International Classification of Diseases-Ninth Revision-Clinical Modification to the ICD-10-CM. Therefore, the ICD-11 (without a CM) should be considered as a candidate to replace the ICD-10-CM for morbidity coding.
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
J. Am. Medical Informatics Assoc.5
2020 The new International Classification of Diseases 11th edition: a comparative analysis with ICD-10 and ICD-10-CM
abstract
OBJECTIVE: To study the newly adopted International Classification of Diseases 11th revision (ICD-11) and compare it to the International Classification of Diseases 10th revision (ICD-10) and International Classification of Diseases 10th revision-Clinical Modification (ICD-10-CM). MATERIALS AND METHODS: : Data files and maps were downloaded from the World Health Organization (WHO) website and through the application programming interfaces. A round trip method based on the WHO maps was used to identify equivalent codes between ICD-10 and ICD-11, which were validated by limited manual review. ICD-11 terms were mapped to ICD-10-CM through normalized lexical mapping. ICD-10-CM codes in 6 disease areas were also manually recoded in ICD-11. RESULTS: Excluding the chapters for traditional medicine, functioning assessment, and extension codes for postcoordination, ICD-11 has 14 622 leaf codes (codes that can be used in coding) compared to ICD-10 and ICD-10-CM, which has 10 607 and 71 932 leaf codes, respectively. We identified 4037 pairs of ICD-10 and ICD-11 codes that were equivalent (estimated accuracy of 96%) by our round trip method. Lexical matching between ICD-11 and ICD-10-CM identified 4059 pairs of possibly equivalent codes. Manual recoding showed that 60% of a sample of 388 ICD-10-CM codes could be fully represented in ICD-11 by precoordinated codes or postcoordination. CONCLUSION: In ICD-11, there is a moderate increase in the number of codes over ICD-10. With postcoordination, it is possible to fully represent the meaning of a high proportion of ICD-10-CM codes, especially with the addition of a limited number of extension codes.
Kin Wah Fung, Julia Xu, Olivier Bodenreider
J. Am. Medical Informatics Assoc.3
2020 Assessing the enrichment of dietary supplement coverage in the Unified Medical Language System
abstract
OBJECTIVE: We sought to assess the need for additional coverage of dietary supplements (DS) in the Unified Medical Language System (UMLS) by investigating (1) the overlap between the integrated DIetary Supplements Knowledge base (iDISK) DS ingredient terminology and the UMLS and (2) the coverage of iDISK and the UMLS over DS mentions in the biomedical literature. MATERIALS AND METHODS: We estimated the overlap between iDISK and the UMLS by mapping iDISK to the UMLS using exact and normalized strings. The coverage of iDISK and the UMLS over DS mentions in the biomedical literature was evaluated via a DS named-entity recognition (NER) task within PubMed abstracts. RESULTS: The coverage analysis revealed that only 30% of iDISK terms can be matched to the UMLS, although these cover over 99% of iDISK concepts. A manual review revealed that a majority of the unmatched terms represented new synonyms, rather than lexical variants. For NER, iDISK nearly doubles the precision and achieves a higher F1 score than the UMLS, while maintaining a competitive recall. DISCUSSION: While iDISK has significant concept overlap with the UMLS, it contains many novel synonyms. Furthermore, almost 3000 of these overlapping UMLS concepts are missing a DS designation, which could be provided by iDISK. The NER experiments show that the specialization of iDISK is useful for identifying DS mentions. CONCLUSIONS: Our results show that the DS representation in the UMLS could be enriched by adding DS designations to many concepts and by adding new synonyms.
Jake Vasilakes, Anusha Bompelli, Jeffrey R. Bishop, Terrence Adam, Olivier Bodenreider, Rui Zhang 0028
J. Am. Medical Informatics Assoc.5
2018 Using RxNav for drug analytics - How to interpret obsolete drug identifiers?
Olivier Bodenreider, Lee B. Peters
AMIA1
2018 Historically Comprehensive Medications Metadata for i2b2 Data Warehouses
Jay Pedersen, Bret J. Gardner, James R. Campbell 0001, Walter S. Campbell, Carol Geary, Lee B. Peters, Olivier Bodenreider, James C. McClay
AMIA7
2018 RxNav-in-a-Box - A locally-installable version of RxNav and related APIs
Lee B. Peters, Richard Rice, Olivier Bodenreider
AMIA3
2018 Auditing SNOMED CT hierarchical relations based on lexical features of concepts in non-lattice subgraphs
Licong Cui, Olivier Bodenreider, Jay Shi, Guo-Qiang Zhang 0001
J. Biomed. Informatics2
2017 RxNorm Concept History Service
Olivier Bodenreider, Lee B. Peters
AMIA1
2017 RxMix - Use of NLM drug APIs by non-programmers
Olivier Bodenreider, Lee B. Peters
AMIA1
2017 Mapping U.S. FDA National Drug Codes to Anatomical-Therapeutic-Chemical Classes using RxNorm
Fabrício S. P. Kury, Olivier Bodenreider
AMIA2
2017 RxNav 2.0 - A web-based, mobile-responsive RxNorm browser
Lee B. Peters, Olivier Bodenreider
AMIA2
2017 Identifying Potentially Missing Hierarchical Relations in SNOMED CT based on Lexical Features - Impact of Synonyms and Lexico-syntactic Constraints
Satyajeet Raje, Olivier Bodenreider
AMIA2
2017 Detecting Adverse Drug Event Safety Signals from MEDLINE Reports: Challenges in Employing Cross-terminology Mapping of MeSH to MedDRA
Abhivyakti Sawarkar, Alfred Sorbello, Anna Ripple, Olivier Bodenreider
AMIA4
2017 Mining non-lattice subgraphs for detecting missing hierarchical relations and concepts in SNOMED CT
abstract
OBJECTIVE: Quality assurance of large ontological systems such as SNOMED CT is an indispensable part of the terminology management lifecycle. We introduce a hybrid structural-lexical method for scalable and systematic discovery of missing hierarchical relations and concepts in SNOMED CT. MATERIAL AND METHODS: All non-lattice subgraphs (the structural part) in SNOMED CT are exhaustively extracted using a scalable MapReduce algorithm. Four lexical patterns (the lexical part) are identified among the extracted non-lattice subgraphs. Non-lattice subgraphs exhibiting such lexical patterns are often indicative of missing hierarchical relations or concepts. Each lexical pattern is associated with a potential specific type of error. RESULTS: Applying the structural-lexical method to SNOMED CT (September 2015 US edition), we found 6801 non-lattice subgraphs that matched these lexical patterns, of which 2046 were amenable to visual inspection. We evaluated a random sample of 100 small subgraphs, of which 59 were reviewed in detail by domain experts. All the subgraphs reviewed contained errors confirmed by the experts. The most frequent type of error was missing is-a relations due to incomplete or inconsistent modeling of the concepts. CONCLUSIONS: Our hybrid structural-lexical method is innovative and proved effective not only in detecting errors in SNOMED CT, but also in suggesting remediation for these errors.
Licong Cui, Wei Zhu 0010, Shiqiang Tao, James T. Case, Olivier Bodenreider, Guo-Qiang Zhang 0001
J. Am. Medical Informatics Assoc.5
2017 Comparison of three commercial knowledge bases for detection of drug-drug interactions in clinical decision support
abstract
OBJECTIVE: To compare 3 commercial knowledge bases (KBs) used for detection and avoidance of potential drug-drug interactions (DDIs) in clinical practice. METHODS: Drugs in the DDI tables from First DataBank (FDB), Micromedex, and Multum were mapped to RxNorm. The KBs were compared at the clinical drug, ingredient, and DDI rule levels. The KBs were evaluated against a reference list of highly significant DDIs from the Office of the National Coordinator for Health Information Technology (ONC). The KBs and the ONC list were applied to a prescription data set to simulate their use in clinical decision support. RESULTS: The KBs contained 1.6 million (FDB), 4.5 million (Micromedex), and 4.8 million (Multum) clinical drug pairs. Altogether, there were 8.6 million unique pairs, of which 79% were found only in 1 KB and 5% in all 3 KBs. However, there was generally more agreement than disagreement in the severity rankings, especially in the contraindicated category. The KBs covered 99.8-99.9% of the alerts of the ONC list and would have generated 25 (FDB), 145 (Micromedex), and 84 (Multum) alerts per 1000 prescriptions. CONCLUSION: The commercial KBs differ considerably in size and quantity of alerts generated. There is less variability in severity ranking of DDIs than suggested by previous studies. All KBs provide very good coverage of the ONC list. More work is needed to standardize the editorial policies and evidence for inclusion of DDIs to reduce variation among knowledge sources and improve relevance. Some DDIs considered contraindicated in all 3 KBs might be possible candidates to add to the ONC list.
Kin Wah Fung, Joan Kapusnik-Uner, Jean Cunningham, Stefanie Higby-Baker, Olivier Bodenreider
J. Am. Medical Informatics Assoc.5
2017 Toward multimodal signal detection of adverse drug reactions
Rave Harpaz, William DuMouchel, Martijn J. Schuemie, Olivier Bodenreider, Carol Friedman, Eric Horvitz, Anna Ripple, Alfred Sorbello, Ryen W. White, Rainer Winnenburg, Nigam H. Shah
J. Biomed. Informatics4
2016 Resources for Analyzing Drug Prescription Datasets
Olivier Bodenreider, Vojtech Huser, Christian G. Reich
AMIA1
2016 Characterizing the semantic composition of the UMLS Metathesaurus over time
Olivier Bodenreider, Lee B. Peters
AMIA1
2016 Assessing the potential risk in drug prescriptions during pregnancy
Ferdinand Dhombres, Vojtech Huser, Laritza Rodriguez, Olivier Bodenreider
AMIA4
2016 Exploring the Use of ClinicalTrials.gov Trial Results Data for Pharmacovigilance
Vojtech Huser, Olivier Bodenreider
AMIA2
2016 NDC Properties in RxNorm
Lee B. Peters, Olivier Bodenreider
AMIA2
2016 PubMed 'Early Alerts': Towards Better Precision of Literature Searching for Pharmacovigilance Information based on an Assessment of Relevance Feedback
Anna Ripple, Alfred Sorbello, Shahrukh Haider, Olivier Bodenreider
AMIA4
2016 The digital revolution in phenotyping
abstract
Phenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data.
Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier
Briefings Bioinform.6
2015 Navigating between Drug Classes and RxNorm Drugs with RxClass
Olivier Bodenreider, Lee B. Peters, Thang Nguyen 0004
AMIA1
2015 Finding Similar Drug Classes using RxClass
Thang Nguyen 0004, Lee B. Peters, Olivier Bodenreider
AMIA3
2015 Approaches to Supporting the Analysis of Historical Medication Datasets with RxNorm
Lee B. Peters, Olivier Bodenreider
AMIA2
2015 PubMed 'Early Alerts': A Pilot Study to Support Prospective Detection of Emerging Adverse Drug Events
Alfred Sorbello, Anna Ripple, Olivier Bodenreider
AMIA3
2015 Context-driven automatic subgraph creation for literature-based discovery
Delroy Cameron, Ramakanth Kavuluru, Thomas C. Rindflesch, Amit P. Sheth, Krishnaprasad Thirunarayan, Olivier Bodenreider
J. Biomed. Informatics6
2015 Leveraging MEDLINE indexing for pharmacovigilance - Inherent limitations and mitigation strategies
Rainer Winnenburg, Alfred Sorbello, Anna Ripple, Rave Harpaz, Joseph M. Tonning, Ana Szarfman, Henry Francis, Olivier Bodenreider
J. Biomed. Informatics8
2014 Analyzing U.S. prescription lists with RxNorm and the ATC/DDD Index
Olivier Bodenreider, Laritza Rodriguez
AMIA1
2014 Coverage of Rare Disease Names in Standard Terminologies and Implications for Patients, Providers, and Research
Kin Wah Fung, Rachel L. Richesson, Olivier Bodenreider
AMIA3
2014 Creating, Maintaining and Publishing Value Sets in the VSAC
Emir Khatipov, Maureen Madden, Pishing Chiang, Philip Chuang, Ivor D'Souza, Rainer Winnenburg, Olivier Bodenreider, Julia L. Skapik, Robert C. McClure, Steven Emrick
AMIA8
2014 RxClass - Navigating between Drug Classes and RxNorm Drugs
Thang Nguyen 0004, Lee B. Peters, Olivier Bodenreider
AMIA3
2014 Automatic coding of Free-Text Medication Data recorded by Research Coordinators
Laritza Rodriguez, Vojtech Huser, Olivier Bodenreider, James J. Cimino
AMIA3
2014 Facilitating Reconciliation of Inter-Annotator Disagreements
Johann Stan, Olivier Bodenreider, Kin Wah Fung, Dina Demner-Fushman
AMIA2
2014 Desiderata for an authoritative Representation of MeSH in RDF
Rainer Winnenburg, Olivier Bodenreider
AMIA2
2014 MaPLE: A MapReduce Pipeline for Lattice-based Evaluation and its application to SNOMED CT
abstract
Non-lattice fragments are often indicative of structural anomalies in ontological systems and, as such, represent possible areas of focus for subsequent quality assurance work. However, extracting the non-lattice fragments in large ontological systems is computationally expensive if not prohibitive, using a traditional sequential approach. In this paper we present a general MapReduce pipeline, called MaPLE (MapReduce Pipeline for Lattice-based Evaluation), for extracting non-lattice fragments in large partially ordered sets and demonstrate its applicability in ontology quality assurance. Using MaPLE in a 30-node Hadoop local cloud, we systematically extracted non-lattice fragments in 8 SNOMED CT versions from 2009 to 2014 (each containing over 300k concepts), with an average total computing time of less than 3 hours per version. With dramatically reduced time, MaPLE makes it feasible not only to perform exhaustive structural analysis of large ontological hierarchies, but also to systematically track structural changes between versions. Our change analysis showed that the average change rates on the non-lattice pairs are up to 38.6 times higher than the change rates of the background structure (concept nodes). This demonstrates that fragments around non-lattice pairs exhibit significantly higher rates of change in the process of ontological evolution.
Guo-Qiang Zhang 0001, Wei Zhu 0010, Shiqiang Tao, Olivier Bodenreider, Licong Cui
IEEE BigData5
2014 Don't like RDF reification?: making statements about statements using singleton property
abstract
Statements about RDF statements, or meta triples, provide additional information about individual triples, such as the source, the occurring time or place, or the certainty. Integrating such meta triples into semantic knowledge bases would enable the querying and reasoning mechanisms to be aware of provenance, time, location, or certainty of triples. However, an efficient RDF representation for such meta knowledge of triples remains challenging. The existing standard reification approach allows such meta knowledge of RDF triples to be expressed using RDF by two steps. The first step is representing the triple by a Statement instance which has subject, predicate, and object indicated separately in three different triples. The second step is creating assertions about that instance as if it is a statement. While reification is simple and intuitive, this approach does not have formal semantics and is not commonly used in practice as described in the RDF Primer. In this paper, we propose a novel approach called Singleton Property for representing statements about statements and provide a formal semantics for it. We explain how this singleton property approach fits well with the existing syntax and formal semantics of RDF, and the syntax of SPARQL query language. We also demonstrate the use of singleton property in the representation and querying of meta knowledge in two examples of Semantic Web knowledge bases: YAGO2 and BKR. Our experiments on the BKR show that the singleton property approach gives a decent performance in terms of number of triples, query length and query execution time compared to existing approaches. This approach, which is also simple and intuitive, can be easily adopted for representing and querying statements about statements in other knowledge bases.
Vinh Nguyen 0002, Olivier Bodenreider, Amit P. Sheth
WWW2
2013 The NLM Value Set Authority Center at 1 year
Steven Emrick, Pishing Chiang, Philip Chuang, Maureen Madden, Rainer Winnenburg, Robert C. McClure, Ivor D'Souza, Olivier Bodenreider
AMIA9
2013 Network Visualization of UMLS Source Vocabularies using Semantic Groups
Thai Le, Bastien Rance, Olivier Bodenreider
AMIA3
2013 Displaying Drug Classes in RxNav
Lee B. Peters, Thang Nguyen 0004, Olivier Bodenreider
AMIA3
2013 Metrics for assessing the quality of value sets in clinical quality measures
Rainer Winnenburg, Olivier Bodenreider
AMIA2
2013 A graph-based recovery and decomposition of Swanson's hypothesis using semantic predications
Delroy Cameron, Olivier Bodenreider, Hima Yalamanchili, Tu Danh, Sreeram Vallabhaneni, Krishnaprasad Thirunarayan, Amit P. Sheth, Thomas C. Rindflesch
J. Biomed. Informatics2
2012 Quality Assurance in LOINC using Description Logic
Tomasz Adamusiak, Olivier Bodenreider
AMIA2
2012 RxBatch - Batch Operations for RxNorm and NDF-RT APIs
Lee B. Peters, Olivier Bodenreider, Thang Nguyen 0004
AMIA2
2012 Issues in Creating and Maintaining Value Sets for Clinical Quality Measures
Rainer Winnenburg, Olivier Bodenreider
AMIA2
2012 A mutation-centric approach to identifying pharmacogenomic relations in text
Bastien Rance, Emily Doughty, Dina Demner-Fushman, Maricel G. Kann, Olivier Bodenreider
J. Biomed. Informatics5
2011 Semantic Predications for Complex Information Needs in Biomedical Literature
abstract
Many complex information needs that arise in biomedical disciplines require exploring multiple documents in order to obtain information. While traditional information retrieval techniques that return a single ranked list of documents are quite common for such tasks, they may not always be adequate. The main issue is that ranked lists typically impose a significant burden on users to filter out irrelevant documents. Additionally, users must intuitively reformulate their search query when relevant documents have not been not highly ranked. Furthermore, even after interesting documents have been selected, very few mechanisms exist that enable document-to-document transitions. In this paper, we demonstrate the utility of assertions extracted from biomedical text (called semantic predications) to facilitate retrieving relevant documents for complex information needs. Our approach offers an alternative to query reformulation by establishing a framework for transitioning from one document to another. We evaluate this novel knowledge-driven approach using precision and recall metrics on the 2006 TREC Genomics Track.
Delroy Cameron, Ramakanth Kavuluru, Olivier Bodenreider, Pablo N. Mendes, Amit P. Sheth, Krishnaprasad Thirunarayan
BIBM3
2011 Toward an automatic method for extracting cancer- and other disease-related point mutations from the biomedical literature
abstract
MOTIVATION: A major goal of biomedical research in personalized medicine is to find relationships between mutations and their corresponding disease phenotypes. However, most of the disease-related mutational data are currently buried in the biomedical literature in textual form and lack the necessary structure to allow easy retrieval and visualization. We introduce a high-throughput computational method for the identification of relevant disease mutations in PubMed abstracts applied to prostate (PCa) and breast cancer (BCa) mutations. RESULTS: We developed the extractor of mutations (EMU) tool to identify mutations and their associated genes. We benchmarked EMU against MutationFinder--a tool to extract point mutations from text. Our results show that both methods achieve comparable performance on two manually curated datasets. We also benchmarked EMU's performance for extracting the complete mutational information and phenotype. Remarkably, we show that one of the steps in our approach, a filter based on sequence analysis, increases the precision for that task from 0.34 to 0.59 (PCa) and from 0.39 to 0.61 (BCa). We also show that this high-throughput approach can be extended to other diseases. DISCUSSION: Our method improves the current status of disease-mutation databases by significantly increasing the number of annotated mutations. We found 51 and 128 mutations manually verified to be related to PCa and Bca, respectively, that are not currently annotated for these cancer types in the OMIM or Swiss-Prot databases. EMU's retrieval performance represents a 2-fold improvement in the number of annotated mutations for PCa and BCa. We further show that our method can benefit from full-text analysis once there is an increase in Open Access availability of full-text articles. AVAILABILITY: Freely available at: http://bioinf.umbc.edu/EMU/ftp.
Emily Doughty, Attila Kertész-Farkas, Olivier Bodenreider, Gary Thompson, Asa Adadey, Thomas A. Peterson, Maricel G. Kann
Bioinform.3
2011 A unified framework for managing provenance information in translational research
abstract
BACKGROUND: A critical aspect of the NIH Translational Research roadmap, which seeks to accelerate the delivery of "bench-side" discoveries to patient's "bedside," is the management of the provenance metadata that keeps track of the origin and history of data resources as they traverse the path from the bench to the bedside and back. A comprehensive provenance framework is essential for researchers to verify the quality of data, reproduce scientific results published in peer-reviewed literature, validate scientific process, and associate trust value with data and results. Traditional approaches to provenance management have focused on only partial sections of the translational research life cycle and they do not incorporate "domain semantics", which is essential to support domain-specific querying and analysis by scientists. RESULTS: We identify a common set of challenges in managing provenance information across the pre-publication and post-publication phases of data in the translational research lifecycle. We define the semantic provenance framework (SPF), underpinned by the Provenir upper-level provenance ontology, to address these challenges in the four stages of provenance metadata:(a) Provenance collection - during data generation(b) Provenance representation - to support interoperability, reasoning, and incorporate domain semantics(c) Provenance storage and propagation - to allow efficient storage and seamless propagation of provenance as the data is transferred across applications(d) Provenance query - to support queries with increasing complexity over large data size and also support knowledge discovery applicationsWe apply the SPF to two exemplar translational research projects, namely the Semantic Problem Solving Environment for Trypanosoma cruzi (T.cruzi SPSE) and the Biomedical Knowledge Repository (BKR) project, to demonstrate its effectiveness. CONCLUSIONS: The SPF provides a unified framework to effectively manage provenance of translational research data during pre and post-publication phases. This framework is underpinned by an upper-level provenance ontology called Provenir that is extended to create domain-specific provenance ontologies to facilitate provenance interoperability, seamless propagation of provenance, automated querying, and analysis.
Satya Sanket Sahoo, Vinh Nguyen 0002, Olivier Bodenreider, Priti Parikh, Todd Minning, Amit P. Sheth
BMC Bioinform.3
2010 Finding Semantic Inconsistencies in UMLS using Answer Set Programming
abstract
We introduce a new method to find semantic inconsistencies (i.e., concepts with erroneous synonymity) in the Unified Medical Language System (UMLS). The idea is to identify the inconsistencies by comparing the semantic groups of hierarchically-related concepts using Answer Set Programming. With this method, we identified several inconsistent concepts in UMLS and discovered an interesting semantic pattern along hierarchies, which seems associated with wrong synonymy.
Halit Erdogan, Olivier Bodenreider, Esra Erdem 0001
AAAI2
2010 Using SPARQL to Test for Lattices: Application to Quality Assurance in Biomedical Ontologies
Guo-Qiang Zhang 0001, Olivier Bodenreider
ISWC (2)2
2010 Provenance Context Entity (PaCE): Scalable Provenance Tracking for Scientific RDF Data
Satya Sanket Sahoo, Olivier Bodenreider, Pascal Hitzler, Amit P. Sheth, Krishnaprasad Thirunarayan
SSDBM2
2010 Extracting Rx information from clinical narrative
abstract
OBJECTIVE: The authors used the i2b2 Medication Extraction Challenge to evaluate their entity extraction methods, contribute to the generation of a publicly available collection of annotated clinical notes, and start developing methods for ontology-based reasoning using structured information generated from the unstructured clinical narrative. DESIGN: Extraction of salient features of medication orders from the text of de-identified hospital discharge summaries was addressed with a knowledge-based approach using simple rules and lookup lists. The entity recognition tool, MetaMap, was combined with dose, frequency, and duration modules specifically developed for the Challenge as well as a prototype module for reason identification. MEASUREMENTS: Evaluation metrics and corresponding results were provided by the Challenge organizers. RESULTS: The results indicate that robust rule-based tools achieve satisfactory results in extraction of simple elements of medication orders, but more sophisticated methods are needed for identification of reasons for the orders and durations. LIMITATIONS: Owing to the time constraints and nature of the Challenge, some obvious follow-on analysis has not been completed yet. CONCLUSIONS: The authors plan to integrate the new modules with MetaMap to enhance its accuracy. This integration effort will provide guidance in retargeting existing tools for better processing of clinical text.
James G. Mork, Olivier Bodenreider, Dina Demner-Fushman, Rezarta Islamaj Dogan, François-Michel Lang, Zhiyong Lu, Aurélie Névéol, Lee B. Peters, Sonya E. Shooshan, Alan R. Aronson
J. Am. Medical Informatics Assoc.2
2010 Comparing and evaluating terminology services application programming interfaces: RxNav, UMLSKS and LexBIG
abstract
To facilitate the integration of terminologies into applications, various terminology services application programming interfaces (API) have been developed in the recent past. In this study, three publicly available terminology services API, RxNav, UMLSKS and LexBIG, are compared and functionally evaluated with respect to the retrieval of information from one biomedical terminology, RxNorm, to which all three services provide access. A list of queries is established covering a wide spectrum of terminology services functionalities such as finding RxNorm concepts by their name, or navigating different types of relationships. Test data were generated from the RxNorm dataset to evaluate the implementation of the functionalities in the three API. The results revealed issues with various aspects of the API implementation (eg, handling of obsolete terms by LexBIG) and documentation (eg, navigational paths used in RxNav) that were subsequently addressed by the development teams of the three API investigated. Knowledge about such discrepancies helps inform the choice of an API for a given use case.
Jyotishman Pathak, Lee B. Peters, Christopher G. Chute, Olivier Bodenreider
J. Am. Medical Informatics Assoc.4
2009 Using SNOMED CT in combination with MedDRA for reporting signal detection and adverse drug reactions reporting
Olivier Bodenreider
AMIA1
2009 Two approaches to integrating phenotype and clinical information
Anita Burgun-Parenthoine, Fleur Mougin, Olivier Bodenreider
AMIA3
2009 Alignment of the UMLS semantic network with BioTop: methodology and assessment
abstract
MOTIVATION: For many years, the Unified Medical Language System (UMLS) semantic network (SN) has been used as an upper-level semantic framework for the categorization of terms from terminological resources in biomedicine. BioTop has recently been developed as an upper-level ontology for the biomedical domain. In contrast to the SN, it is founded upon strict ontological principles, using OWL DL as a formal representation language, which has become standard in the semantic Web. In order to make logic-based reasoning available for the resources annotated or categorized with the SN, a mapping ontology was developed aligning the SN with BioTop. METHODS: The theoretical foundations and the practical realization of the alignment are being described, with a focus on the design decisions taken, the problems encountered and the adaptations of BioTop that became necessary. For evaluation purposes, UMLS concept pairs obtained from MEDLINE abstracts by a named entity recognition system were tested for possible semantic relationships. Furthermore, all semantic-type combinations that occur in the UMLS Metathesaurus were checked for satisfiability. RESULTS: The effort-intensive alignment process required major design changes and enhancements of BioTop and brought up several design errors that could be fixed. A comparison between a human curator and the ontology yielded only a low agreement. Ontology reasoning was also used to successfully identify 133 inconsistent semantic-type combinations. AVAILABILITY: BioTop, the OWL DL representation of the UMLS SN, and the mapping ontology are available at http://www.purl.org/biotop/.
Stefan Schulz 0001, Elena Beisswanger, László van den Hoek, Olivier Bodenreider, Erik M. van Mulligen
Bioinform.4
2009 A graph-based approach to auditing RxNorm
Olivier Bodenreider, Lee B. Peters
J. Biomed. Informatics1
2009 The caBIG terminology review process
abstract
The National Cancer Institute (NCI) is developing an integrated biomedical informatics infrastructure, the cancer Biomedical Informatics Grid (caBIG), to support collaboration within the cancer research community. A key part of the caBIG architecture is the establishment of terminology standards for representing data. In order to evaluate the suitability of existing controlled terminologies, the caBIG Vocabulary and Data Elements Workspace (VCDE WS) working group has developed a set of criteria that serve to assess a terminology's structure, content, documentation, and editorial process. This paper describes the evolution of these criteria and the results of their use in evaluating four standard terminologies: the Gene Ontology (GO), the NCI Thesaurus (NCIt), the Common Terminology for Adverse Events (known as CTCAE), and the laboratory portion of the Logical Objects, Identifiers, Names and Codes (LOINC). The resulting caBIG criteria are presented as a matrix that may be applicable to any terminology standardization effort.
James J. Cimino, Terry F. Hayamizu, Olivier Bodenreider, Grace A. Stafford, Martin Ringwald
J. Biomed. Informatics3
2009 Analyzing polysemous concepts from a clinical perspective: Application to auditing concept categorization in the UMLS
Fleur Mougin, Olivier Bodenreider, Anita Burgun-Parenthoine
J. Biomed. Informatics2
2009 Auditing associative relations across two knowledge sources
Lowell Vizenor, Olivier Bodenreider, Alexa T. McCray
J. Biomed. Informatics2
2008 Issues in Mapping LOINC Laboratory Tests to SNOMED CT
Olivier Bodenreider
AMIA1
2008 Auditing the NCI Thesaurus with Semantic Web Technologies
Fleur Mougin, Olivier Bodenreider
AMIA2
2008 Using the RxNorm Web Services API for Quality Assurance Purposes
Lee B. Peters, Olivier Bodenreider
AMIA2
2008 A Framework for Characterizing Drug Information Sources
Mark Sharp, Olivier Bodenreider, Nina Wacholder
AMIA2
2008 Design and Implementation of a Personal Medication Record - MyMedicationList
Kelly Zeng, Olivier Bodenreider, Stuart J. Nelson
AMIA2
2008 An ontology-driven semantic mashup of gene and biological pathway information: Application to the domain of nicotine dependence
Satya Sanket Sahoo, Olivier Bodenreider, Joni L. Rutter, Karen J. Skinner, Amit P. Sheth
J. Biomed. Informatics2
2007 Identifying Mismatches in Alignments of Large Anatomical Ontologies
Songmao Zhang, Olivier Bodenreider
AMIA2
2007 Investigating subsumption in SNOMED CT: An exploration into large description logic-based biomedical terminologies
Olivier Bodenreider, Barry Smith 0001, Anand Kumar 0005, Anita Burgun-Parenthoine
Artif. Intell. Medicine1
2007 Comparing two approaches for aligning representations of anatomy
Songmao Zhang, Kris Mork, Olivier Bodenreider, Philip A. Bernstein
Artif. Intell. Medicine3
2007 Advancing translational research with the Semantic Web
abstract
BACKGROUND: A fundamental goal of the U.S. National Institute of Health (NIH) "Roadmap" is to strengthen Translational Research, defined as the movement of discoveries in basic research to application at the clinical level. A significant barrier to translational research is the lack of uniformly structured data across related biomedical domains. The Semantic Web is an extension of the current Web that enables navigation and meaningful use of digital resources by automatic processes. It is based on common formats that support aggregation and integration of data drawn from diverse sources. A variety of technologies have been built on this foundation that, together, support identifying, representing, and reasoning across a wide range of biomedical data. The Semantic Web Health Care and Life Sciences Interest Group (HCLSIG), set up within the framework of the World Wide Web Consortium, was launched to explore the application of these technologies in a variety of areas. Subgroups focus on making biomedical data available in RDF, working with biomedical ontologies, prototyping clinical decision support systems, working on drug safety and efficacy communication, and supporting disease researchers navigating and annotating the large amount of potentially relevant literature. RESULTS: We present a scenario that shows the value of the information environment the Semantic Web can support for aiding neuroscience researchers. We then report on several projects by members of the HCLSIG, in the process illustrating the range of Semantic Web technologies that have applications in areas of biomedicine. CONCLUSION: Semantic Web technologies present both promise and challenges. Current tools and standards are already adequate to implement components of the bench-to-bedside vision. On the other hand, these technologies are young. Gaps in standards and implementations still exist and adoption is limited by typical problems with early technology, such as the need for a critical mass of practitioners and installed base, and growing pains as the technology is scaled up. Still, the potential of interoperable knowledge sources for biomedicine, at the scale of the World Wide Web, merits continued work.
Alan Ruttenberg, Tim Clark, William J. Bug, Matthias Samwald, Olivier Bodenreider, Helen Chen, Donald Doherty, Kerstin Forsberg, Vipul Kashyap, June Kinoshita, Joanne S. Luciano, M. Scott Marshall, Chimezie Ogbuji, Jonathan Rees, Susie Stephens, Gwendolyn T. Wong, Elizabeth Wu, Davide Zaccagnini, Tonya Hongsermeier, Eric Neumann, Ivan Herman, Kei-Hoi Cheung
BMC Bioinform.5
2007 Experience in Aligning Anatomical Ontologies
abstract
An ontology is a formal representation of a domain modeling the entities in the domain and their relations. When a domain is represented by multiple ontologies, there is need for creating mappings among these ontologies in order to facilitate the integration of data annotated with these ontologies and reasoning across ontologies. The objective of this paper is to recapitulate our experience in aligning large anatomical ontologies and to reflect on some of the issues and challenges encountered along the way. The four anatomical ontologies under investigation are the Foundational Model of Anatomy, GALEN, the Adult Mouse Anatomical Dictionary and the NCI Thesaurus. Their underlying representation formalisms are all different. Our approach to aligning concepts (directly) is automatic, rule-based, and operates at the schema level, generating mostly point-to-point mappings. It uses a combination of domain-specific lexical techniques and structural and semantic techniques (to validate the mappings suggested lexically). It also takes advantage of domain-specific knowledge (lexical knowledge from external resources such as the Unified Medical Language System, as well as knowledge augmentation and inference techniques). In addition to point-to-point mapping of concepts, we present the alignment of relationships and the mapping of concepts group-to-group. We have also successfully tested an indirect alignment through a domain-specific reference ontology. We present an evaluation of our techniques, both against a gold standard established manually and against a generic schema matching system. The advantages and limitations of our approach are analyzed and discussed throughout the paper.
Songmao Zhang, Olivier Bodenreider
Int. J. Semantic Web Inf. Syst.2
2006 Comparing the Representation of Anatomy in the FMA and SNOMED CT
Olivier Bodenreider, Songmao Zhang
AMIA1
2006 Using WordNet to Improve the Mapping of Data Elements to UMLS for Data Sources Integration
Fleur Mougin, Anita Burgun-Parenthoine, Olivier Bodenreider
AMIA3
2006 Besides Precision & Recall: Exploring Alternative Approaches to Evaluating an Automatic Indexing Tool for MEDLINE
Aurélie Névéol, Kelly Zeng, Olivier Bodenreider
AMIA3
2006 Enhancing Biomedical Ontologies through Alignment of Semantic Relationships: Exploratory Approaches
Lowell Vizenor, Olivier Bodenreider, Lee B. Peters, Alexa T. McCray
AMIA2
2006 RxNav: A Web Service for Standard Drug Information
Kelly Zeng, Olivier Bodenreider, John Kilbourne, Stuart J. Nelson
AMIA2
2006 RxNav: Providing Standard Drug Information
Kelly Zeng, Olivier Bodenreider, John Kilbourne, Stuart J. Nelson
AMIA2
2006 Bio-ontologies: current trends and future directions
abstract
In recent years, as a knowledge-based discipline, bioinformatics has been made more computationally amenable. After its beginnings as a technology advocated by computer scientists to overcome problems of heterogeneity, ontology has been taken up by biologists themselves as a means to consistently annotate features from genotype to phenotype. In medical informatics, artifacts called ontologies have been used for a longer period of time to produce controlled lexicons for coding schemes. In this article, we review the current position in ontologies and how they have become institutionalized within biomedicine. As the field has matured, the much older philosophical aspects of ontology have come into play. With this and the institutionalization of ontology has come greater formality. We review this trend and what benefits it might bring to ontologies and their use within biomedicine.
Olivier Bodenreider, Robert Stevens 0001
Briefings Bioinform.1
2006 Mapping data elements to terminological resources for integrating biomedical data sources
abstract
BACKGROUND: Data integration is a crucial task in the biomedical domain and integrating data sources is one approach to integrating data. Data elements (DEs) in particular play an important role in data integration. We combine schema- and instance-based approaches to mapping DEs to terminological resources in order to facilitate data sources integration. METHODS: We extracted DEs from eleven disparate biomedical sources. We compared these DEs to concepts and/or terms in biomedical controlled vocabularies and to reference DEs. We also exploited DE values to disambiguate underspecified DEs and to identify additional mappings. RESULTS: 82.5% of the 474 DEs studied are mapped to entries of a terminological resource and 74.7% of the whole set can be associated with reference DEs. Only 6.6% of the DEs had values that could be semantically typed. CONCLUSION: Our study suggests that the integration of biomedical sources can be achieved automatically with limited precision and largely facilitated by mapping DEs to terminological resources.
Fleur Mougin, Anita Burgun-Parenthoine, Olivier Bodenreider
BMC Bioinform.3
2006 The foundational model of anatomy in OWL: Experience and perspectives
Christine Golbreich, Songmao Zhang, Olivier Bodenreider
J. Web Semant.3
2005 Of Mice and Men: Aligning Mouse and Human Anatomies
Olivier Bodenreider, Terry F. Hayamizu, Martin Ringwald, Sherri de Coronado, Songmao Zhang
AMIA1
2005 Classifying diseases with respect to anatomy: a study in SNOMED CT
Anita Burgun-Parenthoine, Olivier Bodenreider, Fleur Mougin
AMIA2
2005 Utilizing the UMLS for Semantic Mapping between Terminologies
Kin Wah Fung, Olivier Bodenreider
AMIA2
2005 Approaches to Eliminating Cycles in the UMLS Metathesaurus: Naïve vs. Formal
Fleur Mougin, Olivier Bodenreider
AMIA2
2005 RxNav: A Progress Report
Kelly Zeng, Olivier Bodenreider, Stuart J. Nelson
AMIA2
2005 Alignment of Multiple Ontologies of Anatomy: Deriving Indirect Mappings from Direct Mappings to a Reference
Songmao Zhang, Olivier Bodenreider
AMIA2
2005 An Ontology-Driven Clustering Method for Supporting Gene Expression Analysis
abstract
The gene ontology (GO) is an important knowledge resource for biologists and bioinformaticians. This paper explores the integration of similarity information derived from GO into clustering-based gene expression analysis. A system that integrates GO annotations, similarity patterns and expression data in yeast is assessed. In comparison with a clustering model based only on expression data correlation, the proposed framework not only produces consistent results, but also it offers alternative, potentially meaningful views of the biological problem under study. Moreover, it provides the basis for developing other automated, knowledge-driven data mining systems in this and related application areas.
Haiying Wang 0001, Francisco Azuaje, Olivier Bodenreider
CBMS3
2004 Incorporating Ontology-Driven Similarity Knowledge into Functional Genomics: An Exploratory Study
abstract
This research explores the feasibility of semantic similarity approaches to supporting predictive tasks in functional genomics. It aims to establish potential relationships between ontology-based similarity of gene products and important functional properties, such as gene expression correlation. Similarity measures based on the information content of the Gene Ontology (GO) were analyzed. Models have been implemented using data obtained from well-known studies in S. cerevisiae. Results suggest that there may exist significant relationships between gene expression correlation and semantic similarity. Analyses of protein complex data show that, in general, there is a significant correlation between the semantic similarity exhibited by a pair of genes and the probability of finding them in the same complex. These results can also be interpreted as an assessment of the quality and consistency of the information represented in the GO.
Francisco Azuaje, Olivier Bodenreider
BIBE2
2004 Gene expression correlation and gene ontology-based similarity: an assessment of quantitative relationships
abstract
Genome Database were analyzed to calculate functional similarity of gene products. Three methods for measuring similarity (including a distance-based approach) were implemented. Significant, quantitative relationships between similarity and expression correlation of pairs of genes were detected. Using a known gene expression dataset in yeast, this study compared more than three million pairs of gene products on the basis of these functional properties. Highly correlated genes exhibit strong similarity based on information originating from the gene ontology taxonomies. Such a similarity is significantly stronger than that observed between weakly correlated genes. This study supports the feasibility of applying gene ontology-driven similarity methods to functional prediction tasks, such as the validation of gene expression analyses and the identification of false positives in protein interaction studies.
Haiying Wang 0001, Francisco Azuaje, Olivier Bodenreider, Joaquín Dopazo
CIBCB3
2003 Strength in Numbers: Exploring Redundancy in Hierarchical Relations across Biomedical Terminologies
Olivier Bodenreider
AMIA1
2003 Graphical Visualization and Navigation of Genetic Disease Information
Olivier Bodenreider, Joyce A. Mitchell
AMIA1
2003 Aligning Representations of Anatomy using Lexical and Structural Methods
Songmao Zhang, Olivier Bodenreider
AMIA2
2003 Exploring semantic groups through visual approaches
Olivier Bodenreider, Alexa T. McCray
J. Biomed. Informatics1
2002 GenNav: Visualizing Gene Ontology as a Graph
Olivier Bodenreider
AMIA1
2002 Evaluation of the UMLS as a terminology and knowledge resource for biomedical informatics
Olivier Bodenreider, Joyce A. Mitchell, Alexa T. McCray
AMIA1
2002 Representation of roles in biomedical ontologies: a case study in functional genomics
Anita Burgun-Parenthoine, Olivier Bodenreider, Franck Le Duff, Fouzia Moussouni-Marzolf, Olivier Loréal
AMIA2
2002 The lexical properties of the gene ontology
Alexa T. McCray, Allen C. Browne, Olivier Bodenreider
AMIA3
2002 From Phenotype to Genotype: Experiences in Navigating the Available Information Resources
Joyce A. Mitchell, Alexa T. McCray, Olivier Bodenreider
AMIA3
2001 Circular hierarchical relationships in the UMLS: etiology, diagnosis, treatment, complications and prevention
Olivier Bodenreider
AMIA1
2001 Mapping the UMLS Semantic Network into general ontologies
Anita Burgun-Parenthoine, Olivier Bodenreider
AMIA2
2001 Evaluating UMLS strings for natural language processing
Alexa T. McCray, Olivier Bodenreider, James D. Malley, Allen C. Browne
AMIA2
2001 Aspects of the taxonomic relation in the biomedical domain
abstract
Taxonomies are commonly used for organizing knowledge, particularly in biomedicine where the taxonomy of living organisms and the classification of diseases are central to the domain. The principles used to produce taxonomies are either intrinsic (properties of the partial ordering relation) or added to make knowledge more manageable (opposition of siblings and economy). The applicability of these principles in the biomedical domain is presented using the Unified Medical Language System (UMLS) and issues raised by the application of these principles are illustrated. While intrinsic principles are not challenged, we argue that the opposition of siblings brings to bear excessive constraints on a domain ontology and that the adverse effects of economy may outweigh its benefits. The two-level structure used in the UMLS is discussed.
Anita Burgun-Parenthoine, Olivier Bodenreider
FOIS2
2000 The NLM Indexing Initiative
Alan R. Aronson, Olivier Bodenreider, Florence Chang, Susanne M. Humphrey, James G. Mork, Stuart J. Nelson, Thomas C. Rindflesch, W. John Wilbur
AMIA2
2000 Using UMLS semantics for classification purposes
Olivier Bodenreider
AMIA1
2000 A Semantic Navigation Tool for the UMLS
Olivier Bodenreider
AMIA1
1999 Automated Assignment of Medical Subject Headings
Stuart J. Nelson, Alan R. Aronson, Tamas E. Doszkocs, W. John Wilbur, Olivier Bodenreider, Florence Chang, James G. Mork, Alexa T. McCray
AMIA5
1998 Beyond synonymy: exploiting the UMLS semantics in mapping vocabularies
Olivier Bodenreider, Stuart J. Nelson, William T. Hole, Florence Chang
AMIA1
1998 Case Report: Evaluation of the Unified Medical Language System as a Medical Knowledge Source
abstract
OBJECTIVE: The authors evaluated the use of the Unified Medical Language System (UMLS) as a medical knowledge source for the representation of medical procedures in the MAOUSSC system. DESIGN: MAOUSSC, a multiaxial coding system, was used for the representation of 1500 procedures from 15 clinical specialties, using UMLS concepts (augmented by full sources for three new vocabularies being added to the UMLS) and relationships whenever possible. Evaluation criteria for the UMLS included (1) completeness of representation of concepts and of inter-concept relationships, (2) consistency in the categorization of both concepts and inter-concept relationships, and (3) usability, including adaptability of the UMLS to a foreign language (French), its suitability to a geographic region with different medical practices than the USA, and issues relative to the annual update changes in the test vocabularies. RESULTS: During the MAOUSSC trial, the number of missing concepts or relationships identified in the augmented UMLS sources was deemed to be inconsequential relative to overall project goals. "Missing" UMLS inter-concept relationships were identified, although they were small in number. Some inconsistencies in the UMLS were noted, especially in the area of hierarchic relationships. CONCLUSION: After UMLS was used for five years as a knowledge source for representing 1500 complex medical procedures in MAOUSSC, its value is considered significant. Future editions of the UMLS are expected to improve representation of inter-concept relationships and global consistency.
Olivier Bodenreider, Anita Burgun-Parenthoine, Geneviève Botti, Marius Fieschi, Pierre Le Beux, François Kohler
J. Am. Medical Informatics Assoc.1
1997 Application of Information Technology: A Web Terminology Server Using UMLS for the Description of Medical Procedures
abstract
The Model for Assistance in the Orientation of a User within Coding Systems (MAOUSSC) project has been designed to provide a representation for medical and surgical procedures that allows several applications to be developed from several viewpoints. It is based on a conceptual model, a controlled set of terms, and Web server development. The design includes the UMLS knowledge sources associated with additional knowledge about medico-surgical procedures. The model was implemented using a relational database. The authors developed a complete interface for the Web presentation, with the intermediary layer being written in PERL. The server has been used for the representation of medico-surgical procedures that occur in the discharge summaries of the national survey of hospital activities that is performed by the French Health Statistics Agency in order to produce inpatient profiles. The authors describe the current status of the MAOUSSC server and discuss their interest in using such a server to assist in the coordination of terminology tasks and in the sharing of controlled terminologies.
Anita Burgun-Parenthoine, Patrick Denier, Olivier Bodenreider, Geneviève Botti, Denis Delamarre, Bruno Pouliquen, Philippe Oberlin, Jean M. Lévéque, Bertrand Lukacs, François Kohler, Marius Fieschi, Pierre Le Beux
J. Am. Medical Informatics Assoc.3