EDBT 2026 Demo / reviewers in the wild / expert
James Geller
dblp:g/JamesGeller
· DBLP profile ↗
127ranked-venue papers
22as first author
10since 2021 · last 2026
0000-0002-9120-525XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 79 · 9 first-author · 7 since 2021Databases, data management, data science and information retrieval · 31 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 25 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Security and privacy · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VROOM: Evaluating Role-Based Collaboration for Immersive Ontology ManipulationabstractAn ontology is a structured representation of knowledge that organizes domain-specific concepts and their relationships. As ontologies scale, visualizing and manipulating them using traditional two-dimensional interfaces becomes increasingly challenging due to visual clutter and structural complexity. To address this, we developed a multi-user immersive ontology manipulation system that enables interaction with ontological structures in a three-dimensional space. This work-in-progress study adopts a usability and experience-focused perspective to examine how users navigate, coordinate, and perform ontology-related tasks in a shared immersive setting. We conducted an exploratory user study using qualitative and observational methods, in which participants completed visualization and authoring tasks in co-present, multi-user Virtual Reality (VR). Our analysis focuses on user strategies, interaction patterns, and breakdowns observed during task execution, supported by analysis of audio transcripts. Preliminary findings highlight challenges in spatial orientation, object visibility, and awareness of collaborators, as well as emergent collaborative behaviors, including role adaptation, peer support, and shared task planning. Participants also reported a strong sense of co-presence and engagement in the shared environment. These early results provide insights into the opportunities and limitations of immersive VR for collaborative knowledge work, and inform future directions for interaction design and evaluation in collaborative immersive analytics systems. Aayush Chitransh, Leandro Moises Paulino, Abdul Azeez Shaik, James Geller, Margarita Vinnikov |
IMX | 4 |
| 2026 | Opportunities for informatics to improve patient experiences: observations and reflections of ACMI fellowsabstractOBJECTIVES: We report on findings from a meeting convened by the American College of Medical Informatics (ACMI) to characterize aspects of the patient experience that could be improved using informatics. MATERIALS AND METHODS: The American College of Medical Informatics fellows were invited to share their experiences as patients and suggest informatics approaches that may improve the patient experience. RESULTS: We identified 4 themes: (1) getting the right care, (2) data sharing and data interoperability, (3) guiding low-cost evaluations, and (4) predictive analytics. DISCUSSION: Despite widespread adoption of health IT, patient experiences remain far from optimal. CONCLUSION: The American College of Medical Informatics fellows identified informatics approaches, applications, and research areas that have the potential to improve patient experiences with health care systems. Howard R. Strasberg, Edward P. Hoffer, Ross Koppel, Kevin B. Johnson, William M. Tierney, Geoffrey W. Rutledge, Elmer V. Bernstam, Jos Aarts, Marion J. Ball, Douglas S. Bell, Bernd Blobel, Suzanne Boren, Iain E. Buchan, James J. Cimino, Lawrence M. Fagan, James Geller, María Adela Grando, David A. Hanauer, William R. Hogan, Andrew S. Kanter, Bonnie Kaplan, Casimir A. Kulikowski, Albert Lai, David McCallie, Vimla Patel, Wanda Pratt, Sarah Collins Rossetti, Edward H. Shortliffe, Hardeep Singh 0005, Dean F. Sittig, William W. Stead, Kim M. Unertl, Mark G. Weiner, Kai Zheng 0002 |
J. Am. Medical Informatics Assoc. | 16 |
| 2025 | Ontology enrichment using a large language model: Applying lexical, semantic, and knowledge network-based similarity for concept placementabstractOBJECTIVE: Ontologies are essential for representing the knowledge of a domain. To make ontologies useful, they must encompass a comprehensive domain view. To achieve ontology enrichment, there is a need to discover new concepts to be added, either because they were missed in the first place, or the state-of-the-art has advanced to develop new real-world concepts. Our goal is to develop an automatic enrichment pipeline using a seed ontology, a Large Language Model (LLM), and source of text. The pipeline is applied to the domain of Social Determinants of Health (SDoH), using PubMed as a source of concepts. In this work, the applicability and effectiveness of the enrichment pipeline is demonstrated by extending the SDoH Ontology called SOHOv1, however our methodology could be used in other domains as well. METHODS: We first retrieved PubMed abstracts of candidate articles with existing SOHOv1 concepts as search terms. Next, we used GPT-4-1201 to extract semantic triples from the abstracts. We identified concepts from these triples utilizing lexical, semantic, and knowledge network-based filtering. We also compared the granularity of semantic triples extracted with our method to the triples in the SemMedDB (Semantic MEDLINE Database). The results were evaluated by human experts and standard ontology tools for checking consistency and semantic correctness. RESULTS: We expanded SOHOv1, which contained 173 concepts and 585 axioms, including 207 logical axioms to SOHOv2, which contains 572 concepts, 1,542 axioms, including 725 logical axioms. Our methods identified more concepts than those extracted from SemMedDB for the same task. While we have shown the feasibility of our approach for an SDoH ontology, the methodology is generalizable to other ontologies with an existing seed ontology and text corpus. CONCLUSIONS: The contributions of this work are: Extracting semantic triples from PubMed abstracts using GPT-4-1201 utilizing prompt chaining; showing the superiority of triples from GPT-4-1201 over triples from SemMedDB for SDoH; using lexical and semantic similarity search techniques with knowledge network-based search to identify the concepts to be added to the ontology; confirming the quality of the new concepts with human experts. Navya Martin Kollapally, James Geller, Vipina Kuttichi Keloth, Zhe He 0001, Julia Xu |
J. Biomed. Informatics | 2 |
| 2024 | Using clinical entity recognition for curating an interface terminology to aid fast skimming of EHRsabstractHighlighting of Electronic Health Records (EHRs) involves marking essential content of EHR notes, corresponding to concepts of a clinical terminology. However, employing the best clinical terminology (SNOMED CT) for highlighting EHRs, captures only a portion of their crucial content. In this paper, we describe the curation of a Cardiology Interface Terminology (CIT) dedicated to the application of highlighting EHRs of cardiology patients. We utilize a Clinical-Named Entity Recognition (Clinical NER) approach for extracting phrases, of higher granularity than SNOMED CT concepts, from EHRs, for enriching CIT. For this purpose, we train a neural network model with BIOE-tagged (Beginning, Inside, End, and Outside) cardiology entities. Transfer Learning can be used to facilitate the curation of an interface terminology for highlighting EHRs for other specialties e.g. Nephrology. Large-scale highlighting enables overworked physicians and other healthcare providers to fast skim the dense volume of EHRs they regularly read. Secondary research and EHRs interoperability are other applications that can be supported by highlighting. Navya Martin Kollapally, Mahshad Koohi Habibi Dehkordi, Yehoshua Perl, James Geller, Fadi P. Deek, Hao Liu 0025, Vipina Kuttichi Keloth, Gai Elhanan, Andrew J. Einstein, Shuxin Zhou |
BIBM | 4 |
| 2023 | Using annotation for computerized support for fast skimming of cardiology electronic health record notesabstractUnder the circumstances prevalent in current healthcare, medical professionals such as physicians and nurses need to read large numbers of Electronic Health Record (EHR) notes. This demand fosters a situation in which providers typically do not read a whole note but quickly skim it, to capture its essential content. Consequently, by fast skimming, one may miss a critical medical fact. Annotation highlights important content in EHR notes, and enables healthcare professionals to perform fast skimming, thereby minimizing the risk of missing critical information, which is detrimental to patient care. We designed the Cardiology Interface Terminology (CIT) for the purpose of annotation of cardiology EHRs. We emphasize that by annotation we refer to highlighting important information in the EHR notes for enabling fast skimming rather than just recognizing names of diseases, drugs, etc. from Reference Terminologies as is usually done by existing Named Entity Recognition (NER) systems. The CIT design starts with the cardiology components of SNOMED CT. It is enhanced by mining phrases from cardiology EHRs, as potential CIT concepts, which are of higher granularity than SNOMED concepts. Machine learning (ML), the state of the art technique for mining concepts from EHRs, requires training data. However, there is no training data for designing CIT. In the first stage, we introduce an innovative semi-automatic method for mining concepts from EHRs, to replace costly manual mining. The only manual portion is the review of the automatically mined phrases, before their insertion as CIT concepts. The effectiveness of annotation of cardiology EHRs with CIT was evaluated utilizing proper metrics, and compared to annotation with SNOMED CT. In a future second stage, ML mining techniques will be used for enhancing CIT with extra concepts from EHRs, utilizing the concepts added in the first stage as training data. This work focuses on a novel semi-automated method to design the Cardiology Interface Terminology (CIT) for annotation of cardiology EHRs to support fast skimming of EHR notes. Similar interface terminologies for other medical specialties could be obtained from CIT using Transfer Learning. Mahshad Koohi Habibi Dehkordi, Andrew J. Einstein, Shuxin Zhou, Gai Elhanan, Yehoshua Perl, Vipina Kuttichi Keloth, James Geller, Hao Liu 0025 |
BIBM | 7 |
| 2022 | An Ontology for the Social Determinants of Health DomainabstractSocial Determinants of Health (SDOH) are societal factors, such as where a person was born, grew up, works, lives, etc., along with socio-economic and community factors that affect an individual’s health. SDOH are correlated with many clinical outcomes, hence it is desirable to record SDOH data in Electronic Health Records (EHRs). Besides storing images, text, etc., EHRs rely on coded terms available in standard ontologies and terminologies to record observations and analyses. There is a substantial amount of research on understanding the clinical impact of SDOH, ranging from screening tools to practice-based interventions. However, there is no comprehensive collection of terms for recording SDOH observations in EHRs. Our research goal is to develop an ontology that covers the terms describing SDOH. We present a prototype ontology called Social Determinant of Health Ontology (SOHO) that covers relevant concepts and IS-A relationships describing impacts and associations of social determinants. We describe the evaluation techniques that we applied to SOHO, including human experts’ review and algorithmic evaluation. Navya Martin Kollapally, Yan Chen 0009, Julia Xu, James Geller |
BIBM | 4 |
| 2021 | Detecting, Reporting And Alleviating Racial Biases In Standardized Medical Terminologies And OntologiesabstractRecently, the issue has been raised that personal and systemic biases in organizations, such as some police departments, have also been detected in healthcare organizations. Furthermore, victims of bias incidents often end up in the healthcare system for treatment. Providers use standardized terminologies to record the status of patients in EHRs. To accurately record patient data, these terminologies must contain all the terms that a healthcare provider needs, including terms that might be race-, ethnicity-, or gender-specific. Following reports about gaps in terminologies, we investigated the coverage with respect to such terms in major terminologies such as SNOMED CT, ICD-10, CPT, NCIt and MedDRA. To identify potentially missing terms, we drew on public databases and news articles describing incidents that resulted in minority members requiring medical attention after police interventions. We posit those terms should be added into medical terminologies to improve the ability to record incidents happening inside and outside of the healthcare system. James Geller, Navya Martin Kollapally |
BIBM | 1 |
| 2021 | Knowledge Graph Analysis of Russian Trolls
Alen Chih-Yuan Li, Soon Ae Chun, James Geller |
DATA | 3 |
| 2021 | Health Ontology for Minority Equity (HOME)
Navya Martin Kollapally, Yan Chen 0009, James Geller |
KEOD | 3 |
| 2021 | Visual comprehension and orientation into the COVID-19 CIDO ontology
Yehoshua Perl, Yongqun He, Christopher Ochs, James Geller, Hao Liu 0025, Vipina Kuttichi Keloth |
J. Biomed. Informatics | 5 |
| 2020 | Desiderata for High Quality AMIA Presentation Files
James Geller, Yarou Ding |
AMIA | 1 |
| 2020 | Generating Training Data for Concept-Mining for an 'Interface Terminology' Annotating Cardiology EHRsabstractClinical data stored in EHRs could provide valuable knowledge for research if it were annotated properly. However, almost no EHR notes are currently annotated as the performance of off the shelf annotation tools is unsatisfactory. Concentrating on the cardiology specialty, we propose to design a Cardiology Interface Terminology dedicated to the annotation of EHR notes in cardiology. This interface terminology will be developed by the addition of high granularity concepts, mined from cardiology EHR notes, to an initial version reusing SNOMED CT cardiology subhierarchies. Using text mining NLP tools with machine learning for extending this interface terminology requires proper training data. In this paper, we discuss concept-mining of EHR notes, using concatenation and anchoring operations iteratively to create such training data. This approach can be applied to other medical specialties. Vipina Kuttichi Keloth, Shuxin Zhou, Andrew J. Einstein, Gai Elhanan, Yan Chen 0009, James Geller, Yehoshua Perl |
BIBM | 6 |
| 2020 | Mining Concepts for a COVID Interface Terminology for Annotation of EHRsabstractThe COVID-19 pandemic has overwhelmed the healthcare services of many countries with increased number of patients and also with a deluge of medical data. Furthermore, the emergence and global spread of new infectious diseases are highly likely to continue in the future. Incomplete data about presentations, signs, and symptoms of COVID-19 has had adverse effects on healthcare delivery. The EHRs of US hospitals have ingested huge volumes of relevant, up-to-date data about patients, but the lack of a proper system to annotate this data has greatly reduced its usefulness. We propose to design a COVID interface terminology for the annotation of EHR notes of COVID-19 patients. The initial version of this interface terminology was created by integrating COVID concepts from existing ontologies. Further enrichment of the interface terminology is performed by mining high granularity concepts from EHRs, because such concepts are usually not present in the existing reference terminologies. We use the techniques of concatenation and anchoring iteratively to extract high granularity phrases from the clinical text. In addition to increasing the conceptual base of the COVID interface terminology, this will also help in generating training data for large scale concept mining using machine learning techniques. Having the annotated clinical notes of COVID-19 patients available will help in speeding up research in this field. Vipina Kuttichi Keloth, Shuxin Zhou, Luke Lindemann, Gai Elhanan, Andrew J. Einstein, James Geller, Yehoshua Perl |
IEEE BigData | 6 |
| 2020 | Concept placement using BERT trained by transforming and summarizing biomedical ontology structure
Hao Liu 0025, Yehoshua Perl, James Geller |
J. Biomed. Informatics | 3 |
| 2019 | Transfer Learning from BERT to Support Insertion of New Concepts into SNOMED CT
Hao Liu 0025, Yehoshua Perl, James Geller |
AMIA | 3 |
| 2019 | Training Convolutional Neural Network with Terminology Summarization Data Improves SNOMED CT Enrichment
Hao Liu 0025, Yehoshua Perl, James Geller |
AMIA | 4 |
| 2019 | Measuring and Avoiding Information Loss During Concept Import from a Source to a Target OntologyabstractComparing pairs of ontologies in the same biomedical content domain often uncovers surprising differences. In many cases these differences can be characterized as “density differences,” where one ontology describes the content domain with more concepts in a more detailed manner. Using the Unified Medical Language System across pairs of ontologies contained in it, these differences can be precisely observed and used as the basis for importing concepts from the ontology of higher density into the ontology of lower density. However, such an import can lead to an intuitive loss of information that is hard to formalize. This paper proposes an approach based on information theory that mathematically distinguishes between different methods of concept import and measures the associated avoidance of information loss. James Geller, Shmuel Tomi Klein, Vipina Kuttichi Keloth |
KEOD | 1 |
| 2019 | Detecting Political Bias Trolls in Twitter DataabstractEver since Russian trolls have been brought to light, their interference in the 2016 US Presidential elections has been monitored and studied. These Russian trolls employ fake accounts registered on several major social media sites to influence public opinion in other countries. Our work involves discovering patterns in these tweets and classifying them by training different machine learning models such as Support Vector Machines, Word2vec, Google BERT, and neural network models, and then applying them to several large Twitter datasets to compare the effectiveness of the different models. Two classification tasks are utilized for this purpose. The first one is used to classify any given tweet as either troll or non-troll tweet. The second model classifies specific tweets as coming from left trolls or right trolls, based on apparent extreme political orientations. On the given data sets, Google BERT provides the best results, with an accuracy of 89.4% for the left/right troll detector and 99% for the troll/non-troll detector. Temporal, geographic, and sentiment analyses were also performed and results were visualized. Soon Ae Chun, Richard D. Holowczak, Kannan Dharan, Ruoyu Wang 0013, Soumyadeep Basu, James Geller |
WEBIST | 6 |
| 2019 | Alternative classification of identical concepts in different terminologies: Different ways to view the world
Vipina Kuttichi Keloth, Zhe He 0001, Gai Elhanan, James Geller |
J. Biomed. Informatics | 4 |
| 2018 | How Sustainable are Biomedical Ontologies?
James Geller, Vipina Kuttichi Keloth, Mark A. Musen |
AMIA | 1 |
| 2018 | Leveraging Horizontal Density Differences between Ontologies to Identify Missing Child Concepts: A Proof of Concept
Vipina Kuttichi Keloth, Zhe He 0001, Yan Chen 0009, James Geller |
AMIA | 4 |
| 2018 | Using Convolutional Neural Networks to Support Insertion of New Concepts into SNOMED CT
Hao Liu 0025, James Geller, Michael Halper, Yehoshua Perl |
AMIA | 2 |
| 2018 | Overlapping Complex Concepts Have More Commission Errors, Especially in Intensive Terminology Auditing
Hao Liu 0025, Yehoshua Perl, James Geller, Christopher Ochs, James T. Case |
AMIA | 4 |
| 2018 | Extended Analysis of Topological-Pattern-Based Ontology Enrichment
Zhe He 0001, Vipina Kuttichi Keloth, Yan Chen 0009, James Geller |
BIBM | 4 |
| 2018 | Enrichment of SNOMED CT Ophthalmology Component to Support EHR Coding
Hao Liu 0025, P. Lloyd Hildebrand, Yehoshua Perl, James Geller |
BIBM | 4 |
| 2018 | Quality Assurance of Concept Roles in the National Cancer Institute thesaurus
Yan Chen 0009, Yehoshua Perl, Michael Halper, James Geller, Sherri de Coronado |
BIBM | 5 |
| 2018 | Quality assurance of biomedical terminologies and ontologies
James Geller, Yehoshua Perl, Licong Cui, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 1 |
| 2018 | Complex overlapping concepts: An effective auditing methodology for families of similarly structured BioPortal ontologies
Yan Chen 0009, Gai Elhanan, Yehoshua Perl, James Geller, Christopher Ochs |
J. Biomed. Informatics | 5 |
| 2017 | Applications of the Ontology Abstraction Framework
Christopher Ochs, James Geller, Yehoshua Perl |
AMIA | 2 |
| 2017 | Auditing the assignments of top-level semantic types in the UMLS semantic network to UMLS conceptsabstractThe Unified Medical Language System (UMLS) is an important terminological system. By the policy of its curators, each concept of the UMLS should be assigned the most specific Semantic Types (STs) in the UMLS Semantic Network (SN). Hence, the Semantic Types of most UMLS concepts are assigned at or near the bottom (leaves) of the UMLS Semantic Network. While most ST assignments are correct, some errors do occur. Therefore, Quality Assurance efforts of UMLS curators for ST assignments should concentrate on automatically detected sets of UMLS concepts with higher error rates than random sets. In this paper, we investigate the assignments of top-level semantic types in the UMLS semantic network to concepts, identify potential erroneous assignments, define four categories of errors, and thus provide assistance to curators of the UMLS to avoid these assignments errors. Human experts analyzed samples of concepts assigned 10 of the top-level semantic types and categorized the erroneous ST assignments into these four logical categories. Two thirds of the concepts assigned these 10 top-level semantic types are erroneous. Our results demonstrate that reviewing top-level semantic type assignments to concepts provides an effective way for UMLS quality assurance, comparing to reviewing a random selection of semantic type assignments. Zhe He 0001, Yehoshua Perl, Gai Elhanan, Yan Chen 0009, James Geller, Jiang Bian 0001 |
BIBM | 5 |
| 2017 | Discovering additional complex NCIt gene concepts with high error rateabstractThe Gene hierarchy of the National Cancer Institute (NCI) Thesaurus (NCIt) is of high priority for NCI. It is important to have quality assurance (QA) techniques to improve its content quality. We present a two-step methodology concentrating on auditing the modeling of complex concepts, which are shown to have a higher error rate compared to control concepts. In the first step, we test whether concepts that appear complex in a so called “partial-area taxonomy” have a higher error rate than control concepts. In the second step, we introduce an innovative technique based on a “partial-area sub-taxonomy” (constructed with a subset of roles) to discover additional complex concepts. The results of the QA study show that these concepts are indeed statistically significantly more likely to have more errors than control concepts. This makes it easier for NCI staff to improve the modeling quality of gene concepts in NCIt. Hua Min, Yehoshua Perl, James Geller |
BIBM | 4 |
| 2017 | Enabling Real-Time Drug Abuse Detection in TweetsabstractPrescription drug abuse is one of the fastest growing public health problems in the USA. To address this epidemic, a near real-time monitoring strategy, instead of one resorting to a retrospective health records, may improve detecting the prevalence and patterns of abuse of both illegal drugs and prescription medications. In this paper, our primary goals are to demonstrate the possibility of utilizing social media, e.g., Twitter, for automatic monitoring of illegal drug and prescription medication abuse. We use machine learning methods for an automatic classification that can identify tweets that are indicative of drug abuse. We collected tweets associated with well-known illegal and prescription drugs. We manually annotated 300 tweets that are likely to be related to drug abuse. Our experiment compares a set of classification algorithms, and a decision tree classifier J48, and the SVM outperform others for determining whether tweets contain signals of drug abuse. This automatic supervised classification study results illustrate the utility of Twitter in examining patterns of abuse, and show the feasibility of building the drug abuse detection system that can process large volume data from social media sources in a near real-time. NhatHai Phan, Soon Ae Chun, Manasi Bhole, James Geller |
ICDE | 4 |
| 2017 | An empirical analysis of ontology reuse in BioPortal
Christopher Ochs, Yehoshua Perl, James Geller, Sivaram Arabandi, Tania Tudorache, Mark A. Musen |
J. Biomed. Informatics | 3 |
| 2017 | Quality assurance of chemical ingredient classification for the National Drug File - Reference Terminology
Hasan Yumak, Ling Chen 0007, Christopher Ochs, James Geller, Joan Kapusnik-Uner, Yehoshua Perl |
J. Biomed. Informatics | 5 |
| 2016 | Topological-Pattern-Based Recommendation of UMLS Concepts for National Cancer Institute Thesaurus
Zhe He 0001, Yan Chen 0009, Sherri de Coronado, Katrina Piskorski, James Geller |
AMIA | 5 |
| 2016 | A unified software framework for deriving, visualizing, and exploring abstraction networks for ontologies
Christopher Ochs, James Geller, Yehoshua Perl, Mark A. Musen |
J. Biomed. Informatics | 2 |
| 2016 | Utilizing a structural meta-ontology for family-based quality assurance of the BioPortal ontologies
Christopher Ochs, Zhe He 0001, James Geller, Yehoshua Perl, George Hripcsak, Mark A. Musen |
J. Biomed. Informatics | 4 |
| 2015 | Drug-drug Interaction Discovery Using Abstraction Networks for "National Drug File - Reference Terminology" Chemical Ingredients
Christopher Ochs, Huanying Gu, Yehoshua Perl, James Geller, Joan Kapusnik-Uner, Aleksandr Zakharchenko |
AMIA | 5 |
| 2015 | Collaborative and trajectory prediction models of medical conditions by mining patients' Social DataabstractIn the U.S., 80% of Medicare spending is for managing patients with multiple coexisting conditions. Predicting potentially correlated diseases for an individual patient and correlated disease progression paths are both important research tasks. For example, obese patients are at an increased risk for developing type-2 diabetes and hypertension. This correlation is called comorbidity relationship. Discovering the comorbidity relationships is complex and difficult due to the limited access to Electronic Health Records by privacy laws. In this paper, we present a framework called Social Data-based Prediction of Incidence and Trajectory to predict potential risks for medical conditions as well as its progression trajectory to identify the comorbidity path. The framework utilizes patients' publicly available social media data and presents a collaborative prediction model to predict the ranked list of potential comorbidity incidences, and a trajectory prediction model to reveal different paths of condition progression. The experimental results show that our framework is able to predict future conditions for online patients with a coverage value of 48% and 75% for a top-20 and a top-100 ranked list, respectively. For risk trajectory prediction, our framework is able to reveal each potential progression trajectory between any two conditions and infer the confidence of the future trajectory, given any observed condition. The predicted trajectories are validated with existing comorbidity relations from the medical literature. Xiang Ji 0004, Soon Ae Chun, James Geller, Vincent Oria |
BIBM | 3 |
| 2015 | Using aggregate taxonomies to summarize SNOMED CT evolutionabstractTerminologies are typically large and complex knowledge systems. It is difficult to obtain an orientation into their structure and content. In previous research we designed compact summary networks called partial-area taxonomies to provide a structural summary of a terminology. The sizes of a terminology and of its partial-area taxonomy are defined as their numbers of nodes. While a partial-area taxonomy is typically smaller than the original terminology, it is often not compact enough to provide a clear “big picture,” due to too many nodes that summarize only a small number of terminology concepts. The display of such a partial-area taxonomy is still overwhelming. In this paper, we introduce a more compact summary of a terminology, called an aggregate taxonomy, obtained by aggregating small partial-area taxonomy nodes into larger nodes. We present a parametrized technique to study the design of such an aggregate taxonomy and apply it to the Specimen hierarchy of SNOMED CT. A software tool for creating and displaying aggregate taxonomies is described. We illustrate how aggregate taxonomies derived across multiple SNOMED CT releases can be used to summarize the evolution of the Specimen hierarchy's content over eight years of SNOMED CT releases. Christopher Ochs, Yehoshua Perl, James Geller, Mark A. Musen |
BIBM | 3 |
| 2015 | Identifying Pairs of Terms with Strong Semantic Connections in a Textbook IndexabstractSemantic relationships are important components of ontologies. Specifying these relationships is work-intensive and error-prone when done by experts. Discovering domain concepts and strongly related pairs of concepts in a completely automated way from English text is an unresolved problem. This paper uses index terms from a textbook as domain concepts and suggests pairs of concepts that are likely to be connected by strong semantic relationships. Two textbooks on Cyber Security were used as testbeds. To show the generality of the approach, the index terms from one of the books were used to generate suggestions for where to place semantic relationships using the bodies of both textbooks. A good overlap was found. James Geller, Shmuel Tomi Klein, Yuriy Polyakov |
KEOD | 1 |
| 2015 | Leveraging Social Data for Health Care Behavior Analytics
Xiang Ji 0004, Paolo Cappellari, Soon Ae Chun, James Geller |
ICWE | 4 |
| 2015 | A comparative analysis of the density of the SNOMED CT conceptual content for semantic harmonization
Zhe He 0001, James Geller, Yan Chen 0009 |
Artif. Intell. Medicine | 2 |
| 2015 | A tribal abstraction network for SNOMED CT target hierarchies without attribute relationshipsabstractOBJECTIVE: Large and complex terminologies, such as Systematized Nomenclature of Medicine-Clinical Terms (SNOMED CT), are prone to errors and inconsistencies. Abstraction networks are compact summarizations of the content and structure of a terminology. Abstraction networks have been shown to support terminology quality assurance. In this paper, we introduce an abstraction network derivation methodology which can be applied to SNOMED CT target hierarchies whose classes are defined using only hierarchical relationships (ie, without attribute relationships) and similar description-logic-based terminologies. METHODS: We introduce the tribal abstraction network (TAN), based on the notion of a tribe-a subhierarchy rooted at a child of a hierarchy root, assuming only the existence of concepts with multiple parents. The TAN summarizes a hierarchy that does not have attribute relationships using sets of concepts, called tribal units that belong to exactly the same multiple tribes. Tribal units are further divided into refined tribal units which contain closely related concepts. A quality assurance methodology that utilizes TAN summarizations is introduced. RESULTS: A TAN is derived for the Observable entity hierarchy of SNOMED CT, summarizing its content. A TAN-based quality assurance review of the concepts of the hierarchy is performed, and erroneous concepts are shown to appear more frequently in large refined tribal units than in small refined tribal units. Furthermore, more erroneous concepts appear in large refined tribal units of more tribes than of fewer tribes. CONCLUSIONS: In this paper we introduce the TAN for summarizing SNOMED CT target hierarchies. A TAN was derived for the Observable entity hierarchy of SNOMED CT. A quality assurance methodology utilizing the TAN was introduced and demonstrated. Christopher Ochs, James Geller, Yehoshua Perl, Yan Chen 0009, Ankur Agrawal, James T. Case, George Hripcsak |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Scalable quality assurance for large SNOMED CT hierarchies using subject-based subtaxonomiesabstractOBJECTIVE: Standards terminologies may be large and complex, making their quality assurance challenging. Some terminology quality assurance (TQA) methodologies are based on abstraction networks (AbNs), compact terminology summaries. We have tested AbNs and the performance of related TQA methodologies on small terminology hierarchies. However, some standards terminologies, for example, SNOMED, are composed of very large hierarchies. Scaling AbN TQA techniques to such hierarchies poses a significant challenge. We present a scalable subject-based approach for AbN TQA. METHODS: An innovative technique is presented for scaling TQA by creating a new kind of subject-based AbN called a subtaxonomy for large hierarchies. New hypotheses about concentrations of erroneous concepts within the AbN are introduced to guide scalable TQA. RESULTS: We test the TQA methodology for a subject-based subtaxonomy for the Bleeding subhierarchy in SNOMED's large Clinical finding hierarchy. To test the error concentration hypotheses, three domain experts reviewed a sample of 300 concepts. A consensus-based evaluation identified 87 erroneous concepts. The subtaxonomy-based TQA methodology was shown to uncover statistically significantly more erroneous concepts when compared to a control sample. DISCUSSION: The scalability of TQA methodologies is a challenge for large standards systems like SNOMED. We demonstrated innovative subject-based TQA techniques by identifying groups of concepts with a higher likelihood of having errors within the subtaxonomy. Scalability is achieved by reviewing a large hierarchy by subject. CONCLUSIONS: An innovative methodology for scaling the derivation of AbNs and a TQA methodology was shown to perform successfully for the largest hierarchy of SNOMED. Christopher Ochs, James Geller, Yehoshua Perl, Yan Chen 0009, Junchuan Xu, Hua Min, James T. Case, Zhi Wei 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2015 | Summarizing and visualizing structural changes during the evolution of biomedical ontologies using a Diff Abstraction Network
Christopher Ochs, Yehoshua Perl, James Geller, Melissa A. Haendel, Matthew H. Brush, Sivaram Arabandi, Samson W. Tu |
J. Biomed. Informatics | 3 |
| 2014 | Social InfoButtons for Patient-oriented Healthcare Knowledge Support
Xiang Ji 0004, James Geller, Soon Ae Chun |
AMIA | 2 |
| 2014 | A Hybrid Approach to Developing a Cyber Security OntologyabstractThe process of developing an ontology cannot be fully automated at the current state-of-the-art. However, leaving the tedious, time-consuming and error-prone task of ontology development entirely to humans has problems of its own, including limited staff budgets and semantic disagreements between experts. Thus, a hybrid computer/expert approach is advocated. The research challenge is how to minimize and optimally organize the task of the expert(s) while maximally leveraging the power of the computer and of existing computer-readable documents. The purpose of this paper is two-fold. First we present such a hybrid approach by describing a knowledge acquisition tool that we have developed. This tool makes use of an existing Bootstrap Ontology and proposes likely locations of concepts and semantic relationships, based on a text book, to a domain expert who can decide on them. The tool is attempting to minimize the number of interactions. Secondly we are proposing the notion of an augmented ontology specifically for pedagogical use. The application domain of this work is cyber-security education, but the ontology development methods are applicable to any educational topic. James Geller, Soon Ae Chun, Arwa M. Wali |
DATA | 1 |
| 2014 | Mapping biological entities using the Longest Approximately Common Prefix methodabstractBACKGROUND: The significant growth in the volume of electronic biomedical data in recent decades has pointed to the need for approximate string matching algorithms that can expedite tasks such as named entity recognition, duplicate detection, terminology integration, and spelling correction. The task of source integration in the Unified Medical Language System (UMLS) requires considerable expert effort despite the presence of various computational tools. This problem warrants the search for a new method for approximate string matching and its UMLS-based evaluation. RESULTS: This paper introduces the Longest Approximately Common Prefix (LACP) method as an algorithm for approximate string matching that runs in linear time. We compare the LACP method for performance, precision and speed to nine other well-known string matching algorithms. As test data, we use two multiple-source samples from the Unified Medical Language System (UMLS) and two SNOMED Clinical Terms-based samples. In addition, we present a spell checker based on the LACP method. CONCLUSIONS: The Longest Approximately Common Prefix method completes its string similarity evaluations in less time than all nine string similarity methods used for comparison. The Longest Approximately Common Prefix outperforms these nine approximate string matching methods in its Maximum F1 measure when evaluated on three out of the four datasets, and in its average precision on two of the four datasets. Alex Rudniy, Min Song 0001, James Geller |
BMC Bioinform. | 3 |
| 2013 | A Bootstrapping Approach for Developing a Cyber-security Ontology Using Textbook Index TermsabstractDeveloping a domain ontology with concepts and relationships between them is a challenge, since knowledge engineering is a labor intensive process that can be a bottleneck and is often not scalable. Developing a cyber-security ontology is no exception. A security ontology can improve search for security learning resources that are scattered in different locations in different formats, since it can provide a common controlled vocabulary to annotate the resources with consistent semantics. In this paper, we present a bootstrapping method for developing a cyber-security ontology using both a security textbook index that provides a list of terms in the security domain and an existing security ontology as a scaffold. The bootstrapping approach automatically extracts the textbook index terms (concepts), derives a relationship to a concept in the security ontology for each and classifies them into the existing security ontology. The bootstrapping approach relies on the exact and approximate similarity matching of concepts as well as the category information obtained from external sources such as Wikipedia. The results show feasibility of our method to develop a more comprehensive and scalable cyber-security ontology with rich concepts from a textbook index. We provide criteria used to select a scaffold ontology among existing ontologies. The current approach can be improved by considering synonyms, deep searching in Wikipedia categories, and domain expert validation. Arwa M. Wali, Soon Ae Chun, James Geller |
ARES | 3 |
| 2013 | A Family-Based Framework for Supporting Quality Assurance of Biomedical Ontologies in BioPortal
Zhe He 0001, Christopher Ochs, Ankur Agrawal, Yehoshua Perl, Dimitris Zeginis, Konstantinos A. Tarabanis, Gai Elhanan, Michael Halper, Natasha F. Noy, James Geller |
AMIA | 10 |
| 2013 | Scalability of Abstraction-Network-Based Quality Assurance to Large SNOMED Hierarchies
Christopher Ochs, Yehoshua Perl, James Geller, Michael Halper, Huanying Gu, Yan Chen 0009, Gai Elhanan |
AMIA | 3 |
| 2013 | Rule-based support system for multiple UMLS semantic type assignments
James Geller, Zhe He 0001, Yehoshua Perl, C. Paul Morrey, Julia Xu |
J. Biomed. Informatics | 1 |
| 2012 | Using Gamification and Crowdsourcing to Enhance Terminology Auditing
David Daudelin, James Geller, Yehoshua Perl |
AMIA | 2 |
| 2012 | New Abstraction Networks and a New Visualization Tool in Support of Auditing the SNOMED CT Content
James Geller, Christopher Ochs, Yehoshua Perl, Junchuan Xu |
AMIA | 1 |
| 2012 | Deriving an Abstraction Network to Support Quality Assurance in OCRe
Christopher Ochs, Ankur Agrawal, Yehoshua Perl, Michael Halper, Samson W. Tu, Simona Carini, Ida Sim, Natasha F. Noy, Mark A. Musen, James Geller |
AMIA | 10 |
| 2012 | Overcoming an obstacle in expanding a UMLS semantic type extent
Yan Chen 0009, Huanying Gu, Yehoshua Perl, James Geller |
J. Biomed. Informatics | 4 |
| 2012 | A study of terminology auditors' performance for UMLS semantic type assignments
Huanying Gu, Gai Elhanan, Yehoshua Perl, George Hripcsak, James J. Cimino, Julia Xu, Yan Chen 0009, James Geller, C. Paul Morrey |
J. Biomed. Informatics | 8 |
| 2012 | Abstraction of complex concepts with a refined partial-area taxonomy of SNOMED
Yue Wang 0033, Michael Halper, Duo Helen Wei, Yehoshua Perl, James Geller |
J. Biomed. Informatics | 5 |
| 2011 | A survey of SNOMED CT direct users, 2010: impressions and preferences regarding content and qualityabstractOBJECTIVE: Little information exists concerning SNOMED CT (systematized nomenclature of medicine-clinical terms) users. This report describes current impressions and preferences of direct SNOMED CT users regarding coverage, quality, and concept details, and the change request mechanism. DESIGN: A 43-question anonymous survey distributed electronically to relevant online communities. MEASUREMENTS: Data on user demographic characteristics, modes and purposes of use, means and frequencies of access, satisfaction with SNOMED CT content coverage and quality and with the change request mechanism were recorded. RESULTS: The survey was conducted in January 2010 and elicited 215 responses. Details regarding users' profiles, modes of use and access were reported elsewhere. The coverage of SNOMED CT was perceived to be at least 85% complete by 42% of responders, and 60% were at least satisfied with its quality. Various deficiencies were encountered at least 'somewhat often' by 28-61% of responders. Incorrect data were more bothersome than missing data. Users indicated that significant resources should be allocated to more consistent and complete conceptual representations and to further enhance content coverage. Enhanced synonym coverage and the introduction of textual definitions were important to users (54% and 63%, respectively). LIMITATIONS: A survey format with limited control over recruitment and selection bias. Lack of information regarding the SNOMED CT version used by responders. CONCLUSION: Despite overall satisfaction, direct users indicated a strong desire to improve consistency, quality, and completeness of conceptual representations and concept details, as well as a continued desire to expand coverage. The survey provides much needed data for informed decisions regarding the use and development goals of SNOMED CT. Focused periodical surveys are warranted. Gai Elhanan, Yehoshua Perl, James Geller |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Predicting Web Search Hit CountsabstractKeyword-based search engines often return an unexpected number of results. Zero hits are naturally undesirable, while too many hits are likely to be overwhelming and of low precision. We present an approach for predicting the number of hits for a given set of query terms. Using word frequencies derived from a large corpus, we construct random samples of combinations of these words as search terms. Then we derive a correlation function between the computed probabilities of search terms and the observed hit counts for them. This regression function is used to predict the hit counts for a user's new searches, with the intention of avoiding information overload. We report the results of experiments with Google, Yahoo! and Bing to validate our methodology. We further investigate the monotonicity of search results for negative search terms by those three search engines. Tian Tian 0002, James Geller, Soon Ae Chun |
Web Intelligence | 2 |
| 2009 | Comparing Inconsistent Relationship Configurations Indicating UMLS Errors
James Geller, C. Paul Morrey, Junchuan Xu, Michael Halper, Gai Elhanan, Yehoshua Perl, George Hripcsak |
AMIA | 1 |
| 2009 | Auditing SNOMED Relationships Using a Converse Abstraction Network
Duo Helen Wei, Michael Halper, Gai Elhanan, Yan Chen 0009, Yehoshua Perl, James Geller, Kent A. Spackman |
AMIA | 6 |
| 2009 | Using WordNet synonym substitution to enhance UMLS source integration
Kuo-Chuan Huang, James Geller, Michael Halper, Yehoshua Perl, Junchuan Xu |
Artif. Intell. Medicine | 2 |
| 2009 | Structural group-based auditing of missing hierarchical relationships in UMLS
Yan Chen 0009, Huanying Gu, Yehoshua Perl, James Geller |
J. Biomed. Informatics | 4 |
| 2009 | Structural group auditing of a UMLS semantic type's extent
Yan Chen 0009, Huanying Gu, Yehoshua Perl, James Geller, Michael Halper |
J. Biomed. Informatics | 4 |
| 2009 | Special Issue on Auditing of Terminologies
James Geller, Yehoshua Perl, Michael Halper, Ronald Cornet |
J. Biomed. Informatics | 1 |
| 2009 | The Neighborhood Auditing Tool: A hybrid interface for auditing the UMLS
C. Paul Morrey, James Geller, Michael Halper, Yehoshua Perl |
J. Biomed. Informatics | 2 |
| 2008 | Enriching Ontology for Deep Web Search
Yoo Jung An, Soon Ae Chun, Kuo-Chuan Huang, James Geller |
DEXA | 4 |
| 2008 | Comparing and consolidating two heuristic metaschemas
Yan Chen 0009, Yehoshua Perl, James Geller, George Hripcsak, Li Zhang 0048 |
J. Biomed. Informatics | 3 |
| 2007 | Evaluation of a UMLS Auditing Process of Semantic Type Assignments
Huanying Gu, George Hripcsak, Yan Chen 0009, C. Paul Morrey, Gai Elhanan, James J. Cimino, James Geller, Yehoshua Perl |
AMIA | 7 |
| 2007 | Piecewise Synonyms for Enhanced UMLS Source Terminology Integration
Kuo-Chuan Huang, James Geller, Michael Halper, James J. Cimino |
AMIA | 2 |
| 2007 | Ownership as a conceptual modeling construct
Michael Halper, Li-min Liu, James Geller, Yehoshua Perl |
Data Knowl. Eng. | 3 |
| 2007 | Analysis of a Study of the Users, Uses, and Future Agenda of the UMLSabstractOBJECTIVES: The UMLS constitutes the largest existing collection of medical terms. However, little has been published about the users and uses of the UMLS. This study sheds light on these issues. DESIGN: We designed a questionnaire consisting of 26 questions and distributed it to the UMLS user mailing list. Participants were assured complete confidentiality of their replies. To further encourage list members to respond, we promised to provide them with early results prior to publication. Sector analysis of the responses, according to employment organizations is used to obtain insights into some responses. RESULTS: We received 70 responses. The study confirms two intended uses of the UMLS: access to source terminologies (75%), and mapping among them (44%). However, most access is just to a few sources, led by SNOMED, MeSH, and ICD. Out of 119 reported purposes of use, terminology research (37), information retrieval (19), and terminology translation (14) lead. Four important observations are that the UMLS is widely used as a terminology (77%), even though it was not designed as one; many users (73%) want the NLM to mark concepts with multiple parents in an indented hierarchy and to derive a terminology from the UMLS (73%). Finally, auditing the UMLS is a top budget priority (35%) for users. CONCLUSIONS: The study reports many uses of the UMLS in a variety of subjects from terminology research to decision support and phenotyping. The study confirms that the UMLS is used to access its source terminologies and to map among them. Two primary concerns of the existing user base are auditing the UMLS and the design of a UMLS-based derived terminology. Yan Chen 0009, Yehoshua Perl, James Geller, James J. Cimino |
J. Am. Medical Informatics Assoc. | 3 |
| 2006 | Research Paper: Auditing as Part of the Terminology Design Life CycleabstractOBJECTIVE: To develop and test an auditing methodology for detecting errors in medical terminologies satisfying systematic inheritance. This methodology is based on various abstraction taxonomies that provide high-level views of a terminology and highlight potentially erroneous concepts. DESIGN: Our auditing methodology is based on dividing concepts of a terminology into smaller, more manageable units. First, we divide the terminology's concepts into areas according to their relationships/roles. Then each multi-rooted area is further divided into partial-areas (p-areas) that are singly-rooted. Each p-area contains a set of structurally and semantically uniform concepts. Two kinds of abstraction networks, called the area taxonomy and p-area taxonomy, are derived. These taxonomies form the basis for the auditing approach. Taxonomies tend to highlight potentially erroneous concepts in areas and p-areas. Human reviewers can focus their auditing efforts on the limited number of problematic concepts following two hypotheses on the probable concentration of errors. RESULTS: A sample of the area taxonomy and p-area taxonomy for the Biological Process (BP) hierarchy of the National Cancer Institute Thesaurus (NCIT) was derived from the application of our methodology to its concepts. These views led to the detection of a number of different kinds of errors that are reported, and to confirmation of the hypotheses on error concentration in this hierarchy. CONCLUSION: Our auditing methodology based on area and p-area taxonomies is an efficient tool for detecting errors in terminologies satisfying systematic inheritance of roles, and thus facilitates their maintenance. This methodology concentrates a domain expert's manual review on portions of the concepts with a high likelihood of errors. Hua Min, Yehoshua Perl, Yan Chen 0009, Michael Halper, James Geller, Yue Wang 0033 |
J. Am. Medical Informatics Assoc. | 5 |
| 2006 | Semantic enrichment for medical ontologies
Yugyung Lee, James Geller |
J. Biomed. Informatics | 2 |
| 2005 | An expert study evaluating the UMLS lexical metaschema
Li Zhang 0048, George Hripcsak, Yehoshua Perl, Michael Halper, James Geller |
Artif. Intell. Medicine | 5 |
| 2005 | A lexical metaschema for the UMLS semantic network
Li Zhang 0048, Yehoshua Perl, Michael Halper, James Geller, George Hripcsak |
Artif. Intell. Medicine | 4 |
| 2005 | Raising data for improved support in rule mining: How to raise and how far to raise
James Geller, Xuan Zhou 0007, Kalpana Prathipati, Sripriya Kanigiluppai |
Intell. Data Anal. | 1 |
| 2005 | Model Formulation: Relationship Structures and Semantic Type Assignments of the UMLS Enriched Semantic NetworkabstractOBJECTIVE: The Enriched Semantic Network (ESN) was introduced as an extension of the Unified Medical Language System (UMLS) Semantic Network (SN). Its multiple subsumption configuration and concomitant multiple inheritance make the ESN's relationship structures and semantic type assignments different from those of the SN. A technique for deriving the relationship structures of the ESN's semantic types and an automated technique for deriving the ESN's semantic type assignments from those of the SN are presented. DESIGN: The technique to derive the ESN's relationship structures finds all newly inherited relationships in the ESN. All such relationships are audited for semantic validity, and the blocking mechanism is used to block invalid relationships. The mapping technique to derive the ESN's semantic type assignments uses current SN semantic type assignments and preserves nonredundant categorizations, while preventing new redundant categorizations. RESULTS: Among the 426 newly inherited relationships, 326 are deemed valid. Seven blockings are applied to avoid inheritance of the 100 invalid relationships. Sixteen semantic types have different relationship structures in the ESN as compared to those in the SN. The mapping of semantic type assignments from the SN to the ESN avoids the generation of 26,950 redundant categorizations. The resulting ESN contains 138 semantic types, 149 IS-A links, 7,303 relationships, and 1,013,876 semantic type assignments. CONCLUSION: The ESN's multiple inheritance provides more complete relationship structures than in the SN. The ESN's semantic type assignments avoid the existing redundant categorizations appearing in the SN and prevent new ones that might arise due to multiple parents. Compared to the SN, the ESN provides a more accurate unifying semantic abstraction of the UMLS Metathesaurus. Li Zhang 0048, Michael Halper, Yehoshua Perl, James Geller, James J. Cimino |
J. Am. Medical Informatics Assoc. | 4 |
| 2004 | Ontological and Pragmatic Knowledge Management for Web Service Composition
Soon Ae Chun, Yugyung Lee, James Geller |
DASFAA | 3 |
| 2004 | Towards Intelligent Web Services for Automating Medical Service CompositionabstractThe vision of the Semantic Web is to reduce manual discovery and usage of Web resources (documents and services) and to allow software agents to automatically identify these Web resources, integrate them and execute them for achieving the intended goals of the user. Such a composed Web service may be represented as a workflow, called service flow. Current studies of Web services are not sufficient for automatic composition. This paper presents different types of compositional knowledge required for Web service discovery and composition: syntactic, semantic and pragmatic knowledge. As a proof of concept, we have implemented our framework in a cardiovascular domain which requires advanced service discovery and composition across heterogeneous platforms of multiple organizations. Within this framework, we describe (1) How to represent the compositional knowledge, which plays a role in service discovery and composition, in DAML-S; (2) How heterogeneous medical services intemperate in a composed medical service flow. (3) To solve this knowledge level integration, we build on an ontology integration method called SEMIO (Semantic Interoperability). Yugyung Lee, Chintan Patel, Soon Ae Chun, James Geller |
ICWS | 4 |
| 2004 | Model Formulation: An Enriched Unified Medical Language System Semantic Network with a Multiple Subsumption HierarchyabstractOBJECTIVE: The Unified Medical Language System's (UMLS's) Semantic Network's (SN's) two-tree structure is restrictive because it does not allow a semantic type to be a specialization of several other semantic types. In this article, the SN is expanded into a multiple subsumption structure with a directed acyclic graph (DAG) IS-A hierarchy, allowing a semantic type to have multiple parents. New viable IS-A links are added as warranted. DESIGN: Two methodologies are presented to identify and add new viable IS-A links. The first methodology is based on imposing the characteristic of connectivity on a previously presented partition of the SN. Four transformations are provided to find viable IS-A links in the process of converting the partition's disconnected groups into connected ones. The second methodology identifies new IS-A links through a string matching process involving names and definitions of various semantic types in the SN. A domain expert is needed to review all the results to determine the validity of the new IS-A links. RESULTS: Nineteen new IS-A links are added to the SN, and four new semantic types are also created to support the multiple subsumption framework. The resulting network, called the Enriched Semantic Network (ESN), exhibits a DAG-structured hierarchy. A partition of the ESN containing 19 connected groups is also derived. CONCLUSION: The ESN is an expanded abstraction of the UMLS compared with the original SN. Its multiple subsumption hierarchy can accommodate semantic types with multiple parents. Its representation thus provides direct access to a broader range of subsumption knowledge. Li Zhang 0048, Yehoshua Perl, Michael Halper, James Geller, James J. Cimino |
J. Am. Medical Informatics Assoc. | 4 |
| 2004 | Editorial: Ontology Challenges: A Thumbnail Historical Perspective
James Geller, Yehoshua Perl, Jintae Lee |
Knowl. Inf. Syst. | 1 |
| 2004 | Contextual Partitioning for Comprehension of OODB Schemas
Huanying Gu, Yehoshua Perl, Michael Halper, James Geller, Erich J. Neuhold |
Knowl. Inf. Syst. | 4 |
| 2003 | Using an Interest Ontology for Improved Support in Rule Mining
Xuan Zhou 0007, Richard B. Scherl, James Geller |
DaWaK | 4 |
| 2003 | Frameworks for incorporating semantic relationships into object-oriented database systemsabstractAbstract A semantic relationship is a data modeling construct that connects a pair of classes or categories and has inherent constraints and other functionalities that precisely reflect the characteristics of the specific relationship in an application domain. Examples of semantic relationships include part–whole, ownership, materialization and role‐of. Such relationships are important in the construction of information models for advanced applications, whether one is employing traditional data‐modeling techniques, knowledge‐representation languages or object‐oriented modeling methodologies. This paper focuses on the issue of providing built‐in support for such constructs in the context of object‐oriented database (OODB) systems. Most of the popular object‐oriented modeling approaches include some semantic relationships in their repertoire of data‐modeling primitives. However, commercial OODB systems, which are frequently used as implementation vehicles, tend not to do the same. We will present two frameworks by which a semantic relationship can be incorporated into an existing OODB system. The first only requires that the OODB system support manifest type with respect to its instances. The second assumes that the OODB system has a special kind of metaclass facility. The two frameworks are compared and contrasted. In order to ground our work in existing systems, we show the addition of a part–whole semantic relationship both to the ONTOS DB/Explorer OODB system and the VODAK Model Language. Copyright © 2003 John Wiley & Sons, Ltd. Michael Halper, Li-min Liu, James Geller, Yehoshua Perl |
Concurr. Comput. Pract. Exp. | 3 |
| 2003 | Enhancing OODB semantics to support browsing in an OODB vocabulary representationabstractAbstract In previous work, we have modeled a vocabulary given as a semantic network by an object‐oriented database (OODB). The OODB schema thus obtained provides a compact abstract view of the vocabulary. This enables the fast traversal of the vocabulary by a user. In the semantic network vocabulary, the IS‐A relationships express the specialization hierarchy. In our OODB modeling of the vocabulary, the SUBCLASS relationship expresses the specialization hierarchy of the classes and supports the inheritance of their properties. A typical IS‐A path in the vocabulary has a corresponding shorter SUBCLASS path in the OODB schema. In this paper we expose several cases where the SUBCLASS hierarchy fails to fully correspond to the IS‐A hierarchy of the vocabulary. In these cases there exist traversal paths in the semantic network for which there are no corresponding traversal paths in the OODB schema. The reason for this failure is the existence of some IS‐A relationships between concepts of two classes, which are not connected by a SUBCLASS relationship. This phenomenon weakens the accuracy of our modeling. To rectify the situation we introduce a new OODB semantic relationship IS‐A$'$ to represent the existence of IS‐A relationships between concepts of a pair of classes which are not connected via a SUBCLASS relationship. The resulting schema contains both SUBCLASS relationships and IS‐A$'$ relationships which completely model the IS‐A hierarchy of the vocabulary. We define a mixed‐class level traversal path to contain either SUBCLASS or IS‐A$'$ relationships. Consequently, each traversal path in the semantic network has a corresponding mixed traversal path in the OODB schema. Hence the introduction of the semantic OODB IS‐A$'$ relationship improves the modeling of semantic network vocabularies by OODBs. Copyright © 2003 John Wiley & Sons, Ltd. Li-min Liu, James Geller, Yehoshua Perl |
Concurr. Comput. Pract. Exp. | 2 |
| 2003 | Semantic refinement and error correction in large terminological knowledge bases
James Geller, Huanying Gu, Yehoshua Perl, Michael Halper |
Data Knowl. Eng. | 1 |
| 2003 | Research on structural issues of the UMLS - past, present, and future
Yehoshua Perl, James Geller |
J. Biomed. Informatics | 2 |
| 2003 | Designing metaschemas for the UMLS enriched semantic network
Li Zhang 0048, Yehoshua Perl, Michael Halper, James Geller |
J. Biomed. Informatics | 4 |
| 2002 | Auditing the UMLS for redundant classifications
Michael Halper, Yehoshua Perl, James Geller |
AMIA | 4 |
| 2002 | Enriching the structure of the UMLS semantic network
Li Zhang 0048, Yehoshua Perl, Michael Halper, James Geller, James J. Cimino |
AMIA | 4 |
| 2002 | The cohesive metaschema: a higher-level abstraction of the UMLS Semantic Network
Yehoshua Perl, Zong Chen, Michael Halper, James Geller, Li Zhang 0048 |
J. Biomed. Informatics | 4 |
| 2002 | Efficient Transitive Closure Reasoning in a Combined Class-Part-Containment Hierarchy
Yugyung Lee, James Geller |
Knowl. Inf. Syst. | 2 |
| 2002 | Partitioning the UMLS semantic networkabstractThe unified medical language system (UMLS) integrates many well-established biomedical terminologies. The UMLS semantic network (SN) can help orient users to the vast knowledge content of the UMLS Metathesaurus (META) via its abstract conceptual view. However, the SN itself is large and complex and may still be difficult to comprehend. Our technique partitions the SN into smaller meaningful units amenable to display on limited-sized computer screens. The basis for the partitioning is the distribution of the relationships within the SN. Three rules are applied to transform the original partition into a second more cohesive partition. Zong Chen, Yehoshua Perl, Michael Halper, James Geller, Huanying Gu |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2002 | Evaluation and application of a semantic network partitionabstractSemantic networks (SNs) are excellent knowledge representation structures. However, large semantic networks are hard to comprehend. To overcome this difficulty, several methods of partitioning have been developed that rely on different mixes of structural and semantic methods. However, little has appeared in the literature concerning the question whether a partition of a semantic network creates subnetworks that agree with human insight. We address this issue by presenting a comparison between the results of an algorithmic partitioning method and a partition created by a group of experts. Subsequently, we show how a network partition can be used to generate various partial views of a semantic network, which facilitate user orientation. Examples from the Unified Medical Language System (UMLS) SN are used to demonstrate partial views. James Geller, Yehoshua Perl, Michael Halper, Zong Chen, Huanying Gu |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2002 | Using OODB Modeling to Partition a Vocabulary in Structurally and Semantically Uniform Concept GroupsabstractControlled vocabularies (CVs) are networks of concepts that unify disparate terminologies and facilitate the process of information sharing within an application domain. We describe a general methodology for representing an existing CV as an object-oriented database (OODB), called an object-oriented vocabulary repository (OOVR). A formal description of the OOVR methodology, which is based on a structural abstraction technique, is given, along with an algorithmic description and a number of theorems pertaining to some of the methodology's formal characteristics. An OOVR offers a two-level (concept level and schema level) view of a CV, with the schema-level view serving as an important abstraction that can aid in orientation to the CV's contents. While an OOVR can also assist in traversals of the CV, we have identified certain special CV configurations where such traversals can be problematic. To address this, we introduce - based on the original methodology - an enhanced OOVR methodology that utilizes both structural and semantic features to partition and model a CV's constituent concepts. With its basis in the notions of area and the recursively defined articulation concept, an enhanced OOVR representation provides users with an improved CV view comprising groups of concepts that are uniform both in their structure and semantics. An algorithmic description of the singly-rooted OOVR methodology and theorems describing some of its formal properties are given. The results of applying it to a large existing CV are discussed. Li-min Liu, Michael Halper, James Geller, Yehoshua Perl |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2001 | A metaschema of the UMLS based on a partition of its semantic network
Michael Halper, Zong Chen, James Geller, Yehoshua Perl |
AMIA | 3 |
| 2000 | Partitioning the Semantic Network of the UMLS
Zong Chen, Michael Halper, James Geller, Yehoshua Perl |
AMIA | 3 |
| 2000 | How to Partition a Complex Schema of a Medical Terminology
Huanying Gu, Yehoshua Perl, Michael Halper, James Geller, Feng-shen Kuo, James J. Cimino |
AMIA | 4 |
| 2000 | Research Paper: Representing the UMLS as an Object-oriented Database: Modeling Issues and AdvantagesabstractOBJECTIVE: The Unified Medical Language System (UMLS) combines many well-established authoritative medical informatics terminologies in one knowledge representation system. Such a resource is very valuable to the health care community and industry. However, the UMLS is very large and complex and poses serious comprehension problems for users and maintenance personnel. The authors present a representation to support the user's comprehension and navigation of the UMLS. DESIGN: An object-oriented database (OODB) representation is used to represent the two major components of the UMLS-the Metathesaurus and the Semantic Network-as a unified system. The semantic types of the Semantic Network are modeled as semantic type classes. Intersection classes are defined to model concepts of multiple semantic types, which are removed from the semantic type classes. RESULTS: The authors provide examples of how the intersection classes help expose omissions of concepts, highlight errors of semantic type classification, and uncover ambiguities of concepts in the UMLS. The resulting UMLS OODB schema is deeper and more refined than the Semantic Network, since intersection classes are introduced. The Metathesaurus is classified into more mutually exclusive, uniform sets of concepts. The schema improves the user's comprehension and navigation of the Metathesaurus. CONCLUSIONS: The UMLS OODB schema supports the user's comprehension and navigation of the Metathesaurus. It also helps expose and resolve modeling problems in the UMLS. Huanying Gu, Yehoshua Perl, James Geller, Michael Halper, Li-min Liu, James J. Cimino |
J. Am. Medical Informatics Assoc. | 3 |
| 1999 | Modeling the UMLS using an OODB
Huanying Gu, Yehoshua Perl, James Geller, Michael Halper, Li-min Liu, James J. Cimino |
AMIA | 3 |
| 1999 | Using a Similarity Measurement to Partition a Vocabulary of Medical Concepts
Huanying Gu, James Geller, Li-min Liu, Michael Halper |
DEXA | 2 |
| 1999 | A methodology for partitioning a vocabulary hierarchy into trees
Huanying Gu, Yehoshua Perl, James Geller, Michael Halper, Mansnimar Singh |
Artif. Intell. Medicine | 3 |
| 1999 | Controlled Vocabularies in OODBs: Modeling Issues and Implementation
Li-min Liu, Michael Halper, James Geller, Yehoshua Perl |
Distributed Parallel Databases | 3 |
| 1999 | Guest Editors' Introduction
James Geller, Frederick H. Lochovsky |
Int. J. Cooperative Inf. Syst. | 1 |
| 1999 | Model Formulation: Benefits of an Object-oriented Database Representation for Controlled Medical TerminologiesabstractOBJECTIVE: Controlled medical terminologies (CMTs) have been recognized as important tools in a variety of medical informatics applications, ranging from patient-record systems to decision-support systems. Controlled medical terminologies are typically organized in semantic network structures consisting of tens to hundreds of thousands of concepts. This overwhelming size and complexity can be a serious barrier to their maintenance and widespread utilization. The authors propose the use of object-oriented databases to address the problems posed by the extensive scope and high complexity of most CMTs for maintenance personnel and general users alike. DESIGN: The authors present a methodology that allows an existing CMT, modeled as a semantic network, to be represented as an equivalent object-oriented database. Such a representation is called an object-oriented health care terminology repository (OOHTR). RESULTS: The major benefit of an OOHTR is its schema, which provides an important layer of structural abstraction. Using the high-level view of a CMT afforded by the schema, one can gain insight into the CMT's overarching organization and begin to better comprehend it. The authors' methodology is applied to the Medical Entities Dictionary (MED), a large CMT developed at Columbia-Presbyterian Medical Center. Examples of how the OOHTR schema facilitated updating, correcting, and improving the design of the MED are presented. CONCLUSION: The OOHTR schema can serve as an important abstraction mechanism for enhancing comprehension of a large CMT, and thus promotes its usability. Huanying Gu, Michael Halper, James Geller, Yehoshua Perl |
J. Am. Medical Informatics Assoc. | 3 |
| 1998 | Converting an integrated hospital formulary into an object-oriented database representation
Huanying Gu, Li-min Liu, Michael Halper, James Geller, Yehoshua Perl |
AMIA | 4 |
| 1998 | An OODB Part-Whole Model: Semantics, Notation and Implementation
Michael Halper, James Geller, Yehoshua Perl |
Data Knowl. Eng. | 2 |
| 1998 | The OODB Path-Method Generator (PMG) Using Access Weights and Precomputed Access Relevance
Ashish Mehta, James Geller, Yehoshua Perl, Erich J. Neuhold |
VLDB J. | 2 |
| 1997 | Partitioning a vocabulary's IS-A hierarchy into trees
Huanying Gu, Yehoshua Perl, James Geller, Michael Halper, James J. Cimino, Mansnimar Singh |
AMIA | 3 |
| 1997 | Challenge: How IJCAI 1999 can Prove Value of AI by Using AI
James Geller |
IJCAI (1) | 1 |
| 1996 | Modeling a Vocabulary in an Object-Oriented DatabaseabstractControlled vocabularies have been used as the means for unifying disparate terminologies found within an application field. This unification leads to better administration of information and enhanced communication among various parties. Semantic networks have been shown to be excellent vehicles for modeling controlled vocabularies. However, they often lack the necessary access flexibility and robustness required by external agents such as intelligent information-locators and decision-support systems. In this paper, we describe the process of mapping an existing medical vocabulary based on a semantic network model into an Object-Oriented Database (OODB) system. We first consider two straightforward approaches to carrying out this task and describe their deficiencies. We then present a new approach which yields a very compact OODB schema for the representation of the vocabulary's entire hierarchy and inter-connectivity. We refer to the resulting OODB as the Object-Oriented Healthcare Voc... Li-min Liu, Michael Halper, Huanying Gu, James Geller, Yehoshua Perl |
CIKM | 4 |
| 1996 | Identifying a Forest Hierarchy in an OODB Specification Hierarchy Satisfying Disciplined ModelingabstractThe work is motivated by the desire to develop methods to comprehend large vocabularies and large schemas of object-oriented databases. The ability of a user of a database participating in a federated system to retrieve information from the other database systems will be greatly enhanced by acquiring a better comprehension of these systems. The authors are trying to develop both a theoretical paradigm and a methodology to analyze existing large schemas. Their approach to achieve comprehension is based on combining two concepts: informational thinning (i.e. concentration on the specialization hierarchy of the schema) and partitioning. They present a new technique for modeling which is called disciplined modeling. Based on the rules of disciplined modeling we develop a theoretical paradigm to support the existence of a meaningful forest hierarchy within the specialization hierarchy. Such a hierarchy functions as a skeleton of the schema and supports comprehension and partitioning efforts. Yehoshua Perl, James Geller, Huanying Gu |
CoopIS | 2 |
| 1996 | Parallel Transitive Reasoning in Mixed Relational Hierarchies
Yugyung Lee, James Geller |
KR | 2 |
| 1996 | Computing Access Relevance for Path-Method Generation in OODBs and IM-OODB
Ashid Metha, James Geller, Yehoshua Perl, Peter Fankhauser |
J. Intell. Inf. Syst. | 2 |
| 1994 | Integrating a Part Relationship Into an Open OODB System Using MetaclassesabstractThe part-whole semantic relationship (the part relationship, for short) is an important modeling primitive in many advanced application domains such as manufacturing, design, and document processing. In this paper, we examine the problem of integrating such a construct into an OODB system. Specifically, two questions are addressed in this regard. This first is: Can a part relationship be made an intrinsic construct of an existing OODB system without having to rewrite a substantial portion of the system? The second: Can an “open” OODB system which claims to support such an integration really do so, and, more specifically, can the integration be done using a metaclass mechanism which purports to bring extensibility to the VODAK Model Language (VML)? Michael Halper, James Geller, Yehoshua Perl, Wolfgang Klas |
CIKM | 2 |
| 1993 | Value Propagation in Object-Oriented Database Part HierarchiesabstractDerived schema components are an important aspect of traditional semantic data modeling.In this paper, we address the issue of defining Michael Halper, James Geller, Yehoshua Perl |
CIKM | 2 |
| 1993 | The OODB Path-Method Generator (PMG) Using Precomputed Access RelevanceabstractA path-method is used as a mechanism in object-Our experiments show that the traversal algorithm of PMG is a very successful tool for aiding the user with the dificult task of querying and updating a large OODB. Ashish Mehta, James Geller, Yehoshua Perl, Erich J. Neuhold |
CIKM | 2 |
| 1993 | Design and Implementation of a Knowledge-Based Query ProcessorabstractThis paper deals with query processing using semantic knowledge in relational databases. The Select-Project-Join (SPJ) conjunctive class of queries are dealt with in this paper. We propose to optimize highly repetitive queries by using semantic transformations in addition to syntactic transformations. Thus, we generate a set of pre-optimized queries. This set contains queries that are semantically equivalent to, syntactically different from, and more efficient to process than the user queries that we started with. The issues we address in this paper are: how to map a user query to a query that is in the set of pre-optimized and already optimized queries, how to search efficiently through the set of pre-optimized queries and set of semantic rules, and how to incorporate new queries to the set of pre-optimized queries, so that the number of queries that can be optimized using this method increases with the passage of time. Furthermore, we suggest some ideas of handling queries that do not have any semantically equivalent counterpart in the set of pre-optimized queries. We have tested the performance of the proposed method. An algorithm for mapping is implemented in Prolog. A database schema is implemented in the INGRES database management system. We have adopted a database schema that is widely used for measuring performance in the semantic query optimization literature. Nabil R. Adam, Aryya Gangopadhyay, James Geller |
Int. J. Cooperative Inf. Syst. | 3 |
| 1992 | "Part" Relations for Object-Oriented Databases
Michael Halper, James Geller, Yehoshua Perl |
ER | 2 |
| 1992 | Structural schema integration with full and partial correspondence using the Dual Model
James Geller, Yehoshua Perl, Erich J. Neuhold, Amit P. Sheth |
Inf. Syst. | 1 |
| 1991 | Propositional Representation for Graphical Knowledge
James Geller |
Int. J. Man Mach. Stud. | 1 |
| 1991 | Parallel implementation of a class reasonerabstractOne method to overcome the notorious efficiency problems of logical reasoning algorithms in AI has been to combine a general-purpose reasoner with several special-purpose reasoners for commonly used subtasks. In this paper we are using Schubert's (Schubert et al. 1983, 1987) method of implementing a special-purpose class reasoner. We show that it is possible to replace Schubert's preorder number class tree by a preorder number list without loss of functionality. This form of the algorithm lends itself perfectly towards a parallel implementation,1 and we describe design, coding and testing of such an implementation. Our algorithm is practically independent of the size of the class list, and even with several thousand nodes learning times are under a second and retrieval times are under 500 ms. James Geller, Charles (yaogui) Du |
J. Exp. Theor. Artif. Intell. | 1 |
| 1987 | Graphical Deep Knowledge for Intelligent Machine Drafting
James Geller, Stuart C. Shapiro |
IJCAI | 1 |
| 1985 | The Teachable Letter Recognizer
James Geller |
IJCAI | 1 |