Julia Xu

dblp:122/4285 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-4147-8388ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Ontology enrichment using a large language model: Applying lexical, semantic, and knowledge network-based similarity for concept placement
abstract
OBJECTIVE: Ontologies are essential for representing the knowledge of a domain. To make ontologies useful, they must encompass a comprehensive domain view. To achieve ontology enrichment, there is a need to discover new concepts to be added, either because they were missed in the first place, or the state-of-the-art has advanced to develop new real-world concepts. Our goal is to develop an automatic enrichment pipeline using a seed ontology, a Large Language Model (LLM), and source of text. The pipeline is applied to the domain of Social Determinants of Health (SDoH), using PubMed as a source of concepts. In this work, the applicability and effectiveness of the enrichment pipeline is demonstrated by extending the SDoH Ontology called SOHOv1, however our methodology could be used in other domains as well. METHODS: We first retrieved PubMed abstracts of candidate articles with existing SOHOv1 concepts as search terms. Next, we used GPT-4-1201 to extract semantic triples from the abstracts. We identified concepts from these triples utilizing lexical, semantic, and knowledge network-based filtering. We also compared the granularity of semantic triples extracted with our method to the triples in the SemMedDB (Semantic MEDLINE Database). The results were evaluated by human experts and standard ontology tools for checking consistency and semantic correctness. RESULTS: We expanded SOHOv1, which contained 173 concepts and 585 axioms, including 207 logical axioms to SOHOv2, which contains 572 concepts, 1,542 axioms, including 725 logical axioms. Our methods identified more concepts than those extracted from SemMedDB for the same task. While we have shown the feasibility of our approach for an SDoH ontology, the methodology is generalizable to other ontologies with an existing seed ontology and text corpus. CONCLUSIONS: The contributions of this work are: Extracting semantic triples from PubMed abstracts using GPT-4-1201 utilizing prompt chaining; showing the superiority of triples from GPT-4-1201 over triples from SemMedDB for SDoH; using lexical and semantic similarity search techniques with knowledge network-based search to identify the concepts to be added to the ontology; confirming the quality of the new concepts with human experts.
Navya Martin Kollapally, James Geller, Vipina Kuttichi Keloth, Zhe He 0001, Julia Xu
J. Biomed. Informatics5
2024 Promoting interoperability between SNOMED CT and ICD-11: lessons learned from the pilot project mapping between SNOMED CT and the ICD-11 Foundation
abstract
OBJECTIVE: To explore the feasibility and challenges of mapping between SNOMED CT and the ICD-11 Foundation in both directions, SNOMED International and the World Health Organization conducted a pilot mapping project between September 2021 and August 2022. MATERIALS AND METHODS: Phase 1 mapped ICD-11 Foundation entities from the endocrine diseases chapter, excluding malignant neoplasms, to SNOMED CT. In phase 2, SNOMED CT concepts equivalent to those covered by the ICD-11 entities in phase 1 were mapped to the ICD-11 Foundation. The goal was to identify equivalence between an ICD-11 Foundation entity and a SNOMED CT concept. Postcoordination was used for mapping to ICD-11. Each map was done twice independently, the results were compared, and discrepancies were reconciled. RESULTS: In phase 1, 59% of 637 ICD-11 Foundation entities had an exact match in SNOMED CT. In phase 2, 32% of 1893 SNOMED CT concepts had an exact match in the ICD-11 Foundation, and postcoordination added 15% of exact match. Challenges encountered included non-synonymous synonyms, mismatch in granularity, composite conditions, and residual categories. CONCLUSION: This pilot project shed light on the tremendous amount of effort required to create a map between the 2 coding systems and uncovered some common challenges. Future collaborative work between SNOMED International and WHO will likely benefit from its findings. It is recommended that the 2 organizations should clarify goals and use cases of mapping, provide adequate resources, set up a road map, and reconsider their original proposal of incorporating SNOMED CT into the ICD-11 Foundation ontology.
Kin Wah Fung, Julia Xu, Hazel Brear, Alana Lane, Maggie Lau, Austen Wong, Arabella D'Havé
J. Am. Medical Informatics Assoc.2
2023 REALIMPACT: A Dataset of Impact Sound Fields for Real Objects
abstract
Objects make unique sounds under different perturbations, environment conditions, and poses relative to the listener. While prior works have modeled impact sounds and sound propagation in simulation, we lack a standard dataset of impact sound fields of real objects for audio-visual learning and calibration of the sim-to-real gap. We present Realimpact,a large-scale dataset of real object impact sounds recorded under controlled conditions. Re-alimpactcontains 150,000 recordings of impact sounds of 50 everyday objects with detailed annotations, including their impact locations, microphone locations, contact force profiles, material labels, and RGBD images.**The project page and dataset are available at https://samuelpclarke.com/realimpact/ We make preliminary attempts to use our dataset as a reference to current simulation methods for estimating object impact sounds that match the real world. Moreover, we demon-strate the usefulness of our dataset as a testbed for acoustic and audio-visual learning via the evaluation of two bench-mark tasks, including listener location classification and vi-sual acoustic matching.
Samuel Clarke, Ruohan Gao, Mason L. Wang, Mark Rau, Julia Xu, Jui-Hsien Wang, Doug L. James, Jiajun Wu 0001
CVPR5
2023 Mapping 3 procedure coding systems to the International Classification of Health Interventions (ICHI): coverage and challenges
abstract
OBJECTIVE: To study the coverage and challenges in mapping 3 national and international procedure coding systems to the International Classification of Health Interventions (ICHI). MATERIALS AND METHODS: We identified 300 commonly used codes each from SNOMED CT, ICD-10-PCS, and CCI (Canadian Classification of Health Interventions) and mapped them to ICHI. We evaluated the level of match at the ICHI stem code and Foundation Component levels. We used postcoordination (modification of existing codes by adding other codes) to improve matching. Failure analysis was done for cases where full representation was not achieved. We noted and categorized potential problems that we encountered in ICHI, which could affect the accuracy and consistency of mapping. RESULTS: Overall, among the 900 codes from the 3 sources, 286 (31.8%) had full match with ICHI stem codes, 222 (24.7%) had full match with Foundation entities, and 231 (25.7%) had full match with postcoordination. 143 codes (15.9%) could only be partially represented even with postcoordination. A small number of SNOMED CT and ICD-10-PCS codes (18 codes, 2% of total), could not be mapped because the source codes were underspecified. We noted 4 categories of problems in ICHI-redundancy, missing elements, modeling issues, and naming issues. CONCLUSION: Using the full range of mapping options, at least three-quarters of the commonly used codes in each source system achieved a full match. For the purpose of international statistical reporting, full matching may not be an essential requirement. However, problems in ICHI that could result in suboptimal maps should be addressed.
Kin Wah Fung, Julia Xu, Filip Ameye, Lisa Burelle, Janice Macneil
J. Am. Medical Informatics Assoc.2
2023 A practical strategy to use the ICD-11 for morbidity coding in the United States without a clinical modification
abstract
OBJECTIVE: The aim of this study was to derive and evaluate a practical strategy of replacing ICD-10-CM codes by ICD-11 for morbidity coding in the United States, without the creation of a Clinical Modification. MATERIALS AND METHODS: A stepwise strategy is described, using first the ICD-11 stem codes from the Mortality and Morbidity Statistics (MMS) linearization, followed by exposing Foundation entities, then adding postcoordination (with existing codes and adding new stem codes if necessary), with creating new stem codes as the last resort. The strategy was evaluated by recoding 2 samples of ICD-10-CM codes comprised of frequently used codes and all codes from the digestive diseases chapter. RESULTS: Among the 1725 ICD-10-CM codes examined, the cumulative coverage at the stem code, Foundation, and postcoordination levels are 35.2%, 46.5% and 89.4% respectively. 7.1% of codes require new extension codes and 3.5% require new stem codes. Among the new extension codes, severity scale values and anatomy are the most common categories. 5.5% of codes are not one-to-one matches (1 ICD-10-CM code matched to 1 ICD-11 stem code or Foundation entity) which could be potentially challenging. CONCLUSION: Existing ICD-11 content can achieve full representation of almost 90% of ICD-10-CM codes, provided that postcoordination can be used and the coding guidelines and hierarchical structures of ICD-10-CM and ICD-11 can be harmonized. The various options examined in this study should be carefully considered before embarking on the traditional approach of a full-fledged ICD-11-CM.
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
J. Am. Medical Informatics Assoc.2
2022 An Ontology for the Social Determinants of Health Domain
abstract
Social Determinants of Health (SDOH) are societal factors, such as where a person was born, grew up, works, lives, etc., along with socio-economic and community factors that affect an individual’s health. SDOH are correlated with many clinical outcomes, hence it is desirable to record SDOH data in Electronic Health Records (EHRs). Besides storing images, text, etc., EHRs rely on coded terms available in standard ontologies and terminologies to record observations and analyses. There is a substantial amount of research on understanding the clinical impact of SDOH, ranging from screening tools to practice-based interventions. However, there is no comprehensive collection of terms for recording SDOH observations in EHRs. Our research goal is to develop an ontology that covers the terms describing SDOH. We present a prototype ontology called Social Determinant of Health Ontology (SOHO) that covers relevant concepts and IS-A relationships describing impacts and associations of social determinants. We describe the evaluation techniques that we applied to SOHO, including human experts’ review and algorithmic evaluation.
Navya Martin Kollapally, Yan Chen 0009, Julia Xu, James Geller
BIBM3
2021 Can ICD-11 Replace ICD-10-CM for Morbidity Coding in the U.S.?
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
AMIA2
2021 Evaluation of the International Classification of Health Interventions (ICHI) in the coding of common surgical procedures
abstract
OBJECTIVE: To evaluate the International Classification of Health Interventions (ICHI) in the clinical and statistical use cases. MATERIALS AND METHODS: We identified 300 most-performed surgical procedures as represented by their display names in an electronic health record. For comparison with existing coding systems, we coded the procedures in ICHI, SNOMED CT, International Classification of Diseases (ICD)-10-PCS, and CCI (Canadian Classification of Health Interventions), using postcoordination (modification of existing codes by adding other codes), when applicable. Failure analysis was done for cases where full representation was not achieved. The ICHI encoding was further evaluated for adequacy to support statistical reporting by the Organisation for Economic Co-operation and Development (OECD) and European Union (EU) categories of surgical procedures. RESULTS: After deduplication, 229 distinct procedures remained. Without postcoordination, ICHI achieved full representation in 52.8%. A further 19.2% could be fully represented with postcoordination. SNOMED CT was the best performing overall, with 94.3% full representation without postcoordination, and 99.6% with postcoordination. Failure analysis showed that "method" and "target" constituted most of the missing information for ICHI encoding. For all OECD/EU surgical categories, ICHI coding was adequate to support statistical reporting. One OECD/EU category ("Hip replacement, secondary") required postcoordination for correct assignment. CONCLUSION: In the clinical use case of capturing information in the electronic health record, ICHI was outperformed by the clinically oriented procedure coding systems (SNOMED CT and CCI), but was comparable to ICD-10-PCS. Postcoordination could be an effective and efficient means of improving coverage. ICHI is generally adequate for the collection of international statistics.
Kin Wah Fung, Julia Xu, Filip Ameye, Lisa Burelle, Janice Macneil
J. Am. Medical Informatics Assoc.2
2021 Feasibility of replacing the ICD-10-CM with the ICD-11 for morbidity coding: A content analysis
abstract
OBJECTIVE: The study sought to assess the feasibility of replacing the International Classification of Diseases-Tenth Revision-Clinical Modification (ICD-10-CM) with the International Classification of Diseases-11th Revision (ICD-11) for morbidity coding based on content analysis. MATERIALS AND METHODS: The most frequently used ICD-10-CM codes from each chapter covering 60% of patients were identified from Medicare claims and hospital data. Each ICD-10-CM code was recoded in the ICD-11, using postcoordination (combination of codes) if necessary. Recoding was performed by 2 terminologists independently. Failure analysis was done for cases where full representation was not achieved even with postcoordination. After recoding, the coding guidance (inclusions, exclusions, and index) of the ICD-10-CM and ICD-11 codes were reviewed for conflict. RESULTS: Overall, 23.5% of 943 codes could be fully represented by the ICD-11 without postcoordination. Postcoordination is the potential game changer. It supports the full representation of 8.6% of 943 codes. Moreover, with the addition of only 9 extension codes, postcoordination supports the full representation of 35.2% of 943 codes. Coding guidance review identified potential conflicts in 10% of codes, but mostly not affecting recoding. The majority of the conflicts resulted from differences in granularity and default coding assumptions between the ICD-11 and ICD-10-CM. CONCLUSIONS: With some minor enhancements to postcoordination, the ICD-11 can fully represent almost 60% of the most frequently used ICD-10-CM codes. Even without postcoordination, 23.5% full representation is comparable to the 24.3% of ICD-9-CM codes with exact match in the ICD-10-CM, so migrating from the ICD-10-CM to the ICD-11 is not necessarily more disruptive than from the International Classification of Diseases-Ninth Revision-Clinical Modification to the ICD-10-CM. Therefore, the ICD-11 (without a CM) should be considered as a candidate to replace the ICD-10-CM for morbidity coding.
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider
J. Am. Medical Informatics Assoc.2
2020 Comparison of Three International Terminologies for Medical Interventions and Procedures
Kin Wah Fung, Julia Xu, Filip Ameye
AMIA2
2020 The new International Classification of Diseases 11th edition: a comparative analysis with ICD-10 and ICD-10-CM
abstract
OBJECTIVE: To study the newly adopted International Classification of Diseases 11th revision (ICD-11) and compare it to the International Classification of Diseases 10th revision (ICD-10) and International Classification of Diseases 10th revision-Clinical Modification (ICD-10-CM). MATERIALS AND METHODS: : Data files and maps were downloaded from the World Health Organization (WHO) website and through the application programming interfaces. A round trip method based on the WHO maps was used to identify equivalent codes between ICD-10 and ICD-11, which were validated by limited manual review. ICD-11 terms were mapped to ICD-10-CM through normalized lexical mapping. ICD-10-CM codes in 6 disease areas were also manually recoded in ICD-11. RESULTS: Excluding the chapters for traditional medicine, functioning assessment, and extension codes for postcoordination, ICD-11 has 14 622 leaf codes (codes that can be used in coding) compared to ICD-10 and ICD-10-CM, which has 10 607 and 71 932 leaf codes, respectively. We identified 4037 pairs of ICD-10 and ICD-11 codes that were equivalent (estimated accuracy of 96%) by our round trip method. Lexical matching between ICD-11 and ICD-10-CM identified 4059 pairs of possibly equivalent codes. Manual recoding showed that 60% of a sample of 388 ICD-10-CM codes could be fully represented in ICD-11 by precoordinated codes or postcoordination. CONCLUSION: In ICD-11, there is a moderate increase in the number of codes over ICD-10. With postcoordination, it is possible to fully represent the meaning of a high proportion of ICD-10-CM codes, especially with the addition of a limited number of extension codes.
Kin Wah Fung, Julia Xu, Olivier Bodenreider
J. Am. Medical Informatics Assoc.2
2019 The Use of Inter-terminology Maps for the Creation and Maintenance of Value Sets
Kin Wah Fung, Julia Xu, Sigfried Gold
AMIA2
2018 Re-purposing the ICD-9-CM Procedures Index for Coding in ICD-10-PCS and SNOMED CT
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Ariel Busquets
AMIA2
2017 Achieving Logical Equivalence between SNOMED CT and ICD-10-PCS Surgical Procedures
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Arabella D'Havé
AMIA2
2016 Leveraging Lexical Matching and Ontological Alignment to Map SNOMED CT Surgical Procedures to ICD-10-PCS
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Arabella D'Havé
AMIA2
2015 An exploration of the properties of the CORE problem list subset and how it facilitates the implementation of SNOMED CT
abstract
OBJECTIVE: Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) is the emergent international health terminology standard for encoding clinical information in electronic health records. The CORE Problem List Subset was created to facilitate the terminology's implementation. This study evaluates the CORE Subset's coverage and examines its growth pattern as source datasets are being incorporated. METHODS: Coverage of frequently used terms and the corresponding usage of the covered terms were assessed by "leave-one-out" analysis of the eight datasets constituting the current CORE Subset. The growth pattern was studied using a retrospective experiment, growing the Subset one dataset at a time and examining the relationship between the size of the starting subset and the coverage of frequently used terms in the incoming dataset. Linear regression was used to model that relationship. RESULTS: On average, the CORE Subset covered 80.3% of the frequently used terms of the left-out dataset, and the covered terms accounted for 83.7% of term usage. There was a significant positive correlation between the CORE Subset's size and the coverage of the frequently used terms in an incoming dataset. This implies that the CORE Subset will grow at a progressively slower pace as it gets bigger. CONCLUSION: The CORE Problem List Subset is a useful resource for the implementation of Systematized Nomenclature of Medicine Clinical Terms in electronic health records. It offers good coverage of frequently used terms, which account for a high proportion of term usage. If future datasets are incorporated into the CORE Subset, it is likely that its size will remain small and manageable.
Kin Wah Fung, Julia Xu
J. Am. Medical Informatics Assoc.2
2013 Rule-based support system for multiple UMLS semantic type assignments
James Geller, Zhe He 0001, Yehoshua Perl, C. Paul Morrey, Julia Xu
J. Biomed. Informatics5
2012 A study of terminology auditors' performance for UMLS semantic type assignments
Huanying Gu, Gai Elhanan, Yehoshua Perl, George Hripcsak, James J. Cimino, Julia Xu, Yan Chen 0009, James Geller, C. Paul Morrey
J. Biomed. Informatics6