EDBT 2026 Demo / reviewers in the wild / expert
Kin Wah Fung
dblp:70/7340
· DBLP profile ↗
36ranked-venue papers
23as first author
7since 2021 · last 2024
0000-0003-0593-5377ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 23 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Promoting interoperability between SNOMED CT and ICD-11: lessons learned from the pilot project mapping between SNOMED CT and the ICD-11 FoundationabstractOBJECTIVE: To explore the feasibility and challenges of mapping between SNOMED CT and the ICD-11 Foundation in both directions, SNOMED International and the World Health Organization conducted a pilot mapping project between September 2021 and August 2022. MATERIALS AND METHODS: Phase 1 mapped ICD-11 Foundation entities from the endocrine diseases chapter, excluding malignant neoplasms, to SNOMED CT. In phase 2, SNOMED CT concepts equivalent to those covered by the ICD-11 entities in phase 1 were mapped to the ICD-11 Foundation. The goal was to identify equivalence between an ICD-11 Foundation entity and a SNOMED CT concept. Postcoordination was used for mapping to ICD-11. Each map was done twice independently, the results were compared, and discrepancies were reconciled. RESULTS: In phase 1, 59% of 637 ICD-11 Foundation entities had an exact match in SNOMED CT. In phase 2, 32% of 1893 SNOMED CT concepts had an exact match in the ICD-11 Foundation, and postcoordination added 15% of exact match. Challenges encountered included non-synonymous synonyms, mismatch in granularity, composite conditions, and residual categories. CONCLUSION: This pilot project shed light on the tremendous amount of effort required to create a map between the 2 coding systems and uncovered some common challenges. Future collaborative work between SNOMED International and WHO will likely benefit from its findings. It is recommended that the 2 organizations should clarify goals and use cases of mapping, provide adequate resources, set up a road map, and reconsider their original proposal of incorporating SNOMED CT into the ICD-11 Foundation ontology. Kin Wah Fung, Julia Xu, Hazel Brear, Alana Lane, Maggie Lau, Austen Wong, Arabella D'Havé |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Mapping 3 procedure coding systems to the International Classification of Health Interventions (ICHI): coverage and challengesabstractOBJECTIVE: To study the coverage and challenges in mapping 3 national and international procedure coding systems to the International Classification of Health Interventions (ICHI). MATERIALS AND METHODS: We identified 300 commonly used codes each from SNOMED CT, ICD-10-PCS, and CCI (Canadian Classification of Health Interventions) and mapped them to ICHI. We evaluated the level of match at the ICHI stem code and Foundation Component levels. We used postcoordination (modification of existing codes by adding other codes) to improve matching. Failure analysis was done for cases where full representation was not achieved. We noted and categorized potential problems that we encountered in ICHI, which could affect the accuracy and consistency of mapping. RESULTS: Overall, among the 900 codes from the 3 sources, 286 (31.8%) had full match with ICHI stem codes, 222 (24.7%) had full match with Foundation entities, and 231 (25.7%) had full match with postcoordination. 143 codes (15.9%) could only be partially represented even with postcoordination. A small number of SNOMED CT and ICD-10-PCS codes (18 codes, 2% of total), could not be mapped because the source codes were underspecified. We noted 4 categories of problems in ICHI-redundancy, missing elements, modeling issues, and naming issues. CONCLUSION: Using the full range of mapping options, at least three-quarters of the commonly used codes in each source system achieved a full match. For the purpose of international statistical reporting, full matching may not be an essential requirement. However, problems in ICHI that could result in suboptimal maps should be addressed. Kin Wah Fung, Julia Xu, Filip Ameye, Lisa Burelle, Janice Macneil |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | A practical strategy to use the ICD-11 for morbidity coding in the United States without a clinical modificationabstractOBJECTIVE: The aim of this study was to derive and evaluate a practical strategy of replacing ICD-10-CM codes by ICD-11 for morbidity coding in the United States, without the creation of a Clinical Modification. MATERIALS AND METHODS: A stepwise strategy is described, using first the ICD-11 stem codes from the Mortality and Morbidity Statistics (MMS) linearization, followed by exposing Foundation entities, then adding postcoordination (with existing codes and adding new stem codes if necessary), with creating new stem codes as the last resort. The strategy was evaluated by recoding 2 samples of ICD-10-CM codes comprised of frequently used codes and all codes from the digestive diseases chapter. RESULTS: Among the 1725 ICD-10-CM codes examined, the cumulative coverage at the stem code, Foundation, and postcoordination levels are 35.2%, 46.5% and 89.4% respectively. 7.1% of codes require new extension codes and 3.5% require new stem codes. Among the new extension codes, severity scale values and anatomy are the most common categories. 5.5% of codes are not one-to-one matches (1 ICD-10-CM code matched to 1 ICD-11 stem code or Foundation entity) which could be potentially challenging. CONCLUSION: Existing ICD-11 content can achieve full representation of almost 90% of ICD-10-CM codes, provided that postcoordination can be used and the coding guidelines and hierarchical structures of ICD-10-CM and ICD-11 can be harmonized. The various options examined in this study should be carefully considered before embarking on the traditional approach of a full-fledged ICD-11-CM. Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learningabstractOBJECTIVE: Use heuristic, deep learning (DL), and hybrid AI methods to predict semantic group (SG) assignments for new UMLS Metathesaurus atoms, with target accuracy ≥95%. MATERIALS AND METHODS: We used train-test datasets from successive 2020AA-2022AB UMLS Metathesaurus releases. Our heuristic "waterfall" approach employed a sequence of 7 different SG prediction methods. Atoms not qualifying for a method were passed on to the next method. The DL approach generated BioWordVec and SapBERT embeddings for atom names, BioWordVec embeddings for source vocabulary names, and BioWordVec embeddings for atom names of the second-to-top nodes of an atom's source hierarchy. We fed a concatenation of the 4 embeddings into a fully connected multilayer neural network with an output layer of 15 nodes (one for each SG). For both approaches, we developed methods to estimate the probability that their predicted SG for an atom would be correct. Based on these estimations, we developed 2 hybrid SG prediction methods combining the strengths of heuristic and DL methods. RESULTS: The heuristic waterfall approach accurately predicted 94.3% of SGs for 1 563 692 new unseen atoms. The DL accuracy on the same dataset was also 94.3%. The hybrid approaches achieved an average accuracy of 96.5%. CONCLUSION: Our study demonstrated that AI methods can predict SG assignments for new UMLS atoms with sufficient accuracy to be potentially useful as an intermediate step in the time-consuming task of assigning new atoms to UMLS concepts. We showed that for SG prediction, combining heuristic methods and DL methods can produce better results than either alone. Randolph A. Miller, Olivier Bodenreider, Vinh Nguyen 0002, Kin Wah Fung |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Can ICD-11 Replace ICD-10-CM for Morbidity Coding in the U.S.?
Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider |
AMIA | 1 |
| 2021 | Evaluation of the International Classification of Health Interventions (ICHI) in the coding of common surgical proceduresabstractOBJECTIVE: To evaluate the International Classification of Health Interventions (ICHI) in the clinical and statistical use cases. MATERIALS AND METHODS: We identified 300 most-performed surgical procedures as represented by their display names in an electronic health record. For comparison with existing coding systems, we coded the procedures in ICHI, SNOMED CT, International Classification of Diseases (ICD)-10-PCS, and CCI (Canadian Classification of Health Interventions), using postcoordination (modification of existing codes by adding other codes), when applicable. Failure analysis was done for cases where full representation was not achieved. The ICHI encoding was further evaluated for adequacy to support statistical reporting by the Organisation for Economic Co-operation and Development (OECD) and European Union (EU) categories of surgical procedures. RESULTS: After deduplication, 229 distinct procedures remained. Without postcoordination, ICHI achieved full representation in 52.8%. A further 19.2% could be fully represented with postcoordination. SNOMED CT was the best performing overall, with 94.3% full representation without postcoordination, and 99.6% with postcoordination. Failure analysis showed that "method" and "target" constituted most of the missing information for ICHI encoding. For all OECD/EU surgical categories, ICHI coding was adequate to support statistical reporting. One OECD/EU category ("Hip replacement, secondary") required postcoordination for correct assignment. CONCLUSION: In the clinical use case of capturing information in the electronic health record, ICHI was outperformed by the clinically oriented procedure coding systems (SNOMED CT and CCI), but was comparable to ICD-10-PCS. Postcoordination could be an effective and efficient means of improving coverage. ICHI is generally adequate for the collection of international statistics. Kin Wah Fung, Julia Xu, Filip Ameye, Lisa Burelle, Janice Macneil |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Feasibility of replacing the ICD-10-CM with the ICD-11 for morbidity coding: A content analysisabstractOBJECTIVE: The study sought to assess the feasibility of replacing the International Classification of Diseases-Tenth Revision-Clinical Modification (ICD-10-CM) with the International Classification of Diseases-11th Revision (ICD-11) for morbidity coding based on content analysis. MATERIALS AND METHODS: The most frequently used ICD-10-CM codes from each chapter covering 60% of patients were identified from Medicare claims and hospital data. Each ICD-10-CM code was recoded in the ICD-11, using postcoordination (combination of codes) if necessary. Recoding was performed by 2 terminologists independently. Failure analysis was done for cases where full representation was not achieved even with postcoordination. After recoding, the coding guidance (inclusions, exclusions, and index) of the ICD-10-CM and ICD-11 codes were reviewed for conflict. RESULTS: Overall, 23.5% of 943 codes could be fully represented by the ICD-11 without postcoordination. Postcoordination is the potential game changer. It supports the full representation of 8.6% of 943 codes. Moreover, with the addition of only 9 extension codes, postcoordination supports the full representation of 35.2% of 943 codes. Coding guidance review identified potential conflicts in 10% of codes, but mostly not affecting recoding. The majority of the conflicts resulted from differences in granularity and default coding assumptions between the ICD-11 and ICD-10-CM. CONCLUSIONS: With some minor enhancements to postcoordination, the ICD-11 can fully represent almost 60% of the most frequently used ICD-10-CM codes. Even without postcoordination, 23.5% full representation is comparable to the 24.3% of ICD-9-CM codes with exact match in the ICD-10-CM, so migrating from the ICD-10-CM to the ICD-11 is not necessarily more disruptive than from the International Classification of Diseases-Ninth Revision-Clinical Modification to the ICD-10-CM. Therefore, the ICD-11 (without a CM) should be considered as a candidate to replace the ICD-10-CM for morbidity coding. Kin Wah Fung, Julia Xu, Shannon McConnell-Lamptey, Donna Pickett, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Comparison of Three International Terminologies for Medical Interventions and Procedures
Kin Wah Fung, Julia Xu, Filip Ameye |
AMIA | 1 |
| 2020 | The new International Classification of Diseases 11th edition: a comparative analysis with ICD-10 and ICD-10-CMabstractOBJECTIVE: To study the newly adopted International Classification of Diseases 11th revision (ICD-11) and compare it to the International Classification of Diseases 10th revision (ICD-10) and International Classification of Diseases 10th revision-Clinical Modification (ICD-10-CM). MATERIALS AND METHODS: : Data files and maps were downloaded from the World Health Organization (WHO) website and through the application programming interfaces. A round trip method based on the WHO maps was used to identify equivalent codes between ICD-10 and ICD-11, which were validated by limited manual review. ICD-11 terms were mapped to ICD-10-CM through normalized lexical mapping. ICD-10-CM codes in 6 disease areas were also manually recoded in ICD-11. RESULTS: Excluding the chapters for traditional medicine, functioning assessment, and extension codes for postcoordination, ICD-11 has 14 622 leaf codes (codes that can be used in coding) compared to ICD-10 and ICD-10-CM, which has 10 607 and 71 932 leaf codes, respectively. We identified 4037 pairs of ICD-10 and ICD-11 codes that were equivalent (estimated accuracy of 96%) by our round trip method. Lexical matching between ICD-11 and ICD-10-CM identified 4059 pairs of possibly equivalent codes. Manual recoding showed that 60% of a sample of 388 ICD-10-CM codes could be fully represented in ICD-11 by precoordinated codes or postcoordination. CONCLUSION: In ICD-11, there is a moderate increase in the number of codes over ICD-10. With postcoordination, it is possible to fully represent the meaning of a high proportion of ICD-10-CM codes, especially with the addition of a limited number of extension codes. Kin Wah Fung, Julia Xu, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Use of word and graph embedding to measure semantic relatedness between Unified Medical Language System conceptsabstractOBJECTIVE: The study sought to explore the use of deep learning techniques to measure the semantic relatedness between Unified Medical Language System (UMLS) concepts. MATERIALS AND METHODS: Concept sentence embeddings were generated for UMLS concepts by applying the word embedding models BioWordVec and various flavors of BERT to concept sentences formed by concatenating UMLS terms. Graph embeddings were generated by the graph convolutional networks and 4 knowledge graph embedding models, using graphs built from UMLS hierarchical relations. Semantic relatedness was measured by the cosine between the concepts' embedding vectors. Performance was compared with 2 traditional path-based (shortest path and Leacock-Chodorow) measurements and the publicly available concept embeddings, cui2vec, generated from large biomedical corpora. The concept sentence embeddings were also evaluated on a word sense disambiguation (WSD) task. Reference standards used included the semantic relatedness and semantic similarity datasets from the University of Minnesota, concept pairs generated from the Standardized MedDRA Queries and the MeSH (Medical Subject Headings) WSD corpus. RESULTS: Sentence embeddings generated by BioWordVec outperformed all other methods used individually in semantic relatedness measurements. Graph convolutional network graph embedding uniformly outperformed path-based measurements and was better than some word embeddings for the Standardized MedDRA Queries dataset. When used together, combined word and graph embedding achieved the best performance in all datasets. For WSD, the enhanced versions of BERT outperformed BioWordVec. CONCLUSIONS: Word and graph embedding techniques can be used to harness terms and relations in the UMLS to measure semantic relatedness between concepts. Concept sentence embedding outperforms path-based measurements and cui2vec, and can be further enhanced by combining with graph embedding. Kin Wah Fung |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | The Use of Inter-terminology Maps for the Creation and Maintenance of Value Sets
Kin Wah Fung, Julia Xu, Sigfried Gold |
AMIA | 1 |
| 2019 | Drug-drug Interaction Extraction via Transfer Learning
Kin Wah Fung, Dina Demner-Fushman |
AMIA | 2 |
| 2019 | Sharing of Individual Participant Data from Clinical Trials: General Comparison and HIV Use Case
Craig S. Mayer, Nick Williams, Sigfried Gold, Kin Wah Fung, Vojtech Huser |
AMIA | 4 |
| 2019 | Automated Discovery of Common Data Elements in HIV Clinical Trials Using an Interaction Network
Nick Williams, Sigfried Gold, Vojtech Huser, Kin Wah Fung, Craig S. Mayer |
AMIA | 4 |
| 2019 | A systematic approach for developing a corpus of patient reported adverse drug events: A case study for SSRI and SNRI medications
Maryam Zolnoori, Kin Wah Fung, Timothy B. Patrick, Paul A. Fontelo, Hadi Kharrazi, Anthony Faiola, Yi Shuan Shirley Wu, Christina Eldredge, Jake Luo, Mike Conway, Jiaxi Zhu, Soo Kyung Park, Kelly Xu, Hamideh Moayyed, Somaieh Goudarzvand |
J. Biomed. Informatics | 2 |
| 2018 | Adverse Reactions and Drug-Drug Interaction Extraction tracks at the Text Analysis Conference (TAC)
Dina Demner-Fushman, Joseph M. Tonning, Kin Wah Fung, Phong Do, Richard D. Boyce, Kirk Roberts |
AMIA | 3 |
| 2018 | Re-purposing the ICD-9-CM Procedures Index for Coding in ICD-10-PCS and SNOMED CT
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Ariel Busquets |
AMIA | 1 |
| 2018 | Utilizing Consumer Health Posts to Identify Underlying Factors Associated with Patients' Attitudes towards Antidepressants
Maryam Zolnoori, Kin Wah Fung, Paul A. Fontelo, Hadi Kharrazi, Anthony Faiola, Yi Shuan Shirley Wu, Virginia C. Stoffel, Timothy B. Patrick |
AMIA | 2 |
| 2018 | A value set for documenting adverse reactions in electronic health recordsabstractObjective: To develop a comprehensive value set for documenting and encoding adverse reactions in the allergy module of an electronic health record. Materials and Methods: We analyzed 2 471 004 adverse reactions stored in Partners Healthcare's Enterprise-wide Allergy Repository (PEAR) of 2.7 million patients. Using the Medical Text Extraction, Reasoning, and Mapping System, we processed both structured and free-text reaction entries and mapped them to Systematized Nomenclature of Medicine - Clinical Terms. We calculated the frequencies of reaction concepts, including rare, severe, and hypersensitivity reactions. We compared PEAR concepts to a Federal Health Information Modeling and Standards value set and University of Nebraska Medical Center data, and then created an integrated value set. Results: We identified 787 reaction concepts in PEAR. Frequently reported reactions included: rash (14.0%), hives (8.2%), gastrointestinal irritation (5.5%), itching (3.2%), and anaphylaxis (2.5%). We identified an additional 320 concepts from Federal Health Information Modeling and Standards and the University of Nebraska Medical Center to resolve gaps due to missing and partial matches when comparing these external resources to PEAR. This yielded 1106 concepts in our final integrated value set. The presence of rare, severe, and hypersensitivity reactions was limited in both external datasets. Hypersensitivity reactions represented roughly 20% of the reactions within our data. Discussion: We developed a value set for encoding adverse reactions using a large dataset from one health system, enriched by reactions from 2 large external resources. This integrated value set includes clinically important severe and hypersensitivity reactions. Conclusion: This work contributes a value set, harmonized with existing data, to improve the consistency and accuracy of reaction documentation in electronic health records, providing the necessary building blocks for more intelligent clinical decision support for allergies and adverse reactions. Foster R. Goss, Kenneth H. Lai, Maxim Topaz, Warren W. Acker, Leigh Kowalski, Joseph M. Plasek, Kimberly G. Blumenthal, Diane L. Seger, Sarah P. Slight, Kin Wah Fung, Frank Y. Chang, David W. Bates, Li Zhou 0007 |
J. Am. Medical Informatics Assoc. | 10 |
| 2017 | Achieving Logical Equivalence between SNOMED CT and ICD-10-PCS Surgical Procedures
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Arabella D'Havé |
AMIA | 1 |
| 2017 | Comparison of three commercial knowledge bases for detection of drug-drug interactions in clinical decision supportabstractOBJECTIVE: To compare 3 commercial knowledge bases (KBs) used for detection and avoidance of potential drug-drug interactions (DDIs) in clinical practice. METHODS: Drugs in the DDI tables from First DataBank (FDB), Micromedex, and Multum were mapped to RxNorm. The KBs were compared at the clinical drug, ingredient, and DDI rule levels. The KBs were evaluated against a reference list of highly significant DDIs from the Office of the National Coordinator for Health Information Technology (ONC). The KBs and the ONC list were applied to a prescription data set to simulate their use in clinical decision support. RESULTS: The KBs contained 1.6 million (FDB), 4.5 million (Micromedex), and 4.8 million (Multum) clinical drug pairs. Altogether, there were 8.6 million unique pairs, of which 79% were found only in 1 KB and 5% in all 3 KBs. However, there was generally more agreement than disagreement in the severity rankings, especially in the contraindicated category. The KBs covered 99.8-99.9% of the alerts of the ONC list and would have generated 25 (FDB), 145 (Micromedex), and 84 (Multum) alerts per 1000 prescriptions. CONCLUSION: The commercial KBs differ considerably in size and quantity of alerts generated. There is less variability in severity ranking of DDIs than suggested by previous studies. All KBs provide very good coverage of the ONC list. More work is needed to standardize the editorial policies and evidence for inclusion of DDIs to reduce variation among knowledge sources and improve relevance. Some DDIs considered contraindicated in all 3 KBs might be possible candidates to add to the ONC list. Kin Wah Fung, Joan Kapusnik-Uner, Jean Cunningham, Stefanie Higby-Baker, Olivier Bodenreider |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Leveraging Lexical Matching and Ontological Alignment to Map SNOMED CT Surgical Procedures to ICD-10-PCS
Kin Wah Fung, Julia Xu, Filip Ameye, Arturo Romero-Gutiérrez, Arabella D'Havé |
AMIA | 1 |
| 2016 | Development of an Oncology Subset of SNOMED CT Based on Patient Notes
Sina Madani, Jerry Henderson, Kin Wah Fung |
AMIA | 3 |
| 2016 | Expert Recommendations on Redesigning Drug Allergy Alerts in Electronic Health Record Systems
Maxim Topaz, Foster R. Goss, Kimberly G. Blumenthal, Kenneth H. Lai, Diane L. Seger, Sarah P. Slight, Paige G. Wickner, George A. Robinson, Kin Wah Fung, Robert C. McClure, Shelly Spiro, Warren W. Acker, David W. Bates |
AMIA | 9 |
| 2015 | An exploration of the properties of the CORE problem list subset and how it facilitates the implementation of SNOMED CTabstractOBJECTIVE: Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) is the emergent international health terminology standard for encoding clinical information in electronic health records. The CORE Problem List Subset was created to facilitate the terminology's implementation. This study evaluates the CORE Subset's coverage and examines its growth pattern as source datasets are being incorporated. METHODS: Coverage of frequently used terms and the corresponding usage of the covered terms were assessed by "leave-one-out" analysis of the eight datasets constituting the current CORE Subset. The growth pattern was studied using a retrospective experiment, growing the Subset one dataset at a time and examining the relationship between the size of the starting subset and the coverage of frequently used terms in the incoming dataset. Linear regression was used to model that relationship. RESULTS: On average, the CORE Subset covered 80.3% of the frequently used terms of the left-out dataset, and the covered terms accounted for 83.7% of term usage. There was a significant positive correlation between the CORE Subset's size and the coverage of the frequently used terms in an incoming dataset. This implies that the CORE Subset will grow at a progressively slower pace as it gets bigger. CONCLUSION: The CORE Problem List Subset is a useful resource for the implementation of Systematized Nomenclature of Medicine Clinical Terms in electronic health records. It offers good coverage of frequently used terms, which account for a high proportion of term usage. If future datasets are incorporated into the CORE Subset, it is likely that its size will remain small and manageable. Kin Wah Fung, Julia Xu |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Coverage of Rare Disease Names in Standard Terminologies and Implications for Patients, Providers, and Research
Kin Wah Fung, Rachel L. Richesson, Olivier Bodenreider |
AMIA | 1 |
| 2014 | Facilitating Reconciliation of Inter-Annotator Disagreements
Johann Stan, Olivier Bodenreider, Kin Wah Fung, Dina Demner-Fushman |
AMIA | 3 |
| 2013 | Extracting drug indication information from structured product labels using natural language processingabstractOBJECTIVE: To extract drug indications from structured drug labels and represent the information using codes from standard medical terminologies. MATERIALS AND METHODS: We used MetaMap and other publicly available resources to extract information from the indications section of drug labels. Drugs and indications were encoded by RxNorm and UMLS identifiers respectively. A sample was manually reviewed. We also compared the results with two independent information sources: National Drug File-Reference Terminology and the Semantic Medline project. RESULTS: A total of 6797 drug labels were processed, resulting in 19 473 unique drug-indication pairs. Manual review of 298 most frequently prescribed drugs by seven physicians showed a recall of 0.95 and precision of 0.77. Inter-rater agreement (Fleiss κ) was 0.713. The precision of the subset of results corroborated by Semantic Medline extractions increased to 0.93. DISCUSSION: Correlation of a patient's medical problems and drugs in an electronic health record has been used to improve data quality and reduce medication errors. Authoritative drug indication information is available from drug labels, but not in a format readily usable by computer applications. Our study shows that it is feasible to use publicly available natural language processing resources to extract drug indications from drug labels. The same method can be applied to other sections of the drug label-for example, adverse effects, contraindications. CONCLUSIONS: It is feasible to use publicly available natural language processing tools to extract indication information from freely available drug labels. Named entity recognition sources (eg, MetaMap) provide reasonable recall. Combination with other data sources provides higher precision. Kin Wah Fung, Chiang S. Jao, Dina Demner-Fushman |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Synergism between the Mapping Projects from SNOMED CT to ICD-10 and ICD-10-CM
Kin Wah Fung, Junchuan Xu |
AMIA | 1 |
| 2012 | Handling Age Specification in the SNOMED CT to ICD-10-CM Cross-map
Junchuan Xu, Kin Wah Fung |
AMIA | 2 |
| 2010 | The UMLS-CORE project: a study of the problem list terminologies used in large healthcare institutionsabstractOBJECTIVE: To study existing problem list terminologies (PLTs), and to identify a subset of concepts based on standard terminologies that occur frequently in problem list data. DESIGN: Problem list terms and their usage frequencies were collected from large healthcare institutions. MEASUREMENT: The pattern of usage of the terms was analyzed. The local terms were mapped to the Unified Medical Language System (UMLS). Based on the mapped UMLS concepts, the degree of overlap between the PLTs was analyzed. RESULTS: Six institutions submitted 76,237 terms and their usage frequencies in 14 million patients. The distribution of usage was highly skewed. On average, 21% of unique terms already covered 95% of usage. The most frequently used 14,395 terms, representing the union of terms that covered 95% of usage in each institution, were exhaustively mapped to the UMLS. 13,261 terms were successfully mapped to 6776 UMLS concepts. Less frequently used terms were generally less 'mappable' to the UMLS. The mean pairwise overlap of the PLTs was only 21% (median 19%). Concepts that were shared among institutions were used eight times more often than concepts unique to one institution. A SNOMED Problem List Subset of frequently used problem list concepts was identified. CONCLUSIONS: Most of the frequently used problem list terms could be found in standard terminologies. The overlap between existing PLTs was low. The use of the SNOMED Problem List Subset will save developmental effort, reduce variability of PLTs, and enhance interoperability of problem list data. Kin Wah Fung, Clement J. McDonald, Suresh Srinivasan |
J. Am. Medical Informatics Assoc. | 1 |
| 2008 | RxTerms - a drug interface terminology derived from RxNorm
Kin Wah Fung, Clement J. McDonald, Bruce E. Bray |
AMIA | 1 |
| 2006 | Who is Using the UMLS and How - Insights from the UMLS User Annual Reports
Kin Wah Fung, William T. Hole, Suresh Srinivasan |
AMIA | 1 |
| 2005 | Utilizing the UMLS for Semantic Mapping between Terminologies
Kin Wah Fung, Olivier Bodenreider |
AMIA | 1 |
| 2005 | Research Paper: Integrating SNOMED CT into the UMLS: An Exploration of Different Views of Synonymy and Quality of EditingabstractOBJECTIVE: The integration of SNOMED CT into the Unified Medical Language System (UMLS) involved the alignment of two views of synonymy that were different because the two vocabulary systems have different intended purposes and editing principles. The UMLS is organized according to one view of synonymy, but its structure also represents all the individual views of synonymy present in its source vocabularies. Despite progress in knowledge-based automation of development and maintenance of vocabularies, manual curation is still the main method of determining synonymy. The aim of this study was to investigate the quality of human judgment of synonymy. DESIGN: Sixty pairs of potentially controversial SNOMED CT synonyms were reviewed by 11 domain vocabulary experts (six UMLS editors and five noneditors), and scores were assigned according to the degree of synonymy. MEASUREMENTS: The synonymy scores of each subject were compared to the gold standard (the overall mean synonymy score of all subjects) to assess accuracy. Agreement between UMLS editors and noneditors was measured by comparing the mean synonymy scores of editors to noneditors. RESULTS: Average accuracy was 71% for UMLS editors and 75% for noneditors (difference not statistically significant). Mean scores of editors and noneditors showed significant positive correlation (Spearman's rank correlation coefficient 0.654, two-tailed p < 0.01) with a concurrence rate of 75% and an interrater agreement kappa of 0.43. CONCLUSION: The accuracy in the judgment of synonymy was comparable for UMLS editors and nonediting domain experts. There was reasonable agreement between the two groups. Kin Wah Fung, William T. Hole, Stuart J. Nelson, Suresh Srinivasan, Tammy Powell, Laura Roth |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | Will Decision Support in Medications Order Entry Save Money? A Return On Investment Analysis of the Case of the Hong Kong Hospital Authority
Kin Wah Fung, Lynn Harold Vogel |
AMIA | 1 |