EDBT 2026 Demo / reviewers in the wild / expert
Licong Cui
dblp:51/1932
· DBLP profile ↗
60ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0001-5549-8780ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A conceptual model for provider scheduling: insights from an EHR implementationabstractOBJECTIVE: We present a set of definitions and a conceptual model to support primary and consulting physician schedule integration into the electronic health record (EHR) in an inpatient setting and show how utilization of this functionality supports patient-centered communication. MATERIALS AND METHODS: Our institution transitioned to the Epic EHR and implemented modules to connect primary and consulting provider schedules from external scheduling systems to secure messaging within an inpatient EHR context. We evaluated legacy functionality, met with provider groups to map their shifts to hospital teams, built a crosswalk tool to extract, transform, and load data from the scheduling systems to the EHR, evaluated the utilization, and assessed issues related to the implementation. We used our experience from the project to develop a set of definitions and a conceptual model for provider scheduling. RESULTS: We met with over 100 groups to map over 2000 shifts to nearly 700 teams across 15 facilities in our health system. Utilization was high with an average of 6500 on-call provider searches per day in the 30 days following implementation. The conceptual model for inpatient provider scheduling defines 11 terms. DISCUSSION: Our definitions and conceptual model sufficiently represent the inpatient provider scheduling domain as evidenced by high initial utilization and few reported defects. The standardized terminology for provider scheduling aids integration of scheduling data into the EHR. CONCLUSION: Our successful integration of real-time scheduling data within the EHR guided development of a provider scheduling conceptual model. Standardized provider scheduling terminology promotes interoperability of scheduling systems. Jeremy B. Hill, Elmer V. Bernstam, Licong Cui, Peter V. Killoran |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Temporal Ensemble Logic for Integrative Representation of the Entirety of Clinical Trials
Yan Huang 0034, Rashmie Abeysinghe, Zenan Sun, Pengze Li, Xing He 0003, Shiqiang Tao, Cui Tao, Jiang Bian 0001, Licong Cui, Guo-Qiang Zhang 0001 |
TIME | 11 |
| 2025 | Quantitatively assessing the impact of the quality of SNOMED CT subtype hierarchy on cohort queriesabstractOBJECTIVE: SNOMED CT provides a standardized terminology for clinical concepts, allowing cohort queries over heterogeneous clinical data including Electronic Health Records (EHRs). While it is intuitive that missing and inaccurate subtype (or is-a) relations in SNOMED CT reduce the recall and precision of cohort queries, the extent of these impacts has not been formally assessed. This study fills this gap by developing quantitative metrics to measure these impacts and performing statistical analysis on their significance. MATERIAL AND METHODS: We used the Optum de-identified COVID-19 Electronic Health Record dataset. We defined micro-averaged and macro-averaged recall and precision metrics to assess the impact of missing and inaccurate is-a relations on cohort queries. Both practical and simulated analyses were performed. Practical analyses involved 407 missing and 48 inaccurate is-a relations confirmed by domain experts, with statistical testing using Wilcoxon signed-rank tests. Simulated analyses used two random sets of 400 is-a relations to simulate missing and inaccurate is-a relations. RESULTS: Wilcoxon signed-rank tests from both practical and simulated analyses (P-values < .001) showed that missing is-a relations significantly reduced the micro- and macro-averaged recall, and inaccurate is-a relations significantly reduced the micro- and macro-averaged precision. DISCUSSION: The introduced impact metrics can assist SNOMED CT maintainers in prioritizing critical hierarchical defects for quality enhancement. These metrics are generally applicable for assessing the quality impact of a terminology's subtype hierarchy on its cohort query applications. CONCLUSION: Our results indicate a significant impact of missing and inaccurate is-a relations in SNOMED CT on the recall and precision of cohort queries. Our work highlights the importance of high-quality terminology hierarchy for cohort queries over EHR data and provides valuable insights for prioritizing quality improvements of SNOMED CT's hierarchy. Xubing Hao, Yan Huang 0034, Jay Shi, Rashmie Abeysinghe, Cui Tao, Kirk Roberts, Guo-Qiang Zhang 0001, Licong Cui |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | CDEMapper: enhancing National Institutes of Health common data element use with large language modelsabstractOBJECTIVE: Common Data Elements (CDEs) standardize data collection and sharing across studies, enhancing data interoperability and improving research reproducibility. However, implementing CDEs presents challenges due to the broad range and variety of data elements. This study aims to develop a CDE mapping tool to bridge the gap between local data elements and National Institutes of Health (NIH) CDEs. METHODS: We propose CDEMapper, a large language model (LLM)-powered mapping tool designed to assist in mapping local data elements to NIH CDEs. CDEMapper has 3 core modules: (1) CDE indexing and embeddings. NIH CDEs were indexed and embedded to support semantic search; (2) CDE recommendations. The tool combines Elasticsearch (BM25 methods) with GPT services to recommend candidate CDEs and their permissible values; and (3) Human review. Users review and select the best match for their data elements and value sets. We evaluate the tool's recommendation accuracy and usability against manual annotations and testing. RESULTS: CDEMapper offers a publicly available, LLM-powered, and intuitive user interface that consolidates essential and advanced mapping services into a streamlined pipeline. The evaluation results demonstrated that the augmented BM25 with GPT embeddings and a GPT ranker achieved the overall best performance. The usability test also highlighted the effectiveness and efficiency of our tool. DISCUSSIONS AND CONCLUSIONS: This work opens up the potential of using LLMs to assist with CDE mapping when aligning local data elements with NIH CDEs. Additionally, this effort helps researchers better understand the gaps between their data elements and NIH CDEs while promoting CDE reusability. Yan Wang 0015, Jimin Huang, Yujia Zhou 0003, Xubing Hao, Pritham Ram, Lingfei Qian, Qianqian Xie, Ruey-Ling Weng, Fongci Lin, Licong Cui, Xiaoqian Jiang, Hua Xu 0001, Na Hong |
J. Am. Medical Informatics Assoc. | 13 |
| 2025 | A comparative study of recent large language models on generating hospital discharge summaries for lung cancer patients
Fang Li 0011, Na Hong, Manqi Li, Kirk Roberts, Licong Cui, Cui Tao, Hua Xu 0001 |
J. Biomed. Informatics | 6 |
| 2024 | Exploring Pre-trained Language Models for Vocabulary Alignment in the UMLS
Xubing Hao, Rashmie Abeysinghe, Jay Shi, Licong Cui |
AIME (1) | 4 |
| 2023 | A deep learning approach to identify missing is-a relations in SNOMED CTabstractOBJECTIVE: SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. MATERIALS AND METHODS: Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. RESULTS: We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. CONCLUSIONS: The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT. Rashmie Abeysinghe, Fengbo Zheng, Elmer V. Bernstam, Jay Shi, Olivier Bodenreider, Licong Cui |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | Temporal Cohort Logic
Guo-Qiang Zhang 0001, Yan Huang 0034, Licong Cui |
AMIA | 4 |
| 2022 | Automated Identification of Missing IS-A Relations in the Human Phenotype Ontology
Maryamsadat Mohtashamian, Rashmie Abeysinghe, Xubing Hao, Hua Xu 0001, Licong Cui |
AMIA | 6 |
| 2022 | A substring replacement approach for identifying missing IS-A relations in SNOMED CTabstractBiomedical ontologies provide formalized information and knowledge in the biomedical domain. Over the years, biomedical ontologies have played an important role in facilitating biomedical research and applications. Common quality issues of biomedical ontologies include inconsistent naming of concepts, redundant concepts, redundant relations, incomplete/incorrect concept definitions, and incomplete/incorrect class hierarchies. In this work, we focus on addressing the incompleteness of the class hierarchy in SNOMED CT. We develop a substring replacement approach, leveraging concepts' lexical features and existing IS-A relations to identify potential missing IS-A relations in SNOMED CT. To evaluate the effectiveness of our approach, we performed both automated and manual validation. For the automated evaluation, we leverage relations from external terminologies in the Unified Medical Language System (UMLS) to validate the identified missing IS-A relations. For the manual validation, a randomly selected 100 samples from the results are reviewed by a domain expert. Applying our approach to the March 2022 release of SNOMED CT US Edition, we identified 3,228 potential missing IS-A relations, among which 63 were validated through the UMLS. The evaluation by the domain expert revealed that 89 out of 100 (a precision of 89%) missing IS-A relations are valid cases, showing the effectiveness of this substring replacement approach to facilitate the quality assurance of IS-A relations in SNOMED CT. Xubing Hao, Rashmie Abeysinghe, Jay Shi, Licong Cui |
BIBM | 4 |
| 2022 | Identifying Missing IS-A Relations in Orphanet Rare Disease OntologyabstractThe Orphanet Rare Disease Ontology (ORDO) provides a structured vocabulary encapsulating rare diseases. Downstream applications of ORDO depend on its accuracy to effectively perform their tasks. In this paper, we implement an automated quality assurance pipeline to identify missing is-a relations in ORDO. We first obtain lexical features from concept names. Then we generate related and unrelated feature sharing concept-pairs, where a feature sharing concept-pair can further generate derived term-pairs. If an unrelated and related feature sharing concept-pair generate the same derived term-pair, then we suggest a potential missing is-a relation between the unrelated feature sharing concept-pair. Applying this approach on the 202206-27 release of ORDO, we obtained 705 potential missing is-a relations. Leveraging external ontological information in the Unified Medical Language System, we validated 164 missing is-a relations. This indicates that our approach is a promising way to audit is-a relations in ORDO, even though further domain expert evaluation is still needed to validate the remaining potential missing is-a relations identified. Maryamsadat Mohtashamian, Rashmie Abeysinghe, Xubing Hao, Licong Cui |
BIBM | 4 |
| 2022 | An evidence-based lexical pattern approach for quality assurance of Gene Ontology relationsabstractGene Ontology (GO) is widely used in the biological domain. It is the most comprehensive ontology providing formal representation of gene functions (GO concepts) and relations between them. However, unintentional quality defects (e.g. missing or erroneous relations) in GO may exist due to the large size of GO concepts and complexity of GO structures. Such quality defects would impact the results of GO-based analyses and applications. In this work, we introduce a novel evidence-based lexical pattern approach for quality assurance of GO relations. We leverage two layers of evidence to suggest potentially missing relations in GO as follows. We first utilize related concept pairs (i.e. existing relations) in GO to extract relationship-specific lexical patterns, which serve as the first layer evidence to automatically suggest potentially missing relations between unrelated concept pairs. For each suggested missing relation, we further identify two other existing relations as the second layer of evidence that resemble the difference between the missing relation and the existing relation based on which the missing relation is suggested. Applied to the 15 December 2021 release of GO, this approach suggested a total of 866 potentially missing relations. Local domain experts evaluated the entire set of potentially missing relations, and identified 821 as missing relations and 45 indicate erroneous existing relations. We submitted these findings to the GO consortium for further validation and received encouraging feedback. These indicate that our evidence-based approach can be utilized to uncover missing relations and erroneous existing relations in GO. Rashmie Abeysinghe, Yuntao Yang, Mason Bartels, W. Jim Zheng, Licong Cui |
Briefings Bioinform. | 5 |
| 2022 | Toward a standard formal semantic representation of the model card reportabstractBACKGROUND: Model card reports aim to provide informative and transparent description of machine learning models to stakeholders. This report document is of interest to the National Institutes of Health's Bridge2AI initiative to address the FAIR challenges with artificial intelligence-based machine learning models for biomedical research. We present our early undertaking in developing an ontology for capturing the conceptual-level information embedded in model card reports. RESULTS: Sourcing from existing ontologies and developing the core framework, we generated the Model Card Report Ontology. Our development efforts yielded an OWL2-based artifact that represents and formalizes model card report information. The current release of this ontology utilizes standard concepts and properties from OBO Foundry ontologies. Also, the software reasoner indicated no logical inconsistencies with the ontology. With sample model cards of machine learning models for bioinformatics research (HIV social networks and adverse outcome prediction for stent implantation), we showed the coverage and usefulness of our model in transforming static model card reports to a computable format for machine-based processing. CONCLUSIONS: The benefit of our work is that it utilizes expansive and standard terminologies and scientific rigor promoted by biomedical ontologists, as well as, generating an avenue to make model cards machine-readable using semantic web technology. Our future goal is to assess the veracity of our model and later expand the model to include additional concepts to address terminological gaps. We discuss tools and software that will utilize our ontology for potential application services. Muhammad Amith, Licong Cui, Degui Zhi, Kirk Roberts, Xiaoqian Jiang, Fang Li 0011, Evan Yu, Cui Tao |
BMC Bioinform. | 2 |
| 2022 | Towards quality improvement of vaccine concept mappings in the OMOP vocabulary with a semi-automated methodabstractThe Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) provides a unified model to integrate disparate real-world data (RWD) sources. An integral part of the OMOP CDM is the Standardized Vocabularies (henceforth referred to as the OMOP vocabulary), which enables organization and standardization of medical concepts across various clinical domains of the OMOP CDM. For concepts with the same meaning from different source vocabularies, one is designated as the standard concept, while the others are specified as non-standard or source concepts and mapped to the standard one. However, due to the heterogeneity of source vocabularies, there may exist mapping issues such as erroneous mappings and missing mappings in the OMOP vocabulary, which could affect the results of downstream analyses with RWD. In this paper, we focus on quality assurance of vaccine concept mappings in the OMOP vocabulary, which is necessary to accurately harness the power of RWD on vaccines. We introduce a semi-automated lexical approach to audit vaccine mappings in the OMOP vocabulary. We generated two types of vaccine-pairs: mapped and unmapped, where mapped vaccine-pairs are pairs of vaccine concepts with a "Maps to" relationship, while unmapped vaccine-pairs are those without a "Maps to" relationship. We represented each vaccine concept name as a set of words, and derived term-difference pairs (i.e., name differences) for mapped and unmapped vaccine-pairs. If the same term-difference pair can be obtained by both mapped and unmapped vaccine-pairs, then this is considered as a potential mapping inconsistency. Applying this approach to the vaccine mappings in OMOP, a total of 2087 potentially mapping inconsistencies were obtained. A randomly selected 200 samples were evaluated by domain experts to identify, validate, and categorize the inconsistencies. Experts identified 95 cases revealing valid mapping issues. The remaining 105 cases were found to be invalid due to the external and/or contextual information used in the mappings that were not reflected in the concept names of vaccines. This indicates that our semi-automated approach shows promise in identifying mapping inconsistencies among vaccine concepts in the OMOP vocabulary. Rashmie Abeysinghe, Adam Black, Denys Kaduk, Christian Reich, Lixia Yao, Licong Cui |
J. Biomed. Informatics | 8 |
| 2021 | A Comparison of Exhaustive and Non-lattice-based Methods for Auditing Hierarchical Relations in Gene Ontology
Rashmie Abeysinghe, Fengbo Zheng, Licong Cui |
AMIA | 3 |
| 2021 | Identifying Sleep-Related Factors Associated with Cognitive Function in a Hispanics/Latinos Cohort: A Dual Random Forest Approach
Licong Cui, Paul E. Schulz, Guo-Qiang Zhang 0001 |
AMIA | 2 |
| 2021 | Leveraging non-lattice subgraphs for suggestion of new concepts for SNOMED CTabstractrelations in biomedical ontologies like SNOMED CT. However, little is known about non-lattice subgraphs' capability to uncover new or missing concepts in biomedical ontologies. In this work, we investigate a lexical-based intersection approach based on non-lattice subgraphs to identify potential missing concepts in SNOMED CT. We first construct lexical features of concepts using their fully specified names. Then we generate hierarchically unrelated concept pairs in non-lattice subgraphs as the candidates to derive new concepts. For each candidate pair of concepts, we conduct an order-preserving intersection based on the two concepts' lexical features, with the intersection result serving as the potential new concept name suggested. We further perform automatic validation through terminologies in the Unified Medical Language System (UMLS) and literature in PubMed. Applying this approach to the March 2021 release of SNOMED CT US Edition, we obtained 7,702 potential missing concepts, among which 1,288 were validated through UMLS and 1,309 were validated through PubMed. The results showed that non-lattice subgraphs have the potential to facilitate suggestion of new concepts for SNOMED CT. Xubing Hao, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui |
BIBM | 4 |
| 2020 | A lexical-based approach for exhaustive detection of missing hierarchical IS-A relations in SNOMED CT
Fengbo Zheng, Jay Shi, Licong Cui |
AMIA | 3 |
| 2020 | HRV-Spark: Computing Heart Rate Variability Measures Using Apache SparkabstractHeart rate variability (HRV) analysis has been serving as a significant promising marker in clinical research over the last few decades. The rapidly growing heart rate data generated from various devices, particularly the electrocardiograph (ECG), need to be stored properly and processed timely. There is a pressing need to develop efficient approaches for performing HRV analyses based on ECG signals. In this paper, we introduce a cloud computing approach (called HRV-Spark) to compute HRV measures in parallel by leveraging Apache Spark and a QRS detection algorithm in [1]. We ran HRV-Spark on Amazon Web Services (AWS) clusters using large-scale datasets in the National Sleep Research Resource. We evaluated the performance and scalability of HRV-Spark in terms of the number of computing nodes in the AWS cluster, the size of the input datasets, and the hardware configuration of the computing nodes. The results show that HRV-Spark is an efficient and scalable approach for computing HRV measures. Xufeng Qu, Jinze Liu, Licong Cui |
BIBM | 4 |
| 2020 | A Lexical-based Formal Concept Analysis Method to Identify Missing Concepts in the NCI ThesaurusabstractBiomedical terminologies have been increasingly used in modern biomedical research and applications to facilitate data management and ensure semantic interoperability. As part of the evolution process, new concepts are regularly added to biomedical terminologies in response to the evolving domain knowledge and emerging applications. Most existing concept enrichment methods suggest new concepts via directly importing knowledge from external sources. In this paper, we introduced a lexical method based on formal concept analysis (FCA) to identify potentially missing concepts in a given terminology by leveraging its intrinsic knowledge - concept names. We first construct the FCA formal context based on the lexical features of concepts. Then we perform multistage intersection to formalize new concepts and detect potentially missing concepts. We applied our method to the Disease or Disorder sub-hierarchy in the National Cancer Institute (NCI) Thesaurus (19.08d version) and identified a total of 8,983 potentially missing concepts. As a preliminary evaluation of our method to validate the potentially missing concepts, we further checked whether they were included in any external source terminology in the Unified Medical Language System (UMLS). The result showed that 592 out of 8,937 potentially missing concepts were found in the UMLS. Fengbo Zheng, Licong Cui |
BIBM | 2 |
| 2020 | SSIF: Subsumption-based Sub-term Inference Framework to audit Gene OntologyabstractMOTIVATION: The Gene Ontology (GO) is the unifying biological vocabulary for codifying, managing and sharing biological knowledge. Quality issues in GO, if not addressed, can cause misleading results or missed biological discoveries. Manual identification of potential quality issues in GO is a challenging and arduous task, given its growing size. We introduce an automated auditing approach for suggesting potentially missing is-a relations, which may further reveal erroneous is-a relations. RESULTS: We developed a Subsumption-based Sub-term Inference Framework (SSIF) by leveraging a novel term-algebra on top of a sequence-based representation of GO concepts along with three conditional rules (monotonicity, intersection and sub-concept rules). Applying SSIF to the October 3, 2018 release of GO suggested 1938 unique potentially missing is-a relations. Domain experts evaluated a random sample of 210 potentially missing is-a relations. The results showed SSIF achieved a precision of 60.61, 60.49 and 46.03% for the monotonicity, intersection and sub-concept rules, respectively. AVAILABILITY AND IMPLEMENTATION: SSIF is implemented in Java. The source code is available at https://github.com/rashmie/SSIF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
Bioinform. | 4 |
| 2020 | A transformation-based method for auditing the IS-A hierarchy of biomedical terminologies in the Unified Medical Language SystemabstractOBJECTIVE: The Unified Medical Language System (UMLS) integrates various source terminologies to support interoperability between biomedical information systems. In this article, we introduce a novel transformation-based auditing method that leverages the UMLS knowledge to systematically identify missing hierarchical IS-A relations in the source terminologies. MATERIALS AND METHODS: Given a concept name in the UMLS, we first identify its base and secondary noun chunks. For each identified noun chunk, we generate replacement candidates that are more general than the noun chunk. Then, we replace the noun chunks with their replacement candidates to generate new potential concept names that may serve as supertypes of the original concept. If a newly generated name is an existing concept name in the same source terminology with the original concept, then a potentially missing IS-A relation between the original and the new concept is identified. RESULTS: Applying our transformation-based method to English-language concept names in the UMLS (2019AB release), a total of 39 359 potentially missing IS-A relations were detected in 13 source terminologies. Domain experts evaluated a random sample of 200 potentially missing IS-A relations identified in the SNOMED CT (U.S. edition) and 100 in Gene Ontology. A total of 173 of 200 and 63 of 100 potentially missing IS-A relations were confirmed by domain experts, indicating that our method achieved a precision of 86.5% and 63% for the SNOMED CT and Gene Ontology, respectively. CONCLUSIONS: Our results showed that our transformation-based method is effective in identifying missing IS-A relations in the UMLS source terminologies. Fengbo Zheng, Jay Shi, Yuntao Yang, W. Jim Zheng, Licong Cui |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Leveraging Non-lattice Subgraphs to Audit Hierarchical Relations in NCI Thesaurus
Rashmie Abeysinghe, Michael A. Brooks, Licong Cui |
AMIA | 3 |
| 2019 | Discriminative Sleep Patterns of Alzheimer's Disease via Tensor Factorization
Yejin Kim 0001, Xiaoqian Jiang, Licong Cui |
AMIA | 5 |
| 2019 | SeizureBank: A Repository of Analysis-ready Seizure Signal Data
Yan Huang 0034, Shiqiang Tao, Licong Cui, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
AMIA | 4 |
| 2019 | Ontology of Consumer Health Vocabulary: providing a formal and interoperable semantic resource for linking lay language and medical terminologyabstractThe Consumer Health Vocabulary has been an important contribution to the health informatics field since its introduction in 2006. Many studies have utilized the vocabulary for various scientific research to bridge the gap between consumers and health experts. Given the flat file format of the Consumer Health Vocabulary dataset, we developed a SKOS-based ontology of the dataset. As an ontology, this dataset can be semantically linked to other resources to provide consumer-level meaning. In addition with this artifact, we plan to further expand the terminology. Muhammad Amith, Licong Cui, Kirk Roberts, Hua Xu 0001, Cui Tao |
BIBM | 2 |
| 2019 | Hadoop-EDF: Large-scale Distributed Processing of Electrophysiological Signal Data in Hadoop MapReduceabstractRapidly growing volume of electrophysiological signals has been generated for clinical research in neurological disorders. European Data Format (EDF) is a standard format for storing electrophysiological signals. However, the bottleneck of existing signal analysis tools for handling large-scale datasets is the sequential way of loading large EDF files before performing signal analyses. To overcome this, we develop Hadoop-EDF, a distributed signal processing tool to load EDF data in a parallel manner using Hadoop MapReduce. Hadoop-EDF uses a robust data partition algorithm making EDF data parallelly processable. We evaluate Hadoop-EDF's scalability and performance by leveraging two datasets from the National Sleep Research Resource and running experiments on Amazon Web Service clusters. The performance of Hadoop-EDF on a 20-node cluster achieved about 26 times and 47 times faster than the sequential processing of 200 small-size files and 200 large-size files, respectively. The results demonstrate that Hadoop-EDF is more suitable and effective in processing large EDF files. Jinze Liu, Licong Cui |
BIBM | 4 |
| 2019 | A Hybrid Method to Detect Missing Hierarchical Relations in NCI ThesaurusabstractBiomedical terminologies such as National Cancer Institute thesaurus (NCIt) have been widely used in supporting various biomedical research and applications. Therefore, the quality of biomedical terminologies directly impacts their downstream applications. In this paper, we introduce a hybrid method to identify missing hierarchical IS-A relations in NCIt, by leveraging both role definitions and lexical features of concepts in non-lattice subgraphs. We first extract non-lattice subgraphs in NCIt, problematic areas with quality issue. We model each concept using its role definitions and words in its concept name as well as words in the names of its ancestors. Then we perform a two-step subsumption testing for candidate pairs of concepts in the non-lattice subgraphs to automatically suggest potentially missing IS-A relations. We applied our method to the 19.01d version of NCIt. A total of 9,512 non-lattice subgraphs were extracted, among which 654 of them revealed 268 potentially missing IS-A relations. After the removal of duplication and redundancy, 121 potentially missing IS-A relations were obtained. To evaluate our method, we adopted a retrospective ground truth (RGT)-based idea to use version difference as the reference standard. We constructed a reference standard based on the IS-A changes between the 19.01d and 19.07e versions of NCIt. Among 121 potentially missing IS-A relations suggested by our method, 46 out of them are valid according to the reference standard. The RGT-based evaluation indicates that our hybrid method is promising in detecting missing IS-A relations which motivates us to perform a thorough evaluation by domain experts in future work. Fengbo Zheng, Rashmie Abeysinghe, Licong Cui |
BIBM | 3 |
| 2018 | Identifying Similar Non-Lattice Subgraphs in Gene Ontology based on Structural Isomorphism and Semantic Similarity of Concept Labels
Rashmie Abeysinghe, Xufeng Qu, Licong Cui |
AMIA | 3 |
| 2018 | A Cross-Cohort Query System for the National Sleep Research Resource (NSRR)
Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 1 |
| 2018 | A Lexical Approach to Identifying Subtype Inconsistencies in Biomedical Terminologies
Rashmie Abeysinghe, Fengbo Zheng, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
BIBM | 5 |
| 2018 | Exploring Deep Learning-based Approaches for Predicting Concept Names in SNOMED CT
Fengbo Zheng, Licong Cui |
BIBM | 2 |
| 2018 | The National Sleep Research Resource: towards a sleep data commonsabstractObjective: The gold standard for diagnosing sleep disorders is polysomnography, which generates extensive data about biophysical changes occurring during sleep. We developed the National Sleep Research Resource (NSRR), a comprehensive system for sharing sleep data. The NSRR embodies elements of a data commons aimed at accelerating research to address critical questions about the impact of sleep disorders on important health outcomes. Approach: We used a metadata-guided approach, with a set of common sleep-specific terms enforcing uniform semantic interpretation of data elements across three main components: (1) annotated datasets; (2) user interfaces for accessing data; and (3) computational tools for the analysis of polysomnography recordings. We incorporated the process for managing dataset-specific data use agreements, evidence of Institutional Review Board review, and the corresponding access control in the NSRR web portal. The metadata-guided approach facilitates structural and semantic interoperability, ultimately leading to enhanced data reusability and scientific rigor. Results: The authors curated and deposited retrospective data from 10 large, NIH-funded sleep cohort studies, including several from the Trans-Omics for Precision Medicine (TOPMed) program, into the NSRR. The NSRR currently contains data on 26 808 subjects and 31 166 signal files in European Data Format. Launched in April 2014, over 3000 registered users have downloaded over 130 terabytes of data. Conclusions: The NSRR offers a use case and an example for creating a full-fledged data commons. It provides a single point of access to analysis-ready physiological signals from polysomnography obtained from multiple sources, and a wide variety of clinical data to facilitate sleep research. Guo-Qiang Zhang 0001, Licong Cui, Remo Mueller, Shiqiang Tao, Matthew Kim, Michael Rueschman, Sara Mariani, Daniel R. Mobley, Susan Redline |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | An efficient, large-scale, non-lattice-detection algorithm for exhaustive structural auditing of biomedical ontologies
Guo-Qiang Zhang 0001, Guangming Xing, Licong Cui |
J. Biomed. Informatics | 3 |
| 2018 | Auditing SNOMED CT hierarchical relations based on lexical features of concepts in non-lattice subgraphs
Licong Cui, Olivier Bodenreider, Jay Shi, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 1 |
| 2018 | Quality assurance of biomedical terminologies and ontologies
James Geller, Yehoshua Perl, Licong Cui, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 3 |
| 2018 | HyCLASSS: A Hybrid Classifier for Automatic Sleep Stage ScoringabstractAutomatic identification of sleep stage is an important step in a sleep study. In this paper, we propose a hybrid automatic sleep stage scoring approach, named HyCLASSS, based on single channel electroencephalogram (EEG). HyCLASSS, for the first time, leverages both signal and stage transition features of human sleep for automatic identification of sleep stages. HyCLASSS consists of two parts: A random forest classifier and correction rules. Random forest classifier is trained using 30 EEG signal features, including temporal, frequency, and nonlinear features. The correction rules are constructed based on stage transition feature, importing the continuity property of sleep, and characteristic of sleep stage transition. Compared with the gold standard of manual scoring using Rechtschaffen and Kales criterion, the overall accuracy and kappa coefficient applied on 198 subjects has reached 85.95% and 0.8046 in our experiment, respectively. The performance of HyCLASS compared favorably to previous work, and it could be integrated with sleep evaluation or sleep diagnosis system in the future. Licong Cui, Shiqiang Tao, Jing Chen 0002, Xiang Zhang 0001, Guo-Qiang Zhang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Can SNOMED CT Changes Be Used as a Surrogate Standard for Evaluating the Performance of Its Auditing Methods?
Guo-Qiang Zhang 0001, Yan Huang 0034, Licong Cui |
AMIA | 3 |
| 2017 | Quality Assurance of NCI Thesaurus by Mining Structural-Lexical Patterns
Rashmie Abeysinghe, Michael A. Brooks, Jeffery C. Talbert, Licong Cui |
AMIA | 4 |
| 2017 | SpindleSphere: A Web-based Platform for Large-scale Sleep Spindle Analysis and Visualization
Licong Cui, Shiqiang Tao, Ningzhou Zeng, Guo-Qiang Zhang 0001 |
AMIA | 2 |
| 2017 | Facilitating Cohort Discovery by Enhancing Ontology Exploration, Query Management and Query Sharing for Large Clinical Data Repositories
Shiqiang Tao, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 2 |
| 2017 | Spark-MCA: Large-scale, Exhaustive Formal Concept Analysis for Evaluating the Semantic Completeness of SNOMED CT
Wei Zhu 0010, Guo-Qiang Zhang 0001, Licong Cui |
AMIA | 3 |
| 2017 | Query-constraint-based association rule mining from diverse clinical datasets in the national sleep research resourceabstractSecondary use of biomedical data has gained much attention recently to facilitate rapid knowledge discovery in biomedicine. Association Rule Mining (ARM) has been a popular technique for biomedical researchers to perform exploratory data analysis and discover potential relationships among variables in biomedical datasets. However, ARM of a high-dimensional biomedical dataset may produce a large number of rules that may not be interesting. In this paper, we introduce a query-constraint-based ARM (QARM) approach for exploratory analysis of diverse clinical datasets integrated in the National Sleep Research Resource (NSRR), which enables the rule mining on a subset of data containing items of interest based on a query constraint. In addition, biomedical datasets always contain semantically similar variables, thus we performed similar-variable-merging so that rules with simlar variables are not obtained. Applying QARM on five datasets from NSRR obtained a total of 6,921 rules with a minimum confidence of 60% (using top 50 rules for each query constraint). Rashmie Abeysinghe, Licong Cui |
BIBM | 2 |
| 2017 | Auditing subtype inconsistencies among gene ontology conceptsabstractGene Ontology (GO) provides a controlled vocabulary for describing genes and related gene products. Quality assurance of Gene ontology (GO) is a vital aspect of the terminology management lifecycle. In this paper, we introduce a lexical-based inference approach to detecting subtype (or isa) inconsistencies among GO terms (i.e., biological concepts). We first model the name of each concept as a set of words. Then, we generate hierarchically linked and unlinked pairs of concepts (A, B), where A and B have the same number of words, and contain common words as well as a single different word. Each linked concept-pair infers a linked term-pair, and each unlinked concept-pair infers an unlinked term-pair. A term-pair appearing as both linked and unlinked is considered a potential inconsistency, which may represent a subtype inconsistency between the original linked and unlinked concept-pair. Applying this approach to the 03/28/2017 release of GO, a total of 3,715 potential subtype inconsistencies were obtained. Evaluation of a random sample of potential inconsistencies revealed two types of potential errors: missing subtype relations and incorrect subtype relations in GO, and achieved an accuracy of 56.33% for detecting such errors. This indicates that this lexical-based inference approach using the set-of-words model is a promising way to facilitate quality improvement of GO. Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
BIBM | 4 |
| 2017 | Evaluation of relational and NoSQL approaches for patient cohort identification from heterogeneous data sourcesabstractPatient cohort discovery across heterogeneous data sources is a challenging task, which may involve a complicated process of data loading, harmonization, and querying. Most existing cohort identification tools use a relational database model implemented in SQL for storing patient data. However, SQL databases have restrictions on the maximum number of columns in a table, which necessitates the breaking down of high-dimensional data into multiple tables and affects query performance as a result. In this paper, we proposed two NoSQL-based patient cohort query systems based on an existing SQL-based system for a cross-cohort query interface for the National Sleep Resource Research (NSRR). We used eight NSRR datasets in our experiment to evaluate the performance of NoSQL-based and SQL-based systems in data loading, harmonization, and query. Our experiment showed that NoSQL-based approaches outperformed the SQL-based, and NoSQL-based systems are rather promising for developing patient cohort query systems across heterogeneous data sources. Ningzhou Zeng, Guo-Qiang Zhang 0001, Licong Cui |
BIBM | 4 |
| 2017 | Mining non-lattice subgraphs for detecting missing hierarchical relations and concepts in SNOMED CTabstractOBJECTIVE: Quality assurance of large ontological systems such as SNOMED CT is an indispensable part of the terminology management lifecycle. We introduce a hybrid structural-lexical method for scalable and systematic discovery of missing hierarchical relations and concepts in SNOMED CT. MATERIAL AND METHODS: All non-lattice subgraphs (the structural part) in SNOMED CT are exhaustively extracted using a scalable MapReduce algorithm. Four lexical patterns (the lexical part) are identified among the extracted non-lattice subgraphs. Non-lattice subgraphs exhibiting such lexical patterns are often indicative of missing hierarchical relations or concepts. Each lexical pattern is associated with a potential specific type of error. RESULTS: Applying the structural-lexical method to SNOMED CT (September 2015 US edition), we found 6801 non-lattice subgraphs that matched these lexical patterns, of which 2046 were amenable to visual inspection. We evaluated a random sample of 100 small subgraphs, of which 59 were reviewed in detail by domain experts. All the subgraphs reviewed contained errors confirmed by the experts. The most frequent type of error was missing is-a relations due to incomplete or inconsistent modeling of the concepts. CONCLUSIONS: Our hybrid structural-lexical method is innovative and proved effective not only in detecting errors in SNOMED CT, but also in suggesting remediation for these errors. Licong Cui, Wei Zhu 0010, Shiqiang Tao, James T. Case, Olivier Bodenreider, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | ODaCCI: Ontology-guided Data Curation for Multisite Clinical Research Data Integration in the NINDS Center for SUDEP Research
Licong Cui, Yan Huang 0034, Shiqiang Tao, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
AMIA | 1 |
| 2016 | DCDS: A Real-time Data Capture and Personalized Decision Support System for Heart Failure Patients in Skilled Nursing Facilities
Wei Zhu 0010, Lingyun Luo, Tarun Jain, Rebecca S. Boxer, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 5 |
| 2016 | Biomedical Ontology Quality Assurance Using a Big Data ApproachabstractThis article presents recent progresses made in using scalable cloud computing environment, Hadoop and MapReduce, to perform ontology quality assurance (OQA), and points to areas of future opportunity. The standard sequential approach used for implementing OQA methods can take weeks if not months for exhaustive analyses for large biomedical ontological systems. With OQA methods newly implemented using massively parallel algorithms in the MapReduce framework, several orders of magnitude in speed-up can be achieved (e.g., from three months to three hours). Such dramatically reduced time makes it feasible not only to perform exhaustive structural analysis of large ontological hierarchies, but also to systematically track structural changes between versions for evolutional analysis. As an exemplar, progress is reported in using MapReduce to perform evolutional analysis and visualization on the Systemized Nomenclature of Medicine—Clinical Terms (SNOMED CT), a prominent clinical terminology system. Future opportunities in three areas are described: one is to extend the scope of MapReduce-based approach to existing OQA methods, especially for automated exhaustive structural analysis. The second is to apply our proposed MapReduce Pipeline for Lattice-based Evaluation (MaPLE) approach, demonstrated as an exemplar method for SNOMED CT, to other biomedical ontologies. The third area is to develop interfaces for reviewing results obtained by OQA methods and for visualizing ontological alignment and evolution, which can also take advantage of cloud computing technology to systematically pre-compute computationally intensive jobs in order to increase performance during user interactions with the visualization interface. Advances in these directions are expected to better support the ontological engineering lifecycle. Licong Cui, Shiqiang Tao, Guo-Qiang Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | COHeRE: Cross-Ontology Hierarchical Relation Examination for Ontology Quality Assurance
Licong Cui |
AMIA | 1 |
| 2015 | COBE: A Conjunctive Ontology Browser and Explorer for Visualizing SNOMED CT Fragments
Wei Zhu 0010, Shiqiang Tao, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 4 |
| 2014 | A Semantic-based Approach for Exploring Consumer Health Questions Using UMLS
Licong Cui, Shiqiang Tao, Guo-Qiang Zhang 0001 |
AMIA | 1 |
| 2014 | MEDCIS: Multi-Modality Epilepsy Data Capture and Integration System
Guo-Qiang Zhang 0001, Licong Cui, Samden D. Lhatoo, Satya Sanket Sahoo |
AMIA | 2 |
| 2014 | MaPLE: A MapReduce Pipeline for Lattice-based Evaluation and its application to SNOMED CTabstractNon-lattice fragments are often indicative of structural anomalies in ontological systems and, as such, represent possible areas of focus for subsequent quality assurance work. However, extracting the non-lattice fragments in large ontological systems is computationally expensive if not prohibitive, using a traditional sequential approach. In this paper we present a general MapReduce pipeline, called MaPLE (MapReduce Pipeline for Lattice-based Evaluation), for extracting non-lattice fragments in large partially ordered sets and demonstrate its applicability in ontology quality assurance. Using MaPLE in a 30-node Hadoop local cloud, we systematically extracted non-lattice fragments in 8 SNOMED CT versions from 2009 to 2014 (each containing over 300k concepts), with an average total computing time of less than 3 hours per version. With dramatically reduced time, MaPLE makes it feasible not only to perform exhaustive structural analysis of large ontological hierarchies, but also to systematically track structural changes between versions. Our change analysis showed that the average change rates on the non-lattice pairs are up to 38.6 times higher than the change rates of the background structure (concept nodes). This demonstrates that fragments around non-lattice pairs exhibit significantly higher rates of change in the process of ontological evolution. Guo-Qiang Zhang 0001, Wei Zhu 0010, Shiqiang Tao, Olivier Bodenreider, Licong Cui |
IEEE BigData | 6 |
| 2014 | Epilepsy and seizure ontology: towards an epilepsy informatics infrastructure for clinical research and patient careabstractOBJECTIVE: Epilepsy encompasses an extensive array of clinical and research subdomains, many of which emphasize multi-modal physiological measurements such as electroencephalography and neuroimaging. The integration of structured, unstructured, and signal data into a coherent structure for patient care as well as clinical research requires an effective informatics infrastructure that is underpinned by a formal domain ontology. METHODS: We have developed an epilepsy and seizure ontology (EpSO) using a four-dimensional epilepsy classification system that integrates the latest International League Against Epilepsy terminology recommendations and National Institute of Neurological Disorders and Stroke (NINDS) common data elements. It imports concepts from existing ontologies, including the Neural ElectroMagnetic Ontologies, and uses formal concept analysis to create a taxonomy of epilepsy syndromes based on their seizure semiology and anatomical location. RESULTS: EpSO is used in a suite of informatics tools for (a) patient data entry, (b) epilepsy focused clinical free text processing, and (c) patient cohort identification as part of the multi-center NINDS-funded study on sudden unexpected death in epilepsy. EpSO is available for download at http://prism.case.edu/prism/index.php/EpilepsyOntology. DISCUSSION: An epilepsy ontology consortium is being created for community-driven extension, review, and adoption of EpSO. We are in the process of submitting EpSO to the BioPortal repository. CONCLUSIONS: EpSO plays a critical role in informatics tools for epilepsy patient care and multi-center clinical research. Satya Sanket Sahoo, Samden D. Lhatoo, Licong Cui, Catherine P. Jayapandian, Alireza Bozorgi, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Complex epilepsy phenotype extraction from narrative clinical discharge summaries
Licong Cui, Satya Sanket Sahoo, Samden D. Lhatoo, Prashant Rai, Alireza Bozorgi, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 1 |
| 2012 | EpiDEA: Extracting Structured Epilepsy and Seizure Information from Patient Discharge Summaries for Cohort Identification
Licong Cui, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo, Alireza Bozorgi |
AMIA | 1 |
| 2010 | A set coverage problem
Guo-Qiang Zhang 0001, Licong Cui |
Inf. Process. Lett. | 2 |
| 2009 | Intuitionistic Fuzzy Linguistic Quantifiers Based on Intuitionistic Fuzzy-Valued Fuzzy Measures and integralsabstractIn this paper, we generalize Ying's model of linguistic quantifiers [M.S. Ying, Linguistic quantifiers modeled by Sugeno integrals, Artificial Intelligence, 170 (2006) 581-606] to intuitionistic linguistic quantifiers. An intuitionistic linguistic quantifier is represented by a family of intuitionistic fuzzy-valued fuzzy measures and the intuitionistic truth value (the degrees of satisfaction and non-satisfaction) of a quantified proposition is calculated by using intuitionistic fuzzy-valued fuzzy integral. Description of a quantifier by intuitionistic fuzzy-valued fuzzy measures allows us to take into account differences in understanding the meaning of the quantifier by different persons. If the intuitionistic fuzzy linguistic quantifiers are taken to be linguistic fuzzy quantifiers, then our model reduces to Ying's model. Some excellent logical properties of intuitionistic linguistic quantifiers are obtained including a prenex norm form theorem. A simple example is presented to illustrate the use of intuitionistic linguistic quantifiers. Licong Cui, Yongming Li 0001, Xiaohong Zhang 0001 |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |
| 2008 | Linguistic quantifiers based on Choquet integrals
Licong Cui, Yongming Li 0001 |
Int. J. Approx. Reason. | 1 |