EDBT 2026 Demo / reviewers in the wild / expert
Rashmie Abeysinghe
dblp:211/4173
· DBLP profile ↗
19ranked-venue papers
11as first author
11since 2021 · last 2025
0000-0001-9297-9540ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporal Ensemble Logic for Integrative Representation of the Entirety of Clinical Trials
Yan Huang 0034, Rashmie Abeysinghe, Zenan Sun, Pengze Li, Xing He 0003, Shiqiang Tao, Cui Tao, Jiang Bian 0001, Licong Cui, Guo-Qiang Zhang 0001 |
TIME | 3 |
| 2025 | Quantitatively assessing the impact of the quality of SNOMED CT subtype hierarchy on cohort queriesabstractOBJECTIVE: SNOMED CT provides a standardized terminology for clinical concepts, allowing cohort queries over heterogeneous clinical data including Electronic Health Records (EHRs). While it is intuitive that missing and inaccurate subtype (or is-a) relations in SNOMED CT reduce the recall and precision of cohort queries, the extent of these impacts has not been formally assessed. This study fills this gap by developing quantitative metrics to measure these impacts and performing statistical analysis on their significance. MATERIAL AND METHODS: We used the Optum de-identified COVID-19 Electronic Health Record dataset. We defined micro-averaged and macro-averaged recall and precision metrics to assess the impact of missing and inaccurate is-a relations on cohort queries. Both practical and simulated analyses were performed. Practical analyses involved 407 missing and 48 inaccurate is-a relations confirmed by domain experts, with statistical testing using Wilcoxon signed-rank tests. Simulated analyses used two random sets of 400 is-a relations to simulate missing and inaccurate is-a relations. RESULTS: Wilcoxon signed-rank tests from both practical and simulated analyses (P-values < .001) showed that missing is-a relations significantly reduced the micro- and macro-averaged recall, and inaccurate is-a relations significantly reduced the micro- and macro-averaged precision. DISCUSSION: The introduced impact metrics can assist SNOMED CT maintainers in prioritizing critical hierarchical defects for quality enhancement. These metrics are generally applicable for assessing the quality impact of a terminology's subtype hierarchy on its cohort query applications. CONCLUSION: Our results indicate a significant impact of missing and inaccurate is-a relations in SNOMED CT on the recall and precision of cohort queries. Our work highlights the importance of high-quality terminology hierarchy for cohort queries over EHR data and provides valuable insights for prioritizing quality improvements of SNOMED CT's hierarchy. Xubing Hao, Yan Huang 0034, Jay Shi, Rashmie Abeysinghe, Cui Tao, Kirk Roberts, Guo-Qiang Zhang 0001, Licong Cui |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | Exploring Pre-trained Language Models for Vocabulary Alignment in the UMLS
Xubing Hao, Rashmie Abeysinghe, Jay Shi, Licong Cui |
AIME (1) | 2 |
| 2023 | A deep learning approach to identify missing is-a relations in SNOMED CTabstractOBJECTIVE: SNOMED CT is the largest clinical terminology worldwide. Quality assurance of SNOMED CT is of utmost importance to ensure that it provides accurate domain knowledge to various SNOMED CT-based applications. In this work, we introduce a deep learning-based approach to uncover missing is-a relations in SNOMED CT. MATERIALS AND METHODS: Our focus is to identify missing is-a relations between concept-pairs exhibiting a containment pattern (ie, the set of words of one concept being a proper subset of that of the other concept). We use hierarchically related containment concept-pairs as positive instances and hierarchically unrelated containment concept-pairs as negative instances to train a model predicting whether an is-a relation exists between 2 concepts with containment pattern. The model is a binary classifier leveraging concept name features, hierarchical features, enriched lexical attribute features, and logical definition features. We introduce a cross-validation inspired approach to identify missing is-a relations among all hierarchically unrelated containment concept-pairs. RESULTS: We trained and applied our model on the Clinical finding subhierarchy of SNOMED CT (September 2019 US edition). Our model (based on the validation sets) achieved a precision of 0.8164, recall of 0.8397, and F1 score of 0.8279. Applying the model to predict actual missing is-a relations, we obtained a total of 1661 potential candidates. Domain experts performed evaluation on randomly selected 230 samples and verified that 192 (83.48%) are valid. CONCLUSIONS: The results showed that our deep learning approach is effective in uncovering missing is-a relations between containment concept-pairs in SNOMED CT. Rashmie Abeysinghe, Fengbo Zheng, Elmer V. Bernstam, Jay Shi, Olivier Bodenreider, Licong Cui |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Automated Identification of Missing IS-A Relations in the Human Phenotype Ontology
Maryamsadat Mohtashamian, Rashmie Abeysinghe, Xubing Hao, Hua Xu 0001, Licong Cui |
AMIA | 3 |
| 2022 | A substring replacement approach for identifying missing IS-A relations in SNOMED CTabstractBiomedical ontologies provide formalized information and knowledge in the biomedical domain. Over the years, biomedical ontologies have played an important role in facilitating biomedical research and applications. Common quality issues of biomedical ontologies include inconsistent naming of concepts, redundant concepts, redundant relations, incomplete/incorrect concept definitions, and incomplete/incorrect class hierarchies. In this work, we focus on addressing the incompleteness of the class hierarchy in SNOMED CT. We develop a substring replacement approach, leveraging concepts' lexical features and existing IS-A relations to identify potential missing IS-A relations in SNOMED CT. To evaluate the effectiveness of our approach, we performed both automated and manual validation. For the automated evaluation, we leverage relations from external terminologies in the Unified Medical Language System (UMLS) to validate the identified missing IS-A relations. For the manual validation, a randomly selected 100 samples from the results are reviewed by a domain expert. Applying our approach to the March 2022 release of SNOMED CT US Edition, we identified 3,228 potential missing IS-A relations, among which 63 were validated through the UMLS. The evaluation by the domain expert revealed that 89 out of 100 (a precision of 89%) missing IS-A relations are valid cases, showing the effectiveness of this substring replacement approach to facilitate the quality assurance of IS-A relations in SNOMED CT. Xubing Hao, Rashmie Abeysinghe, Jay Shi, Licong Cui |
BIBM | 2 |
| 2022 | Identifying Missing IS-A Relations in Orphanet Rare Disease OntologyabstractThe Orphanet Rare Disease Ontology (ORDO) provides a structured vocabulary encapsulating rare diseases. Downstream applications of ORDO depend on its accuracy to effectively perform their tasks. In this paper, we implement an automated quality assurance pipeline to identify missing is-a relations in ORDO. We first obtain lexical features from concept names. Then we generate related and unrelated feature sharing concept-pairs, where a feature sharing concept-pair can further generate derived term-pairs. If an unrelated and related feature sharing concept-pair generate the same derived term-pair, then we suggest a potential missing is-a relation between the unrelated feature sharing concept-pair. Applying this approach on the 202206-27 release of ORDO, we obtained 705 potential missing is-a relations. Leveraging external ontological information in the Unified Medical Language System, we validated 164 missing is-a relations. This indicates that our approach is a promising way to audit is-a relations in ORDO, even though further domain expert evaluation is still needed to validate the remaining potential missing is-a relations identified. Maryamsadat Mohtashamian, Rashmie Abeysinghe, Xubing Hao, Licong Cui |
BIBM | 2 |
| 2022 | An evidence-based lexical pattern approach for quality assurance of Gene Ontology relationsabstractGene Ontology (GO) is widely used in the biological domain. It is the most comprehensive ontology providing formal representation of gene functions (GO concepts) and relations between them. However, unintentional quality defects (e.g. missing or erroneous relations) in GO may exist due to the large size of GO concepts and complexity of GO structures. Such quality defects would impact the results of GO-based analyses and applications. In this work, we introduce a novel evidence-based lexical pattern approach for quality assurance of GO relations. We leverage two layers of evidence to suggest potentially missing relations in GO as follows. We first utilize related concept pairs (i.e. existing relations) in GO to extract relationship-specific lexical patterns, which serve as the first layer evidence to automatically suggest potentially missing relations between unrelated concept pairs. For each suggested missing relation, we further identify two other existing relations as the second layer of evidence that resemble the difference between the missing relation and the existing relation based on which the missing relation is suggested. Applied to the 15 December 2021 release of GO, this approach suggested a total of 866 potentially missing relations. Local domain experts evaluated the entire set of potentially missing relations, and identified 821 as missing relations and 45 indicate erroneous existing relations. We submitted these findings to the GO consortium for further validation and received encouraging feedback. These indicate that our evidence-based approach can be utilized to uncover missing relations and erroneous existing relations in GO. Rashmie Abeysinghe, Yuntao Yang, Mason Bartels, W. Jim Zheng, Licong Cui |
Briefings Bioinform. | 1 |
| 2022 | Towards quality improvement of vaccine concept mappings in the OMOP vocabulary with a semi-automated methodabstractThe Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) provides a unified model to integrate disparate real-world data (RWD) sources. An integral part of the OMOP CDM is the Standardized Vocabularies (henceforth referred to as the OMOP vocabulary), which enables organization and standardization of medical concepts across various clinical domains of the OMOP CDM. For concepts with the same meaning from different source vocabularies, one is designated as the standard concept, while the others are specified as non-standard or source concepts and mapped to the standard one. However, due to the heterogeneity of source vocabularies, there may exist mapping issues such as erroneous mappings and missing mappings in the OMOP vocabulary, which could affect the results of downstream analyses with RWD. In this paper, we focus on quality assurance of vaccine concept mappings in the OMOP vocabulary, which is necessary to accurately harness the power of RWD on vaccines. We introduce a semi-automated lexical approach to audit vaccine mappings in the OMOP vocabulary. We generated two types of vaccine-pairs: mapped and unmapped, where mapped vaccine-pairs are pairs of vaccine concepts with a "Maps to" relationship, while unmapped vaccine-pairs are those without a "Maps to" relationship. We represented each vaccine concept name as a set of words, and derived term-difference pairs (i.e., name differences) for mapped and unmapped vaccine-pairs. If the same term-difference pair can be obtained by both mapped and unmapped vaccine-pairs, then this is considered as a potential mapping inconsistency. Applying this approach to the vaccine mappings in OMOP, a total of 2087 potentially mapping inconsistencies were obtained. A randomly selected 200 samples were evaluated by domain experts to identify, validate, and categorize the inconsistencies. Experts identified 95 cases revealing valid mapping issues. The remaining 105 cases were found to be invalid due to the external and/or contextual information used in the mappings that were not reflected in the concept names of vaccines. This indicates that our semi-automated approach shows promise in identifying mapping inconsistencies among vaccine concepts in the OMOP vocabulary. Rashmie Abeysinghe, Adam Black, Denys Kaduk, Christian Reich, Lixia Yao, Licong Cui |
J. Biomed. Informatics | 1 |
| 2021 | A Comparison of Exhaustive and Non-lattice-based Methods for Auditing Hierarchical Relations in Gene Ontology
Rashmie Abeysinghe, Fengbo Zheng, Licong Cui |
AMIA | 1 |
| 2021 | Leveraging non-lattice subgraphs for suggestion of new concepts for SNOMED CTabstractrelations in biomedical ontologies like SNOMED CT. However, little is known about non-lattice subgraphs' capability to uncover new or missing concepts in biomedical ontologies. In this work, we investigate a lexical-based intersection approach based on non-lattice subgraphs to identify potential missing concepts in SNOMED CT. We first construct lexical features of concepts using their fully specified names. Then we generate hierarchically unrelated concept pairs in non-lattice subgraphs as the candidates to derive new concepts. For each candidate pair of concepts, we conduct an order-preserving intersection based on the two concepts' lexical features, with the intersection result serving as the potential new concept name suggested. We further perform automatic validation through terminologies in the Unified Medical Language System (UMLS) and literature in PubMed. Applying this approach to the March 2021 release of SNOMED CT US Edition, we obtained 7,702 potential missing concepts, among which 1,288 were validated through UMLS and 1,309 were validated through PubMed. The results showed that non-lattice subgraphs have the potential to facilitate suggestion of new concepts for SNOMED CT. Xubing Hao, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui |
BIBM | 2 |
| 2020 | SSIF: Subsumption-based Sub-term Inference Framework to audit Gene OntologyabstractMOTIVATION: The Gene Ontology (GO) is the unifying biological vocabulary for codifying, managing and sharing biological knowledge. Quality issues in GO, if not addressed, can cause misleading results or missed biological discoveries. Manual identification of potential quality issues in GO is a challenging and arduous task, given its growing size. We introduce an automated auditing approach for suggesting potentially missing is-a relations, which may further reveal erroneous is-a relations. RESULTS: We developed a Subsumption-based Sub-term Inference Framework (SSIF) by leveraging a novel term-algebra on top of a sequence-based representation of GO concepts along with three conditional rules (monotonicity, intersection and sub-concept rules). Applying SSIF to the October 3, 2018 release of GO suggested 1938 unique potentially missing is-a relations. Domain experts evaluated a random sample of 210 potentially missing is-a relations. The results showed SSIF achieved a precision of 60.61, 60.49 and 46.03% for the monotonicity, intersection and sub-concept rules, respectively. AVAILABILITY AND IMPLEMENTATION: SSIF is implemented in Java. The source code is available at https://github.com/rashmie/SSIF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
Bioinform. | 1 |
| 2019 | Leveraging Non-lattice Subgraphs to Audit Hierarchical Relations in NCI Thesaurus
Rashmie Abeysinghe, Michael A. Brooks, Licong Cui |
AMIA | 1 |
| 2019 | A Hybrid Method to Detect Missing Hierarchical Relations in NCI ThesaurusabstractBiomedical terminologies such as National Cancer Institute thesaurus (NCIt) have been widely used in supporting various biomedical research and applications. Therefore, the quality of biomedical terminologies directly impacts their downstream applications. In this paper, we introduce a hybrid method to identify missing hierarchical IS-A relations in NCIt, by leveraging both role definitions and lexical features of concepts in non-lattice subgraphs. We first extract non-lattice subgraphs in NCIt, problematic areas with quality issue. We model each concept using its role definitions and words in its concept name as well as words in the names of its ancestors. Then we perform a two-step subsumption testing for candidate pairs of concepts in the non-lattice subgraphs to automatically suggest potentially missing IS-A relations. We applied our method to the 19.01d version of NCIt. A total of 9,512 non-lattice subgraphs were extracted, among which 654 of them revealed 268 potentially missing IS-A relations. After the removal of duplication and redundancy, 121 potentially missing IS-A relations were obtained. To evaluate our method, we adopted a retrospective ground truth (RGT)-based idea to use version difference as the reference standard. We constructed a reference standard based on the IS-A changes between the 19.01d and 19.07e versions of NCIt. Among 121 potentially missing IS-A relations suggested by our method, 46 out of them are valid according to the reference standard. The RGT-based evaluation indicates that our hybrid method is promising in detecting missing IS-A relations which motivates us to perform a thorough evaluation by domain experts in future work. Fengbo Zheng, Rashmie Abeysinghe, Licong Cui |
BIBM | 2 |
| 2018 | Identifying Similar Non-Lattice Subgraphs in Gene Ontology based on Structural Isomorphism and Semantic Similarity of Concept Labels
Rashmie Abeysinghe, Xufeng Qu, Licong Cui |
AMIA | 1 |
| 2018 | A Lexical Approach to Identifying Subtype Inconsistencies in Biomedical Terminologies
Rashmie Abeysinghe, Fengbo Zheng, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
BIBM | 1 |
| 2017 | Quality Assurance of NCI Thesaurus by Mining Structural-Lexical Patterns
Rashmie Abeysinghe, Michael A. Brooks, Jeffery C. Talbert, Licong Cui |
AMIA | 1 |
| 2017 | Query-constraint-based association rule mining from diverse clinical datasets in the national sleep research resourceabstractSecondary use of biomedical data has gained much attention recently to facilitate rapid knowledge discovery in biomedicine. Association Rule Mining (ARM) has been a popular technique for biomedical researchers to perform exploratory data analysis and discover potential relationships among variables in biomedical datasets. However, ARM of a high-dimensional biomedical dataset may produce a large number of rules that may not be interesting. In this paper, we introduce a query-constraint-based ARM (QARM) approach for exploratory analysis of diverse clinical datasets integrated in the National Sleep Research Resource (NSRR), which enables the rule mining on a subset of data containing items of interest based on a query constraint. In addition, biomedical datasets always contain semantically similar variables, thus we performed similar-variable-merging so that rules with simlar variables are not obtained. Applying QARM on five datasets from NSRR obtained a total of 6,921 rules with a minimum confidence of 60% (using top 50 rules for each query constraint). Rashmie Abeysinghe, Licong Cui |
BIBM | 1 |
| 2017 | Auditing subtype inconsistencies among gene ontology conceptsabstractGene Ontology (GO) provides a controlled vocabulary for describing genes and related gene products. Quality assurance of Gene ontology (GO) is a vital aspect of the terminology management lifecycle. In this paper, we introduce a lexical-based inference approach to detecting subtype (or isa) inconsistencies among GO terms (i.e., biological concepts). We first model the name of each concept as a set of words. Then, we generate hierarchically linked and unlinked pairs of concepts (A, B), where A and B have the same number of words, and contain common words as well as a single different word. Each linked concept-pair infers a linked term-pair, and each unlinked concept-pair infers an unlinked term-pair. A term-pair appearing as both linked and unlinked is considered a potential inconsistency, which may represent a subtype inconsistency between the original linked and unlinked concept-pair. Applying this approach to the 03/28/2017 release of GO, a total of 3,715 potential subtype inconsistencies were obtained. Evaluation of a random sample of potential inconsistencies revealed two types of potential errors: missing subtype relations and incorrect subtype relations in GO, and achieved an accuracy of 56.33% for detecting such errors. This indicates that this lexical-based inference approach using the set-of-words model is a promising way to facilitate quality improvement of GO. Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui |
BIBM | 1 |