Hang Dong 0002

dblp:135/8614-2 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
12since 2021 · last 2024
0000-0001-6828-6891ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Ontology Text Alignment: Aligning Textual Content to Terminological Axioms
abstract
Despite the impressive advancements in Large Language Models (LLMs), their ability to perform reasoning and provide explainable outcomes remains a challenge, underscoring the continued relevance of ontologies in certain areas, particularly due to the reasoning and validation capabilities of ontologies. Ontology modelling and semantic search, due to their inherent complexity, still demand considerable human effort and expertise. Addressing this gap, our paper introduces the problem of ontology text alignment, which involves finding the most relevant axioms with respect to the given reference text. We propose an advanced Retrieval Augmented Generation framework that leverages BERT models and generative LLMs, together with ontology semantic enhancement based on atomic decomposition. Additionally, we have developed benchmarks in geology and biomedical areas. Our evaluation demonstrates the positive impact of our framework.
Jieying Chen 0001, Hang Dong 0002, Jiaoyan Chen 0001, Ian Horrocks 0001
ECAI2
2024 A Language Model Based Framework for New Concept Placement in Ontologies
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yongsheng Gao 0005, Ian Horrocks 0001
ESWC (1)1
2024 Taxonomy Completion via Implicit Concept Insertion
abstract
\beginabstract High quality taxonomies play a critical role in various domains such as e-commerce, web search and ontology engineering. While there has been extensive work on expanding taxonomies from externally mined data, there has been less attention paid to enriching taxonomies by exploiting existing concepts and structure within the taxonomy. In this work, we show the usefulness of this kind of enrichment, and explore its viability with a new taxonomy completion system ICON (I mplicit CON cept Insertion). ICON generates new concepts by identifying implicit concepts based on the existing concept structure, generating names for such concepts and inserting them in appropriate positions within the taxonomy. ICON integrates techniques from entity retrieval, text summary, and subsumption prediction; this modular architecture offers high flexibility while achieving state-of-the-art performance. We have evaluated ICON on two e-commerce taxonomies, and the results show that it offers significant advantages over strong baselines including recent taxonomy completion models and the large language model, ChatGPT.
Jingchuan Shi, Hang Dong 0002, Jiaoyan Chen 0001, Ian Horrocks 0001
WWW2
2024 Can GPT-3.5 generate and code discharge summaries?
abstract
OBJECTIVES: The aim of this study was to investigate GPT-3.5 in generating and coding medical documents with International Classification of Diseases (ICD)-10 codes for data augmentation on low-resource labels. MATERIALS AND METHODS: Employing GPT-3.5 we generated and coded 9606 discharge summaries based on lists of ICD-10 code descriptions of patients with infrequent (or generation) codes within the MIMIC-IV dataset. Combined with the baseline training set, this formed an augmented training set. Neural coding models were trained on baseline and augmented data and evaluated on an MIMIC-IV test set. We report micro- and macro-F1 scores on the full codeset, generation codes, and their families. Weak Hierarchical Confusion Matrices determined within-family and outside-of-family coding errors in the latter codesets. The coding performance of GPT-3.5 was evaluated on prompt-guided self-generated data and real MIMIC-IV data. Clinicians evaluated the clinical acceptability of the generated documents. RESULTS: Data augmentation results in slightly lower overall model performance but improves performance for the generation candidate codes and their families, including 1 absent from the baseline training data. Augmented models display lower out-of-family error rates. GPT-3.5 identifies ICD-10 codes by their prompted descriptions but underperforms on real data. Evaluators highlight the correctness of generated concepts while suffering in variety, supporting information, and narrative. DISCUSSION AND CONCLUSION: While GPT-3.5 alone given our prompt setting is unsuitable for ICD-10 coding, it supports data augmentation for training neural models. Augmentation positively affects generation code families but mainly benefits codes with existing examples. Augmentation reduces out-of-family errors. Documents generated by GPT-3.5 state prompted concepts correctly but lack variety, and authenticity in narratives.
Matús Falis, Aryo Pradipta Gema, Hang Dong 0002, Luke Daines, Siddharth Basetti, Michael Holder, Rose S. Penfold, Alexandra Birch, Beatrice Alex
J. Am. Medical Informatics Assoc.3
2023 Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement
abstract
Mentions of new concepts appear regularly in texts and require automated approaches to harvest and place them into Knowledge Bases (KB), e.g., ontologies and taxonomies. Existing datasets suffer from three issues, (i) mostly assuming that a new concept is pre-discovered and cannot support out-of-KB mention discovery; (ii) only using the concept label as the input along with the KB and thus lacking the contexts of a concept label; and (iii) mostly focusing on concept placement w.r.t a taxonomy of atomic concepts, instead of complex concepts, i.e., with logical operators. To address these issues, we propose a new benchmark, adapting MedMentions dataset (PubMed abstracts) with SNOMED CT versions in 2014 and 2017 under the Diseases sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic product. We provide usage on the evaluation with the dataset for out-of-KB mention discovery and concept placement, adapting recent Large Language Model based methods.
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Ian Horrocks 0001
CIKM1
2023 Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking
abstract
Discovering entity mentions that are out of a Knowledge Base (KB) from texts plays a critical role in KB maintenance, but has not yet been fully explored. The current methods are mostly limited to the simple threshold-based approach and feature-based classification, and the datasets for evaluation are relatively rare. We propose BLINKout, a new BERT-based Entity Linking (EL) method which can identify mentions that do not have corresponding KB entities by matching them to a special NIL entity. To better utilize BERT, we propose new techniques including NIL entity representation and classification, with synonym enhancement. We also apply KB Pruning and Versioning strategies to automatically construct out-of-KB datasets from common in-KB EL datasets. Results on five datasets of clinical notes, biomedical publications, and Wikipedia articles in various domains show the advantages of BLINKout over existing methods to identify out-of-KB mentions for the medical ontologies, UMLS, SNOMED CT, and the general KB, WikiData.
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yinan Liu 0001, Ian Horrocks 0001
CIKM1
2023 Subsumption Prediction for E-Commerce Taxonomies
Jingchuan Shi, Jiaoyan Chen 0001, Hang Dong 0002, Ishita K. Khan, Lizzie Liang, Qunzhi Zhou, Ian Horrocks 0001
ESWC3
2023 Contextual semantic embeddings for ontology subsumption prediction
abstract
Automating ontology construction and curation is an important but challenging task in knowledge engineering and artificial intelligence. Prediction by machine learning techniques such as contextual semantic embedding is a promising direction, but the relevant research is still preliminary especially for expressive ontologies in Web Ontology Language (OWL). In this paper, we present a new subsumption prediction method named BERTSubs for classes of OWL ontology. It exploits the pre-trained language model BERT to compute contextual embeddings of a class, where customized templates are proposed to incorporate the class context (e.g., neighbouring classes) and the logical existential restriction. BERTSubs is able to predict multiple kinds of subsumers including named classes from the same ontology or another ontology, and existential restrictions from the same ontology. Extensive evaluation on five real-world ontologies for three different subsumption tasks has shown the effectiveness of the templates and that BERTSubs can dramatically outperform the baselines that use (literal-aware) knowledge graph embeddings, non-contextual word embeddings and the state-of-the-art OWL ontology embeddings.
Jiaoyan Chen 0001, Yuan He 0008, Yuxia Geng, Ernesto Jiménez-Ruiz, Hang Dong 0002, Ian Horrocks 0001
World Wide Web (WWW)5
2022 Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching
abstract
Ontology Matching (OM) plays an important role in many domains such as bioinformatics and the Semantic Web, and its research is becoming increasingly popular, especially with the application of machine learning (ML) techniques. Although the Ontology Alignment Evaluation Initiative (OAEI) represents an impressive effort for the systematic evaluation of OM systems, it still suffers from several limitations including limited evaluation of subsumption mappings, suboptimal reference mappings, and limited support for the evaluation of ML-based systems. To tackle these limitations, we introduce five new biomedical OM tasks involving ontologies extracted from Mondo and UMLS. Each task includes both equivalence and subsumption matching; the quality of reference mappings is ensured by human curation, ontology pruning, etc.; and a comprehensive evaluation framework is proposed to measure OM performance from various perspectives for both ML-based and non-ML-based OM systems. We report evaluation results for OM systems of different types to demonstrate the usage of these resources, all of which are publicly available as part of the new Bio-ML track at OAEI 2022. Resource type: Ontology Matching Dataset License: CC BY 4.0 International DOI: https://doi.org/10.5281/zenodo.6510086 Documentation: https://krr-oxford.github.io/DeepOnto/#/om_resources OAEI track: https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/
Yuan He 0008, Jiaoyan Chen 0001, Hang Dong 0002, Ernesto Jiménez-Ruiz, Ali Hadian 0001, Ian Horrocks 0001
ISWC3
2021 CoPHE: A Count-Preserving Hierarchical Evaluation Metric in Large-Scale Multi-Label Text Classification
abstract
Large-Scale Multi-Label Text Classification (LMTC) includes tasks with hierarchical label spaces, such as automatic assignment of ICD-9 codes to discharge summaries.Performance of models in prior art is evaluated with standard precision, recall, and F 1 measures without regard for the rich hierarchical structure.In this work we argue for hierarchical evaluation of the predictions of neural LMTC models.With the example of the ICD-9 ontology we describe a structural issue in the representation of the structured label space in prior art, and propose an alternative representation based on the depth of the ontology.We propose a set of metrics for hierarchical evaluation using the depthbased representation.We compare the evaluation scores from the proposed metrics with previously used metrics on prior art LMTC models for ICD-9 coding in MIMIC-III.We also propose further avenues of research involving the proposed ontological representation.
Matús Falis, Hang Dong 0002, Alexandra Birch, Beatrice Alex
EMNLP (1)2
2021 Explainable automated coding of clinical notes using hierarchical label-wise attention networks and label embedding initialisation
Hang Dong 0002, Víctor Suárez-Paniagua, William Whiteley, Honghan Wu
J. Biomed. Informatics1
2021 Automated Social Text Annotation With Joint Multilabel Attention Networks
abstract
Automated social text annotation is the task of suggesting a set of tags for shared documents on social media platforms. The automated annotation process can reduce users' cognitive overhead in tagging and improve tag management for better search, browsing, and recommendation of documents. It can be formulated as a multilabel classification problem. We propose a novel deep learning-based method for this problem and design an attention-based neural network with semantic-based regularization, which can mimic users' reading and annotation behavior to formulate better document representation, leveraging the semantic relations among labels. The network separately models the title and the content of each document and injects an explicit, title-guided attention mechanism into each sentence. To exploit the correlation among labels, we propose two semantic-based loss regularizers, i.e., similarity and subsumption, which enforce the output of the network to conform to label semantics. The model with the semantic-based loss regularizers is referred to as the joint multilabel attention network (JMAN). We conducted a comprehensive evaluation study and compared JMAN to the state-of-the-art baseline models, using four large, real-world social media data sets. In terms of F1, JMAN significantly outperformed bidirectional gated recurrent unit (Bi-GRU) relatively by around 12.8%-78.6% and the hierarchical attention network (HAN) by around 3.9%-23.8%. The JMAN model demonstrates advantages in convergence and training speed. Further improvement of performance was observed against latent Dirichlet allocation (LDA) and support vector machine (SVM). When applying the semantic-based loss regularizers, the performance of HAN and Bi-GRU in terms of F1was also boosted. It is also found that dynamic update of the label semantic matrices (JMANd) has the potential to further improve the performance of JMAN but at the cost of substantial memory and warrants further study.
Hang Dong 0002, Wei Wang 0042, Kaizhu Huang, Frans Coenen
IEEE Trans. Neural Networks Learn. Syst.1
2020 Knowledge base enrichment by relation learning from social tagging data
Hang Dong 0002, Wei Wang 0042, Frans Coenen, Kaizhu Huang
Inf. Sci.1
2019 Motivations for self-archiving on an academic social networking site: A study on researchgate
abstract
This study investigates motivations for self‐archiving research items on academic social networking sites (ASNSs). A model of these motivations was developed based on two existing motivation models: motivation for self‐archiving in academia and motivations for information sharing in social media. The proposed model is composed of 18 factors drawn from personal, social, professional, and external contexts, including enjoyment, personal/professional gain, reputation, learning, self‐efficacy, altruism, reciprocity, trust, community interest, social engagement, publicity, accessibility, self‐archiving culture, influence of external actors, credibility, system stability, copyright concerns, additional time, and effort. Two hundred and twenty‐six ResearchGate users participated in the survey. Accessibility was the most highly rated factor, followed by altruism, reciprocity, trust, self‐efficacy, reputation, publicity, and others. Personal, social, and professional factors were also highly rated, while external factors were rated relatively low. Motivations were correlated with one another, demonstrating that RG motivations for self‐archiving could increase or decrease based on several factors in combination with motivations from the personal, social, professional, and external contexts. We believe the findings from this study can increase our understanding of users' motivations in sharing their research and provide useful implications for the development and improvement of ASNS services, thereby attracting more active users.
Jongwook Lee, Sanghee Oh, Hang Dong 0002, Gary Burnett
J. Assoc. Inf. Sci. Technol.3
2018 Learning Relations from Social Tagging Data
Hang Dong 0002, Wei Wang 0042, Frans Coenen
PRICAI (1)1