Hang Dong 0002

dblp:135/8614-2 · DBLP profile ↗
← Back
8ranked-venue papers in the field
4as first author
6since 2021 · last 2024
0000-0001-6828-6891ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 4 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)
YearPublicationVenuePosition
2024 A Language Model Based Framework for New Concept Placement in Ontologies
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yongsheng Gao 0005, Ian Horrocks 0001
ESWC (1)1
2024 Taxonomy Completion via Implicit Concept Insertion
abstract
\beginabstract High quality taxonomies play a critical role in various domains such as e-commerce, web search and ontology engineering. While there has been extensive work on expanding taxonomies from externally mined data, there has been less attention paid to enriching taxonomies by exploiting existing concepts and structure within the taxonomy. In this work, we show the usefulness of this kind of enrichment, and explore its viability with a new taxonomy completion system ICON (I mplicit CON cept Insertion). ICON generates new concepts by identifying implicit concepts based on the existing concept structure, generating names for such concepts and inserting them in appropriate positions within the taxonomy. ICON integrates techniques from entity retrieval, text summary, and subsumption prediction; this modular architecture offers high flexibility while achieving state-of-the-art performance. We have evaluated ICON on two e-commerce taxonomies, and the results show that it offers significant advantages over strong baselines including recent taxonomy completion models and the large language model, ChatGPT.
Jingchuan Shi, Hang Dong 0002, Jiaoyan Chen 0001, Ian Horrocks 0001
WWW2
2023 Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement
abstract
Mentions of new concepts appear regularly in texts and require automated approaches to harvest and place them into Knowledge Bases (KB), e.g., ontologies and taxonomies. Existing datasets suffer from three issues, (i) mostly assuming that a new concept is pre-discovered and cannot support out-of-KB mention discovery; (ii) only using the concept label as the input along with the KB and thus lacking the contexts of a concept label; and (iii) mostly focusing on concept placement w.r.t a taxonomy of atomic concepts, instead of complex concepts, i.e., with logical operators. To address these issues, we propose a new benchmark, adapting MedMentions dataset (PubMed abstracts) with SNOMED CT versions in 2014 and 2017 under the Diseases sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic product. We provide usage on the evaluation with the dataset for out-of-KB mention discovery and concept placement, adapting recent Large Language Model based methods.
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Ian Horrocks 0001
CIKM1
2023 Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking
abstract
Discovering entity mentions that are out of a Knowledge Base (KB) from texts plays a critical role in KB maintenance, but has not yet been fully explored. The current methods are mostly limited to the simple threshold-based approach and feature-based classification, and the datasets for evaluation are relatively rare. We propose BLINKout, a new BERT-based Entity Linking (EL) method which can identify mentions that do not have corresponding KB entities by matching them to a special NIL entity. To better utilize BERT, we propose new techniques including NIL entity representation and classification, with synonym enhancement. We also apply KB Pruning and Versioning strategies to automatically construct out-of-KB datasets from common in-KB EL datasets. Results on five datasets of clinical notes, biomedical publications, and Wikipedia articles in various domains show the advantages of BLINKout over existing methods to identify out-of-KB mentions for the medical ontologies, UMLS, SNOMED CT, and the general KB, WikiData.
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yinan Liu 0001, Ian Horrocks 0001
CIKM1
2023 Subsumption Prediction for E-Commerce Taxonomies
Jingchuan Shi, Jiaoyan Chen 0001, Hang Dong 0002, Ishita K. Khan, Lizzie Liang, Qunzhi Zhou, Ian Horrocks 0001
ESWC3
2022 Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology Matching
abstract
Ontology Matching (OM) plays an important role in many domains such as bioinformatics and the Semantic Web, and its research is becoming increasingly popular, especially with the application of machine learning (ML) techniques. Although the Ontology Alignment Evaluation Initiative (OAEI) represents an impressive effort for the systematic evaluation of OM systems, it still suffers from several limitations including limited evaluation of subsumption mappings, suboptimal reference mappings, and limited support for the evaluation of ML-based systems. To tackle these limitations, we introduce five new biomedical OM tasks involving ontologies extracted from Mondo and UMLS. Each task includes both equivalence and subsumption matching; the quality of reference mappings is ensured by human curation, ontology pruning, etc.; and a comprehensive evaluation framework is proposed to measure OM performance from various perspectives for both ML-based and non-ML-based OM systems. We report evaluation results for OM systems of different types to demonstrate the usage of these resources, all of which are publicly available as part of the new Bio-ML track at OAEI 2022. Resource type: Ontology Matching Dataset License: CC BY 4.0 International DOI: https://doi.org/10.5281/zenodo.6510086 Documentation: https://krr-oxford.github.io/DeepOnto/#/om_resources OAEI track: https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/
Yuan He 0008, Jiaoyan Chen 0001, Hang Dong 0002, Ernesto Jiménez-Ruiz, Ali Hadian 0001, Ian Horrocks 0001
ISWC3
2020 Knowledge base enrichment by relation learning from social tagging data
Hang Dong 0002, Wei Wang 0042, Frans Coenen, Kaizhu Huang
Inf. Sci.1
2019 Motivations for self-archiving on an academic social networking site: A study on researchgate
abstract
This study investigates motivations for self‐archiving research items on academic social networking sites (ASNSs). A model of these motivations was developed based on two existing motivation models: motivation for self‐archiving in academia and motivations for information sharing in social media. The proposed model is composed of 18 factors drawn from personal, social, professional, and external contexts, including enjoyment, personal/professional gain, reputation, learning, self‐efficacy, altruism, reciprocity, trust, community interest, social engagement, publicity, accessibility, self‐archiving culture, influence of external actors, credibility, system stability, copyright concerns, additional time, and effort. Two hundred and twenty‐six ResearchGate users participated in the survey. Accessibility was the most highly rated factor, followed by altruism, reciprocity, trust, self‐efficacy, reputation, publicity, and others. Personal, social, and professional factors were also highly rated, while external factors were rated relatively low. Motivations were correlated with one another, demonstrating that RG motivations for self‐archiving could increase or decrease based on several factors in combination with motivations from the personal, social, professional, and external contexts. We believe the findings from this study can increase our understanding of users' motivations in sharing their research and provide useful implications for the development and improvement of ASNS services, thereby attracting more active users.
Jongwook Lee, Sanghee Oh, Hang Dong 0002, Gary Burnett
J. Assoc. Inf. Sci. Technol.3