VLDB 2026 Research / reviewers in the wild / expert
Yuan He 0008
dblp:11/1735-8
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-4486-1262ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar ArgumentationabstractRetrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, yet suffers from critical limitations in high-stakes domains—namely, sensitivity to noisy or contradictory evidence and opaque, stochastic decision-making. We propose \textsc{ArgRAG}, an explainable, and contestable alternative that replaces black-box reasoning with structured inference using a Quantitative Bipolar Argumentation Framework (QBAF). \textsc{ArgRAG} constructs a QBAF from retrieved documents and performs deterministic reasoning under gradual semantics. This allows faithfully explanaining and contesting decisions. Evaluated on two fact verification benchmarks, PubHealth and RAGuard, \textsc{ArgRAG} achieves strong accuracy while significantly improving transparency. Yuqicheng Zhu, Nico Potyka, Daniel Hernández 0002, Yuan He 0008, Zifeng Ding, Bo Xiong 0001, Dongzhuoran Zhou, Evgeny Kharlamov, Steffen Staab |
NeSy | 4 |
| 2025 | Language Models as Ontology Encoders
Jiaoyan Chen 0001, Yuan He 0008, Yongsheng Gao 0005, Ian Horrocks 0001 |
ISWC (1) | 3 |
| 2025 | Ontology Embedding: A Survey of Methods, Applications and ResourcesabstractOntologies are widely used for representing domain knowledge and meta data, playing an increasingly important role in Information Systems, the Semantic Web, Bioinformatics and many other domains. However, logical reasoning that ontologies can directly support are quite limited in learning, approximation and prediction. One straightforward solution is to integrate statistical analysis and machine learning. To this end, automatically learning vector representation for knowledge of an ontology i.e.,ontology embeddinghas been widely investigated. Numerous papers have been published on ontology embedding, but a lack of systematic reviews hinders researchers from gaining a comprehensive understanding of this field. To bridge this gap, we write this survey paper, which first introduces different kinds of semantics of ontologies and formally defines ontology embedding as well as its property of faithfulness. Based on this, it systematically categorizes and analyses a relatively complete set of over 80 papers, according to the ontologies they aim at and their technical solutions including geometric modeling, sequence modeling and graph propagation. This survey also introduces the applications of ontology embedding in ontology engineering, machine learning augmentation and life sciences, presents a new library mOWL and discusses the challenges and future directions. Jiaoyan Chen 0001, Olga Mashkova, Fernando Zhapa-Camacho, Robert Hoehndorf, Yuan He 0008, Ian Horrocks 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | A Language Model Based Framework for New Concept Placement in Ontologies
Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yongsheng Gao 0005, Ian Horrocks 0001 |
ESWC (1) | 3 |
| 2024 | Language Models as Hierarchy EncodersabstractInterpreting hierarchical structures latent in language is a key limitation of current language models (LMs). While previous research has implicitly leveraged these hierarchies to enhance LMs, approaches for their explicit encoding are yet to be explored. To address this, we introduce a novel approach to re-train transformer encoder-based LMs as Hierarchy Transformer encoders (HiTs), harnessing the expansive nature of hyperbolic space. Our method situates the output embedding space of pre-trained LMs within a Poincaré ball with a curvature that adapts to the embedding dimension, followed by re-training on hyperbolic clustering and centripetal losses. These losses are designed to effectively cluster related entities (input as texts) and organise them hierarchically. We evaluate HiTs against pre-trained LMs, standard fine-tuned LMs, and several hyperbolic embedding baselines, focusing on their capabilities in simulating transitive inference, predicting subsumptions, and transferring knowledge across hierarchies. The results demonstrate that HiTs consistently outperform all baselines in these tasks, underscoring the effectiveness and transferability of our re-trained hierarchy encoders. Yuan He 0008, Moy Yuan, Jiaoyan Chen 0001, Ian Horrocks 0001 |
NeurIPS | 1 |
| 2023 | Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and PlacementabstractMentions of new concepts appear regularly in texts and require automated approaches to harvest and place them into Knowledge Bases (KB), e.g., ontologies and taxonomies. Existing datasets suffer from three issues, (i) mostly assuming that a new concept is pre-discovered and cannot support out-of-KB mention discovery; (ii) only using the concept label as the input along with the KB and thus lacking the contexts of a concept label; and (iii) mostly focusing on concept placement w.r.t a taxonomy of atomic concepts, instead of complex concepts, i.e., with logical operators. To address these issues, we propose a new benchmark, adapting MedMentions dataset (PubMed abstracts) with SNOMED CT versions in 2014 and 2017 under the Diseases sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic product. We provide usage on the evaluation with the dataset for out-of-KB mention discovery and concept placement, adapting recent Large Language Model based methods. Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Ian Horrocks 0001 |
CIKM | 3 |
| 2023 | Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity LinkingabstractDiscovering entity mentions that are out of a Knowledge Base (KB) from texts plays a critical role in KB maintenance, but has not yet been fully explored. The current methods are mostly limited to the simple threshold-based approach and feature-based classification, and the datasets for evaluation are relatively rare. We propose BLINKout, a new BERT-based Entity Linking (EL) method which can identify mentions that do not have corresponding KB entities by matching them to a special NIL entity. To better utilize BERT, we propose new techniques including NIL entity representation and classification, with synonym enhancement. We also apply KB Pruning and Versioning strategies to automatically construct out-of-KB datasets from common in-KB EL datasets. Results on five datasets of clinical notes, biomedical publications, and Wikipedia articles in various domains show the advantages of BLINKout over existing methods to identify out-of-KB mentions for the medical ontologies, UMLS, SNOMED CT, and the general KB, WikiData. Hang Dong 0002, Jiaoyan Chen 0001, Yuan He 0008, Yinan Liu 0001, Ian Horrocks 0001 |
CIKM | 3 |
| 2023 | Zero-Shot and Few-Shot Learning With Knowledge Graphs: A Comprehensive SurveyabstractMachine learning (ML), especially deep neural networks, has achieved great success, but many of them often rely on a number of labeled samples for supervision. As sufficient labeled training data are not always ready due to, e.g., continuously emerging prediction targets and costly sample annotation in real-world applications, ML with sample shortage is now being widely investigated. Among all these studies, many prefer to utilize auxiliary information including those in the form of knowledge graph (KG) to reduce the reliance on labeled samples. In this survey, we have comprehensively reviewed over 90 articles about KG-aware research for two major sample shortage settings—zero-shot learning (ZSL) where some classes to be predicted have no labeled samples and few-shot learning (FSL) where some classes to be predicted have only a small number of labeled samples that are available. We first introduce KGs used in ZSL and FSL as well as their construction methods and then systematically categorize and summarize KG-aware ZSL and FSL methods, dividing them into different paradigms, such as the mapping-based, the data augmentation, the propagation-based, and the optimization-based. We next present different applications, including not only KG augmented prediction tasks such as image classification, question answering, text classification, and knowledge extraction but also KG completion tasks and some typical evaluation resources for each task. We eventually discuss some challenges and open problems from different perspectives. Jiaoyan Chen 0001, Yuxia Geng, Zhuo Chen 0007, Jeff Z. Pan, Yuan He 0008, Wen Zhang 0015, Ian Horrocks 0001, Huajun Chen |
Proc. IEEE | 5 |
| 2023 | Contextual semantic embeddings for ontology subsumption predictionabstractAutomating ontology construction and curation is an important but challenging task in knowledge engineering and artificial intelligence. Prediction by machine learning techniques such as contextual semantic embedding is a promising direction, but the relevant research is still preliminary especially for expressive ontologies in Web Ontology Language (OWL). In this paper, we present a new subsumption prediction method named BERTSubs for classes of OWL ontology. It exploits the pre-trained language model BERT to compute contextual embeddings of a class, where customized templates are proposed to incorporate the class context (e.g., neighbouring classes) and the logical existential restriction. BERTSubs is able to predict multiple kinds of subsumers including named classes from the same ontology or another ontology, and existential restrictions from the same ontology. Extensive evaluation on five real-world ontologies for three different subsumption tasks has shown the effectiveness of the templates and that BERTSubs can dramatically outperform the baselines that use (literal-aware) knowledge graph embeddings, non-contextual word embeddings and the state-of-the-art OWL ontology embeddings. Jiaoyan Chen 0001, Yuan He 0008, Yuxia Geng, Ernesto Jiménez-Ruiz, Hang Dong 0002, Ian Horrocks 0001 |
World Wide Web (WWW) | 2 |
| 2022 | BERTMap: A BERT-Based Ontology Alignment SystemabstractOntology alignment (a.k.a ontology matching (OM)) plays a critical role in knowledge integration. Owing to the success of machine learning in many domains, it has been applied in OM. However, the existing methods, which often adopt ad-hoc feature engineering or non-contextual word embeddings, have not yet outperformed rule-based systems especially in an unsupervised setting. In this paper, we propose a novel OM system named BERTMap which can support both unsupervised and semi-supervised settings. It first predicts mappings using a classifier based on fine-tuning the contextual embedding model BERT on text semantics corpora extracted from ontologies, and then refines the mappings through extension and repair by utilizing the ontology structure and logic. Our evaluation with three alignment tasks on biomedical ontologies demonstrates that BERTMap can often perform better than the leading OM systems LogMap and AML. Yuan He 0008, Jiaoyan Chen 0001, Denvar Antonyrajah, Ian Horrocks 0001 |
AAAI | 1 |
| 2022 | Machine Learning-Friendly Biomedical Datasets for Equivalence and Subsumption Ontology MatchingabstractOntology Matching (OM) plays an important role in many domains such as bioinformatics and the Semantic Web, and its research is becoming increasingly popular, especially with the application of machine learning (ML) techniques. Although the Ontology Alignment Evaluation Initiative (OAEI) represents an impressive effort for the systematic evaluation of OM systems, it still suffers from several limitations including limited evaluation of subsumption mappings, suboptimal reference mappings, and limited support for the evaluation of ML-based systems. To tackle these limitations, we introduce five new biomedical OM tasks involving ontologies extracted from Mondo and UMLS. Each task includes both equivalence and subsumption matching; the quality of reference mappings is ensured by human curation, ontology pruning, etc.; and a comprehensive evaluation framework is proposed to measure OM performance from various perspectives for both ML-based and non-ML-based OM systems. We report evaluation results for OM systems of different types to demonstrate the usage of these resources, all of which are publicly available as part of the new Bio-ML track at OAEI 2022. Resource type: Ontology Matching Dataset License: CC BY 4.0 International DOI: https://doi.org/10.5281/zenodo.6510086 Documentation: https://krr-oxford.github.io/DeepOnto/#/om_resources OAEI track: https://www.cs.ox.ac.uk/isg/projects/ConCur/oaei/ Yuan He 0008, Jiaoyan Chen 0001, Hang Dong 0002, Ernesto Jiménez-Ruiz, Ali Hadian 0001, Ian Horrocks 0001 |
ISWC | 1 |