VLDB 2026 Research / reviewers in the wild / expert
Ying He 0010
dblp:39/2405-10
· DBLP profile ↗
8ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0002-9994-2051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language ModelsabstractYing He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui HE, Shimin Tao, Mahongxia, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ying He 0010, Sihang Jiang 0001, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui He, Shimin Tao, Hongxia Ma, Yanghua Xiao |
ACL (1) | 1 |
| 2026 | Caf4AVC: LLM-Enhanced Collaborative Framework for Attribute Value Canonicalization in Open KBsabstractOpen Knowledge Bases (Open KBs) are fundamental to knowledge-driven applications, including semantic search, knowledge reasoning, and recommendation systems. However, the presence of redundant and ambiguous expressions within Open KBs significantly hinders their application. This highlights the urgent need for Open KB canonicalization, particularly of attribute values, which comprise nearly 40% of the facts within Open KBs. Unlike entities and predicates, attribute values are inherently sparse and diverse, posing unique challenges for their canonicalization. However, existing studies mainly focus on entities or predicates, leaving attribute value-level noun phrase canonicalization (NPC-AV) underexplored. Large language models (LLMs), with their strengths in common-sense reasoning and fault tolerance, have shown promise in Open KB canonicalization. Yet, current LLM-based approaches often rely heavily on LLM responses, overlooking their high computational cost and potential errors. In this paper, we introduce Caf4AVC, a collaborative framework that integrates clustering-based methods and LLMs for the NPC-AV task. We further propose an innovative two-factor authentication correction mechanism and an adaptive threshold-based selection strategy to address these limitations. Extensive experiments on multiple real-world Open KB datasets demonstrate the effectiveness of our framework, achieving a 17.52% reduction in LLM call costs and a 6.3% average performance improvement compared to competitive methods. The code and dataset are available athttps://github.com/hedyHe/Caf4AV. Ying He 0010, Qiang Yang 0015, Zhouhong Gu, Zhixu Li, Yanghua Xiao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | A Joint Link-Retrieve Framework for Open Table-and-Text Question Answering
Jiaan Wang, Ying He 0010, Jianfeng Qu, Zhixu Li, Pengpeng Zhao 0001, An Liu 0002, Lei Zhao 0001 |
DASFAA (3) | 3 |
| 2022 | RT-KGD: Relation Transition Aware Knowledge-Grounded Dialogue Generation
Zhixu Li, Jiaan Wang, Jianfeng Qu, Ying He 0010, An Liu 0002, Lei Zhao 0001 |
ISWC | 5 |
| 2020 | Fine-Grained Entity Typing for Relation-Sparsity Entities
Lei Niu, Binbin Gu, Zhixu Li, Wei Chen 0070, Ying He 0010, Zhaoyin Zhang, Zhigang Chen 0003 |
DASFAA (2) | 5 |
| 2020 | End-to-end relation extraction based on bootstrapped multi-level distant supervision
Ying He 0010, Zhixu Li, Qiang Yang 0015, Zhigang Chen 0003, An Liu 0002, Lei Zhao 0001, Xiaofang Zhou 0001 |
World Wide Web | 1 |
| 2018 | Bootstrapped Multi-level Distant Supervision for Relation Extraction
Ying He 0010, Zhixu Li, Guanfeng Liu 0001, Fangfei Cao, Zhigang Chen 0003 |
WISE (1) | 1 |
| 2018 | Diagnosing and Minimizing Semantic Drift in Iterative Bootstrapping ExtractionabstractSemantic drift is a common problem in iterative information extraction. Previous approaches for minimizing semantic drift may incur substantial loss in recall. We observe that most semantic drifts are introduced by a small number of questionable extractions in the earlier rounds of iterations. These extractions subsequently introduce a large number of questionable results, which lead to the semantic drift phenomenon. We call these questionable extractions Drifting Points (DPs). If erroneous extractions are the “symptoms” of semantic drift, then DPs are the “causes” of semantic drift. In this paper, we propose a method to minimize semantic drift by identifying the DPs and removing the effect introduced by the DPs. We use isA (concept-instance) extraction as an example to describe our approach in cleaning information extraction errors caused by semantic drift, but we perform experiments on different relation extraction processes on three large real data extraction collections. The experimental results show that our DP cleaning method enables us to clean around 90 percent incorrect instances or patterns with about 90 percent precision, which outperforms the previous approaches we compare with. Zhixu Li, Ying He 0010, Binbin Gu, An Liu 0002, Hongsong Li, Haixun Wang, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |