Ying He 0010

dblp:39/2405-10 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0002-9994-2051ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models
abstract
Ying He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui HE, Shimin Tao, Mahongxia, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ying He 0010, Sihang Jiang 0001, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui He, Shimin Tao, Hongxia Ma, Yanghua Xiao
ACL (1)1
2026 Caf4AVC: LLM-Enhanced Collaborative Framework for Attribute Value Canonicalization in Open KBs
abstract
Open Knowledge Bases (Open KBs) are fundamental to knowledge-driven applications, including semantic search, knowledge reasoning, and recommendation systems. However, the presence of redundant and ambiguous expressions within Open KBs significantly hinders their application. This highlights the urgent need for Open KB canonicalization, particularly of attribute values, which comprise nearly 40% of the facts within Open KBs. Unlike entities and predicates, attribute values are inherently sparse and diverse, posing unique challenges for their canonicalization. However, existing studies mainly focus on entities or predicates, leaving attribute value-level noun phrase canonicalization (NPC-AV) underexplored. Large language models (LLMs), with their strengths in common-sense reasoning and fault tolerance, have shown promise in Open KB canonicalization. Yet, current LLM-based approaches often rely heavily on LLM responses, overlooking their high computational cost and potential errors. In this paper, we introduce Caf4AVC, a collaborative framework that integrates clustering-based methods and LLMs for the NPC-AV task. We further propose an innovative two-factor authentication correction mechanism and an adaptive threshold-based selection strategy to address these limitations. Extensive experiments on multiple real-world Open KB datasets demonstrate the effectiveness of our framework, achieving a 17.52% reduction in LLM call costs and a 6.3% average performance improvement compared to competitive methods. The code and dataset are available athttps://github.com/hedyHe/Caf4AV.
Ying He 0010, Qiang Yang 0015, Zhouhong Gu, Zhixu Li, Yanghua Xiao
IEEE Trans. Knowl. Data Eng.1
2023 A Joint Link-Retrieve Framework for Open Table-and-Text Question Answering
Jiaan Wang, Ying He 0010, Jianfeng Qu, Zhixu Li, Pengpeng Zhao 0001, An Liu 0002, Lei Zhao 0001
DASFAA (3)3
2022 RT-KGD: Relation Transition Aware Knowledge-Grounded Dialogue Generation
Zhixu Li, Jiaan Wang, Jianfeng Qu, Ying He 0010, An Liu 0002, Lei Zhao 0001
ISWC5
2020 Fine-Grained Entity Typing for Relation-Sparsity Entities
Lei Niu, Binbin Gu, Zhixu Li, Wei Chen 0070, Ying He 0010, Zhaoyin Zhang, Zhigang Chen 0003
DASFAA (2)5
2020 End-to-end relation extraction based on bootstrapped multi-level distant supervision
Ying He 0010, Zhixu Li, Qiang Yang 0015, Zhigang Chen 0003, An Liu 0002, Lei Zhao 0001, Xiaofang Zhou 0001
World Wide Web1
2018 Bootstrapped Multi-level Distant Supervision for Relation Extraction
Ying He 0010, Zhixu Li, Guanfeng Liu 0001, Fangfei Cao, Zhigang Chen 0003
WISE (1)1
2018 Diagnosing and Minimizing Semantic Drift in Iterative Bootstrapping Extraction
abstract
Semantic drift is a common problem in iterative information extraction. Previous approaches for minimizing semantic drift may incur substantial loss in recall. We observe that most semantic drifts are introduced by a small number of questionable extractions in the earlier rounds of iterations. These extractions subsequently introduce a large number of questionable results, which lead to the semantic drift phenomenon. We call these questionable extractions Drifting Points (DPs). If erroneous extractions are the “symptoms” of semantic drift, then DPs are the “causes” of semantic drift. In this paper, we propose a method to minimize semantic drift by identifying the DPs and removing the effect introduced by the DPs. We use isA (concept-instance) extraction as an example to describe our approach in cleaning information extraction errors caused by semantic drift, but we perform experiments on different relation extraction processes on three large real data extraction collections. The experimental results show that our DP cleaning method enables us to clean around 90 percent incorrect instances or patterns with about 90 percent precision, which outperforms the previous approaches we compare with.
Zhixu Li, Ying He 0010, Binbin Gu, An Liu 0002, Hongsong Li, Haixun Wang, Xiaofang Zhou 0001
IEEE Trans. Knowl. Data Eng.2