Yikai Guo

dblp:334/4154 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0003-0345-1686ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ConsistEAE: Enhancing low-resource event argument extraction with linguistically consistent demonstrations
Yikai Guo, Xuemeng Tian, Bin Ge 0006, Wenjun Ke 0002, Yanyang Li, Haoran Luo 0001
Neurocomputing1
2026 CGDA: Enhancing low-resource event causality extraction via consistency-guided data augmentation
Zhi Fang, Xuemeng Tian, Yikai Guo
Neurocomputing7
2025 KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search
abstract
Knowledge Base Question Answering (KBQA) aims to answer natural language questions with a large-scale structured knowledge base (KB). Despite advancements with large language models (LLMs), KBQA still faces challenges in weak KB awareness, imbalance between effectiveness and efficiency, and high reliance on annotated data. To address these challenges, we propose KBQA-o1, a novel agentic KBQA method with Monte Carlo Tree Search (MCTS). It introduces a ReAct-based agent process for stepwise logical form generation with KB environment exploration. Moreover, it employs MCTS, a heuristic search method driven by policy and reward models, to balance agentic exploration’s performance and search space. With heuristic exploration, KBQA-o1 generates high-quality annotations for further improvement by incremental fine-tuning. Experimental results show that KBQA-o1 outperforms previous low-resource KBQA methods with limited annotated data, boosting Llama-3.1-8B model’s GrailQA F1 performance to 78.5% compared to 48.5% of the previous sota method with GPT-3.5-turbo. Our code is publicly available.
Haoran Luo 0001, Haihong E, Yikai Guo, Qika Lin, Xiaobao Wu, Xinyu Mu, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
ICML3
2025 HyperGraphRAG: Retrieval-Augmented Generation via Hypergraph-Structured Knowledge Representation
abstract
Standard Retrieval-Augmented Generation (RAG) relies on chunk-based retrieval, whereas GraphRAG advances this approach by graph-based knowledge representation. However, existing graph-based RAG approaches are constrained by binary relations, as each edge in an ordinary graph connects only two entities, limiting their ability to represent the n-ary relations (n >= 2) in real-world knowledge. In this work, we propose HyperGraphRAG, the first hypergraph-based RAG method that represents n-ary relational facts via hyperedges. HyperGraphRAG consists of a comprehensive pipeline, including knowledge hypergraph construction, retrieval, and generation. Experiments across medicine, agriculture, computer science, and law demonstrate that HyperGraphRAG outperforms both standard RAG and previous graph-based RAG methods in answer accuracy, retrieval efficiency, and generation quality.
Haoran Luo 0001, Haihong E, Guanting Chen 0004, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng 0015, Zemin Kuang, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS6
2024 Boosting LLMS with Ontology-Aware Prompt for Ner Data Augmentation
abstract
Named Entity Recognition (NER) data augmentation (DA) aims to improve the performance and generalization capabilities of NER models by generating scalable training data. The key challenge lies in ensuring the generated samples maintain contextual diversity while preserving label consistency. However, existing dominant methods fail to simultaneously satisfy both criteria. Inspired by the extensive generative capabilities of large language models (LLMs), we propose ANGEL, a frAmework integrating the oNtoloGy structure and instructivE prompting within LLMs. Specifically, the hierarchical ontology structure guides prompt ranking, while instructive prompting enhances LLMs’ mastery of domain knowledge, empowering synthetic sample generation and annotation. Experiments show ANGEL surpasses state-of-the-art (SOTA) baselines, conferring absolute F1 increases of 2.86% and 0.93% on two benchmark datasets, respectively.
Zhizhao Luo, Youchen Wang, Wenjun Ke 0002, Yikai Guo, Peng Wang 0004
ICASSP5
2024 Recall, Retrieve and Reason: Towards Better In-Context Relation Extraction
Peng Wang 0004, Wenjun Ke 0002, Yikai Guo, Ke Ji, Ziyu Shang, Jiajun Liu 0005, Zijie Xu 0003
IJCAI4
2024 Meta In-Context Learning Makes Large Language Models Better Zero and Few-Shot Relation Extractors
Peng Wang 0004, Jiajun Liu 0005, Yikai Guo, Ke Ji, Ziyu Shang, Zijie Xu 0003
IJCAI4
2024 Empirical Analysis of Dialogue Relation Extraction with Large Language Models
Zijie Xu 0003, Ziyu Shang, Jiajun Liu 0005, Ke Ji, Yikai Guo
IJCAI6
2024 Text2NKG: Fine-Grained N-ary Relation Extraction for N-ary relational Knowledge Graph Construction
abstract
Beyond traditional binary relational facts, n-ary relational knowledge graphs (NKGs) are comprised of n-ary relational facts containing more than two entities, which are closer to real-world facts with broader applications. However, the construction of NKGs remains at a coarse-grained level, which is always in a single schema, ignoring the order and variable arity of entities. To address these restrictions, we propose Text2NKG, a novel fine-grained n-ary relation extraction framework for n-ary relational knowledge graph construction. We introduce a span-tuple classification approach with hetero-ordered merging and output merging to accomplish fine-grained n-ary relation extraction in different arity. Furthermore, Text2NKG supports four typical NKG schemas: hyper-relational schema, event-based schema, role-based schema, and hypergraph-based schema, with high flexibility and practicality. The experimental results demonstrate that Text2NKG achieves state-of-the-art performance in F1 scores on the fine-grained n-ary relation extraction benchmark. Our code and datasets are publicly available.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Tianyu Yao, Yikai Guo, Zichen Tang, Wentai Zhang 0004, Shiyao Peng, Kaiyang Wan, Meina Song, Yifan Zhu 0001, Anh Tuan Luu
NeurIPS5
2024 Unveiling factuality and injecting knowledge for LLMs via reinforcement learning and data proportion
Wenjun Ke 0002, Ziyu Shang, Zhizhao Luo, Peng Wang 0004, Yikai Guo, Qi Liu 0056
Sci. China Inf. Sci.5
2024 MDM: Meta diffusion model for hard-constrained text generation
Wenjun Ke 0002, Yikai Guo, Qi Liu 0056, Peng Wang 0004, Haoran Luo 0001, Zhizhao Luo
Knowl. Based Syst.2
2024 Agent-DA: Enhancing low-resource event extraction with collaborative multi-agent data augmentation
Xuemeng Tian, Yikai Guo, Bin Ge 0006, Wenjun Ke 0002
Knowl. Based Syst.2
2023 NQE: N-ary Query Embedding for Complex Query Answering over Hyper-Relational Knowledge Graphs
abstract
Complex query answering (CQA) is an essential task for multi-hop and logical reasoning on knowledge graphs (KGs). Currently, most approaches are limited to queries among binary relational facts and pay less attention to n-ary facts (n≥2) containing more than two entities, which are more prevalent in the real world. Moreover, previous CQA methods can only make predictions for a few given types of queries and cannot be flexibly extended to more complex logical queries, which significantly limits their applications. To overcome these challenges, in this work, we propose a novel N-ary Query Embedding (NQE) model for CQA over hyper-relational knowledge graphs (HKGs), which include massive n-ary facts. The NQE utilizes a dual-heterogeneous Transformer encoder and fuzzy logic theory to satisfy all n-ary FOL queries, including existential quantifiers (∃), conjunction (∧), disjunction (∨), and negation (¬). We also propose a parallel processing algorithm that can train or predict arbitrary n-ary FOL queries in a single batch, regardless of the kind of each query, with good flexibility and extensibility. In addition, we generate a new CQA dataset WD50K-NFOL, including diverse n-ary FOL queries over WD50K. Experimental results on WD50K-NFOL and other standard CQA datasets show that NQE is the state-of-the-art CQA method over HKGs with good generalization capability. Our code and dataset are publicly available.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Gengxian Zhou, Yikai Guo, Tianyu Yao, Zichen Tang, Xueyuan Lin, Kaiyang Wan
AAAI5
2023 HAHE: Hierarchical Attention for Hyper-Relational Knowledge Graphs in Global and Local Level
abstract
Haoran Luo, Haihong E, Yuhao Yang, Yikai Guo, Mingzhi Sun, Tianyu Yao, Zichen Tang, Kaiyang Wan, Meina Song, Wei Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Haoran Luo 0001, Haihong E, Yuhao Yang 0006, Yikai Guo, Mingzhi Sun, Tianyu Yao, Zichen Tang, Kaiyang Wan, Meina Song
ACL (1)4