Xiaohe Bo

dblp:353/7419 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0006-3853-3218ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph Retrieval-Augmented Generation: A Survey
abstract
Recently, Retrieval-Augmented Generation (RAG) has achieved remarkable success in addressing the challenges of Large Language Models (LLMs) without necessitating retraining. By referencing an external knowledge base, RAG refines LLM outputs, effectively mitigating issues such as “hallucination,” lack of domain-specific knowledge, and outdated information. However, the complex structure of relationships among different entities in databases presents challenges for RAG systems. In response, GraphRAG leverages structural information across entities to enable more precise and comprehensive retrieval, capturing relational knowledge and facilitating more accurate, context-aware responses. Given the novelty and potential of GraphRAG, a systematic review of current technologies is imperative. This article provides the first comprehensive overview of GraphRAG methodologies. We formalize the GraphRAG workflow, encompassing Graph-Based Indexing, Graph-Guided Retrieval, and Graph-Enhanced Generation. We then outline the core technologies and training methods at each stage. Additionally, we examine downstream tasks, application domains, evaluation methodologies, and industrial use cases of GraphRAG. Finally, we explore future research directions to inspire further inquiries and advance progress in the field. In order to track recent progress, we set up a repository at https://github.com/pengboci/GraphRAG-Survey .
Boci Peng, Yun Zhu 0007, Yongchao Liu 0004, Xiaohe Bo, Haizhou Shi, Chuntao Hong, Yan Zhang 0117, Siliang Tang
ACM Trans. Inf. Syst.4
2025 M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering Benchmark
abstract
Boci Peng, Yongchao Liu, Xiaohe Bo, Jiaxin Guo, Yun Zhu, Xuanbo Fan, Chuntao Hong, Yan Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Boci Peng, Yongchao Liu 0004, Xiaohe Bo, Yun Zhu 0007, Xuanbo Fan, Chuntao Hong, Yan Zhang 0117
ACL (1)3
2025 Incorporating Review-missing Interactions for Generative Explainable Recommendation
abstract
Explainable recommendation has attracted much attention from the academic and industry communities. Traditional models usually leverage user reviews as ground truths for model training, and the interactions without reviews are totally ignored. However, in practice, a large amount of users may not leave reviews after purchasing items. In this paper, we argue that the interactions without reviews may also contain comprehensive user preferences, and incorporating them to build explainable recommender model may further improve the explanation quality. To follow such intuition, we first leverage generative models to predict the missing reviews, and then train the recommender model based on all the predicted and original reviews. In specific, since the reviews are discrete tokens, we regard the review generation process as a reinforcement learning problem, where each token is an action at one step. We hope that the generated reviews are indistinguishable with the real ones. Thus, we introduce an discriminator as a reward model to evaluate the quality of the generated reviews. At last, to smooth the review generation process, we introduce a self-paced learning strategy to first generate shorter reviews and then predict the longer ones. We conduct extensive experiments on three publicly available datasets to demonstrate the effectiveness of our model.
Xiaohe Bo, Chen Ma 0001, Xu Chen 0017
COLING2
2025 CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
abstract
Current Large Language Models (LLMs) are confronted with overwhelming information volume when comprehending long-form documents. This challenge raises the imperative of a cohesive memory module, which can elevate vanilla LLMs into autonomous reading agents. Despite the emergence of some heuristic approaches, a systematic design principle remains absent. To fill this void, we draw inspiration from Jean Piaget's Constructivist Theory, illuminating three traits of the agentic memory---structured schemata, flexible assimilation, and dynamic accommodation. This blueprint forges a clear path toward a more robust and efficient memory system for LLM-based reading comprehension. To this end, we develop CAM, a prototype implementation of Constructivist Agentic Memory that simultaneously embodies the structurality, flexibility, and dynamicity. At its core, CAM is endowed with an incremental overlapping clustering algorithm for structured memory development, supporting both coherent hierarchical summarization and online batch integration. During inference, CAM adaptively explores the memory structure to activate query-relevant information for contextual response, akin to the human associative process. Compared to existing approaches, our design demonstrates dual advantages in both performance and efficiency across diverse long-text reading comprehension tasks, including question answering, query-based summarization, and claim verification.
Rui Li 0086, Zeyu Zhang 0007, Xiaohe Bo, Zihang Tian, Xu Chen 0017, Quanyu Dai, Zhenhua Dong, Ruiming Tang
NeurIPS3
2025 A Survey on the Memory Mechanism of Large Language Model-based Agents
abstract
Large language model (LLM)-based agents have recently attracted much attention from the research and industry communities. Compared with original LLMs, LLM-based agents are featured in their self-evolving capability, which is the basis for solving real-world problems that need long-term and complex agent-environment interactions. The key component to support agent-environment interactions is the memory of the agents. While previous studies have proposed many promising memory mechanisms, they are scattered in different papers, and there lacks a systematical review to summarize and compare these works from a holistic perspective, failing to abstract common and effective designing patterns for inspiring future studies. To bridge this gap, in this article, we propose a comprehensive survey on the memory mechanism of LLM-based agents. In specific, we first discuss “what is” and “why do we need” the memory in LLM-based agents. Then, we systematically review previous studies on how to design and evaluate the memory module. In addition, we also present many agent applications, where the memory module plays an important role. At last, we analyze the limitations of existing work and show important future directions. To keep up with the latest advances in this field, we create a repository at https://github.com/nuster1128/LLM_Agent_Memory_Survey .
Zeyu Zhang 0007, Quanyu Dai, Xiaohe Bo, Chen Ma 0001, Rui Li 0086, Xu Chen 0017, Jieming Zhu, Zhenhua Dong, Ji-Rong Wen
ACM Trans. Inf. Syst.3
2024 A Diffusion Model with User Preference Guidance for Recommendation
Boci Peng, Xiaohe Bo, Jiayan Guo
DASFAA (3)2
2024 Active Explainable Recommendation with Limited Labeling Budgets
abstract
Explainable recommendation has gained significant attention due to its potential to enhance user trust and system transparency. Previous studies primarily focus on refining model architectures to generate more informative explanations, assuming that the explanation data is sufficient and easy to acquire. However, in practice, obtaining the ground truth for explanations can be costly since individuals may not be inclined to put additional efforts to provide behavior explanations. In this paper, we study a novel problem in the field of explainable recommendation, that is, “given a limited budget to incentivize users to provide behavior explanations, how to effectively collect data such that the downstream models can be better optimized?” To solve this problem, we propose an active learning framework for recommender system, which consists of an acquisition function for sample collection and an explainable recommendation model to provide the final results. We consider both uncertainty and influence based strategies to design the acquisition function, which can determine the sample effectiveness from complementary perspectives. To demonstrate the effectiveness of our framework, we conduct extensive experiments based on real-world datasets.
Jingsen Zhang, Xiaohe Bo, Quanyu Dai, Zhenhua Dong, Ruiming Tang, Xu Chen 0017
ICASSP2
2024 Reflective Multi-Agent Collaboration based on Large Language Models
abstract
Benefiting from the powerful language expression and planning capabilities of Large Language Models (LLMs), LLM-based autonomous agents have achieved promising performance in various downstream tasks. Recently, based on the development of single-agent systems, researchers propose to construct LLM-based multi-agent systems to tackle more complicated tasks. In this paper, we propose a novel framework, named COPPER, to enhance the collaborative capabilities of LLM-based agents with the self-reflection mechanism. To improve the quality of reflections, we propose to fine-tune a shared reflector, which automatically tunes the prompts of actor models using our counterfactual PPO mechanism. On the one hand, we propose counterfactual rewards to assess the contribution of a single agent’s reflection within the system, alleviating the credit assignment problem. On the other hand, we propose to train a shared reflector, which enables the reflector to generate personalized reflections according to agent roles, while reducing the computational resource requirements and improving training stability. We conduct experiments on three datasets to evaluate the performance of our model in multi-hop question answering, mathematics, and chess scenarios. Experimental results show that COPPER possesses stronger reflection capabilities and exhibits excellent generalization performance across different actor models.
Xiaohe Bo, Zeyu Zhang 0007, Quanyu Dai, Xueyang Feng, Lei Wang 0198, Rui Li 0086, Xu Chen 0017, Ji-Rong Wen
NeurIPS1
2024 Subgraph Retrieval Enhanced by Graph-Text Alignment for Commonsense Question Answering
Boci Peng, Yongchao Liu 0004, Xiaohe Bo, Baokun Wang, Chuntao Hong, Yan Zhang 0117
ECML/PKDD (6)3
2023 Recommendation with Dynamic Natural Language Explanations
abstract
Explaining recommendation with natural languages has shown to be an effective strategy to improve the recommendation pervasiveness and user satisfaction. While recent years have witnessed many promising models, they mostly consider the user-item interactions as independent samples. However, in real-world scenarios, the user preference is always dynamic and evolving, and the current user behaviors may have strong correlations with the previous ones. To bridge this gap, in this paper, we propose to build an explainable recommender model by considering the user dynamic preference. Our general idea is to build a sequential model to capture the user history behaviors, and then the explanations are generated by summarizing all the past interactions. In specific, we firstly deploy two independent components to model the user sequential interactions and reviews separately. Then, we design a duration-aware attention mechanism to discriminate the importance of different items and reviews. For more effectively modeling the history information, we introduce a denoising module to remove the user behaviors which are less important for the current prediction. We conduct extensive experiments to demonstrate the effectiveness of our model based on three real-world datasets, in which the best performance can be improved by about 13.3%, 6.5%, 5.0% and 1.9% on the metrics of BLEU-1, ROUGE-1, ROUGE-2 and MAE, respectively. In addition, we also evaluate the generated explanations from both qualitative and qualitative perspectives.
Jingsen Zhang, Xiaohe Bo, Lei Wang 0198, Xu Chen 0017
IJCNN3