Zhanpeng Guan

dblp:24/4600 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-6024-921XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Bootstrapping in the Loop: Multi-hop Question Answering via Alternating Decomposition and Retrieval
abstract
Multi-hop question answering (QA) typically involves retrieving multiple relevant passages as evidence, which a reader then uses to derive the final answer. Existing methods often decompose complex questions into subquestions, leveraging large language models for step-by-step inference. However, these approaches fail to address the interdependence between retrieval and decomposition, resulting in suboptimal performance: Incomplete or inaccurate evidence can lead to poorly crafted subquestions, which, in turn, amplify the retrieval of irrelevant information and impede the reasoning process. To tackle this, we introduce BidLoop, a multi-step reasoning framework that explicitly models Bid irectional Loop between decomposition and retrieval. BidLoop employs four specialized modules: the Planner, Evaluator, Retriever, and Reader. In each reasoning round, the Planner generates a new subquestion based on prior evidence and subquestion-answer pairs. The Evaluator assesses whether the gathered evidence is sufficient to produce the final answer or if further reasoning is needed. Guided by the Planner's subquestion, the Retriever fetches relevant evidence, while the Reader answers the subquestion, adding new evidence for the next round. Our approach excels in generalization, performing strongly on unseen datasets without training on them. Extensive experiments across diverse multi-hop QA datasets demonstrate that BidLoop significantly surpasses existing state-of-the-art models.
Zhanpeng Guan, Zhao Zhang 0011, Yongjun Xu 0001
WSDM2
2025 Should We Use a Fixed Embedding Size? Customized Dimension Sizes for Knowledge Graph Embedding
abstract
Knowledge Graph Embedding (KGE) aims to project entities and relations into a low-dimensional space, so as to enable Knowledge Graphs (KGs) to be effectively used by downstream AI tasks. Most existing KGs (e.g. Wikidata) suffer from the data imbalance issue, i.e., the occurrence frequencies vary significantly among different entities. Current KGE models use a fixed embedding size, leading to overfitting for low-frequency entities and underfitting for high-frequency ones. A simple method is to manually set embedding sizes based on frequency, but this is not feasible due to the complexity and the large number of entities. To this end, we propose CustomizE, which customizes embedding sizes in a data-driven way, assigning larger sizes for high-frequency entities and smaller sizes for low-frequency ones. We use bilevel optimization for stable learning of representations and sizes. It is noteworthy that our framework is universal and flexible, which is suitable for various KGE models. Experiments on link prediction tasks show its superiority over state-of-the-art baselines.
Zhanpeng Guan, Zhao Zhang 0011, Yiqing Wu, Yongjun Xu 0001
COLING1
2025 AdaE: Knowledge Graph Embedding With Adaptive Embedding Sizes
abstract
Knowledge Graph Embedding (KGE) aims to learn dense embeddings as the representations for entities and relations in KGs. Indeed, the entities in existing KGs suffer from the data imbalance issue, i.e., there exists a substantial disparity in the occurrence frequencies among various entities. Existing KGE models pre-define a unified and fixed dimension size for all entity embeddings. However, embedding sizes of entities are highly desired for their frequencies, while a uniform embedding size may result in inadequate expression of entities, i.e., leading to overfitting for low-frequency entities and underfitting for high-frequency ones. A straight-forward idea is to set the embedding sizes for each entity before KGE training. However, manually selecting different embedding sizes is labor-intensive and time-consuming, which is difficult to achieve in real-world scenarios. To tackle this problem, we propose AdaE, which adaptively learns KG embeddings with different embedding sizes during training. In particular, AdaE is capable of selecting appropriate dimension sizes for each entity from a continuous integer space. To this end, we specially tailor bilevel optimization for the KGE task, which alternately learns representations and embedding sizes of entities. Moreover, it is worth noting that our framework is general and flexible, which is suitable for various existing KGE models. Extensive experiments demonstrate the effectiveness and compatibility of AdaE.
Zhanpeng Guan, Zhao Zhang 0011, Fuzhen Zhuang, Fei Wang 0014, Zhulin An, Yongjun Xu 0001
IEEE Trans. Knowl. Data Eng.1
2023 Weighted Knowledge Graph Embedding
abstract
Knowledge graph embedding (KGE) aims to project both entities and relations in a knowledge graph (KG) into low-dimensional vectors. Indeed, existing KGs suffer from the data imbalance issue, i.e., entities and relations conform to a long-tail distribution, only a small portion of entities and relations occur frequently, while the vast majority of entities and relations only have a few training samples. Existing KGE methods assign equal weights to each entity and relation during the training process. Under this setting, long-tail entities and relations are not fully trained during training, leading to unreliable representations. In this paper, we propose WeightE, which attends differentially to different entities and relations. Specifically, WeightE is able to endow lower weights to frequent entities and relations, and higher weights to infrequent ones. In such manner, WeightE is capable of increasing the weights of long-tail entities and relations, and learning better representations for them. In particular, WeightE tailors bilevel optimization for the KGE task, where the inner level aims to learn reliable entity and relation embeddings, and the outer level attempts to assign appropriate weights for each entity and relation. Moreover, it is worth noting that our technique of applying weights to different entities and relations is general and flexible, which can be applied to a number of existing KGE models. Finally, we extensively validate the superiority of WeightE against various state-of-the-art baselines.
Zhao Zhang 0011, Zhanpeng Guan, Fuzhen Zhuang, Zhulin An, Fei Wang 0014, Yongjun Xu 0001
SIGIR2
2005 Modeling and Analysis of Worm and Killer-Worm Propagation Using the Divide-and-Conquer Strategy
Dongyang Long, Changji Wang, Zhanpeng Guan
ICA3PP4