Canzhi Chen

dblp:410/8613 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-3888-5719ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Recommender systems · 42% Information retrieval · 39% Web and social media mining · 20%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining › citation network analysis
citation prediction
1.012026
What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026
Recommender systems › content recommendation
citation recommendation
1.012026
What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026
Information retrieval
retrieval-augmented generation
1.012026
What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026
Information retrieval
retrieval models
1.012026
What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026
Recommender systems › debiased recommendation › selection bias
exposure bias
0.912025
Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning · SIGIR 2025
Natural language and speech › Language models and text generation
large language model
0.312026
What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026
Recommender systems › beyond-accuracy recommendation
long-tail recommendation
0.312025
Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0fine-tuning · 2.0contrastive learning · 2.0inverse propensity scoring · 0.9graph contrastive learning · 0.9
YearPublicationVenuePosition
2026 What Should I Cite? A RAG Benchmark for Academic Citation Prediction
abstract
With the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation prediction aims to automatically suggest appropriate references, helping scholars navigate the expanding scientific literature. Here we present CiteRAG, the first comprehensive retrieval-augmented generation (RAG)-integrated benchmark for evaluating large language models on academic citation prediction, featuring a multi-level retrieval strategy, specialized retrievers, and generators. Our benchmark makes four core contributions: (1) We establish two instances of the citation prediction task with different granularity. Task 1 focuses on coarse-grained list-specific citation prediction, while Task 2 targets fine-grained position-specific citation prediction. To enhance these two tasks, we build a dataset containing 7,267 instances for Task 1 and 8,541 instances for Task 2, enabling comprehensive evaluation of both retrieval and generation. (2) We construct a three-level large-scale corpus with 554k papers spanning many major subfields, using an incremental pipeline. (3) We propose a multi-level hybrid RAG approach to citation prediction, fine-tuning embedding models with contrastive learning to capture complex citation relationships, paired with specialized generation models. (4) We conduct extensive experiments across state-of-the-art language models, including closed-source APIs, open-source models, and our fine-tuned generators, demonstrating the effectiveness of our framework. Our open-source toolkit enables reproducible evaluation and focuses on academic literature, providing the first comprehensive evaluation framework for citation prediction and serving as a methodological template for other scientific domains. Our source code and data are released at https://github.com/LQgdwind/CiteRAG.
Leqi Zheng, Jiajun Zhang 0012, Canzhi Chen, Chaokun Wang, Hongwei Li 0032, Yuying Li 0006, Yaoxin Mao, Shannan Yan, Zixin Song, Zhiyuan Feng, Zhaolu Kang, Zirong Chen, Hang Zhang 0032, Qiang Liu 0006, Liang Wang 0001, Ziyang Liu 0004
WWW3
2025 Emergency Evacuation Map Guided Navigation via Topological Alignment and VLM Reasoning
Canzhi Chen, Weiqi Huang, Huijun Di, Wei Liang 0008
PRCV (6)1
2025 Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning
abstract
Recent advances in graph contrastive learning (GCL) have significantly enhanced recommendation systems. However, most existing approaches predominantly focus on optimizing training data fit while overlooking exposure bias, a critical issue that can substantially impact recommendation effectiveness. Drawing inspiration from sociological theories of human interaction patterns-specifically how individuals balance self-presentation and self-hiding behaviors in social contexts-this paper proposes BPH4Rec, a novel Balancing self-Presentation and self-Hiding approach for exposure-aware Recommendation based on GCL. Within the GCL framework, BPH4Rec introduces two complementary mechanisms: (1) a self-hiding mechanism that modifies the adjacency matrix of contrastive views through custom inverse propensity scoring (IPS), effectively addressing exposure bias, and (2) a self-presentation mechanism that incorporates densification factors during matrix reconstruction to mitigate sparsity-induced biases. Through extensive evaluation on six public benchmark datasets, BPH4Rec demonstrates substantial improvements over state-of-the-art baselines, particularly in promoting long-tail item discovery while maintaining recommendation accuracy.
Leqi Zheng, Chaokun Wang, Ziyang Liu 0004, Canzhi Chen, Cheng Wu 0004, Hongwei Li 0032
SIGIR4