VLDB 2026 Research / reviewers in the wild / expert
Canzhi Chen
dblp:410/8613
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-3888-5719ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Recommender systems · 42% Information retrieval · 39% Web and social media mining · 20% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and social media mining › citation network analysis
citation prediction |
1.0 | 1 | 2026 | What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026 |
Recommender systems › content recommendation
citation recommendation |
1.0 | 1 | 2026 | What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026 |
Information retrieval
retrieval-augmented generation |
1.0 | 1 | 2026 | What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026 |
Information retrieval
retrieval models |
1.0 | 1 | 2026 | What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026 |
Recommender systems › debiased recommendation › selection bias
exposure bias |
0.9 | 1 | 2025 | Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning · SIGIR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2026 | What Should I Cite? A RAG Benchmark for Academic Citation Prediction · WWW 2026 |
Recommender systems › beyond-accuracy recommendation
long-tail recommendation |
0.3 | 1 | 2025 | Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 2.0fine-tuning · 2.0contrastive learning · 2.0inverse propensity scoring · 0.9graph contrastive learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Should I Cite? A RAG Benchmark for Academic Citation PredictionabstractWith the rapid growth of Web-based academic publications, more and more papers are being published annually, making it increasingly difficult to find relevant prior work. Citation prediction aims to automatically suggest appropriate references, helping scholars navigate the expanding scientific literature. Here we present CiteRAG, the first comprehensive retrieval-augmented generation (RAG)-integrated benchmark for evaluating large language models on academic citation prediction, featuring a multi-level retrieval strategy, specialized retrievers, and generators. Our benchmark makes four core contributions: (1) We establish two instances of the citation prediction task with different granularity. Task 1 focuses on coarse-grained list-specific citation prediction, while Task 2 targets fine-grained position-specific citation prediction. To enhance these two tasks, we build a dataset containing 7,267 instances for Task 1 and 8,541 instances for Task 2, enabling comprehensive evaluation of both retrieval and generation. (2) We construct a three-level large-scale corpus with 554k papers spanning many major subfields, using an incremental pipeline. (3) We propose a multi-level hybrid RAG approach to citation prediction, fine-tuning embedding models with contrastive learning to capture complex citation relationships, paired with specialized generation models. (4) We conduct extensive experiments across state-of-the-art language models, including closed-source APIs, open-source models, and our fine-tuned generators, demonstrating the effectiveness of our framework. Our open-source toolkit enables reproducible evaluation and focuses on academic literature, providing the first comprehensive evaluation framework for citation prediction and serving as a methodological template for other scientific domains. Our source code and data are released at https://github.com/LQgdwind/CiteRAG. Leqi Zheng, Jiajun Zhang 0012, Canzhi Chen, Chaokun Wang, Hongwei Li 0032, Yuying Li 0006, Yaoxin Mao, Shannan Yan, Zixin Song, Zhiyuan Feng, Zhaolu Kang, Zirong Chen, Hang Zhang 0032, Qiang Liu 0006, Liang Wang 0001, Ziyang Liu 0004 |
WWW | 3 |
| 2025 | Emergency Evacuation Map Guided Navigation via Topological Alignment and VLM Reasoning
Canzhi Chen, Weiqi Huang, Huijun Di, Wei Liang 0008 |
PRCV (6) | 1 |
| 2025 | Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive LearningabstractRecent advances in graph contrastive learning (GCL) have significantly enhanced recommendation systems. However, most existing approaches predominantly focus on optimizing training data fit while overlooking exposure bias, a critical issue that can substantially impact recommendation effectiveness. Drawing inspiration from sociological theories of human interaction patterns-specifically how individuals balance self-presentation and self-hiding behaviors in social contexts-this paper proposes BPH4Rec, a novel Balancing self-Presentation and self-Hiding approach for exposure-aware Recommendation based on GCL. Within the GCL framework, BPH4Rec introduces two complementary mechanisms: (1) a self-hiding mechanism that modifies the adjacency matrix of contrastive views through custom inverse propensity scoring (IPS), effectively addressing exposure bias, and (2) a self-presentation mechanism that incorporates densification factors during matrix reconstruction to mitigate sparsity-induced biases. Through extensive evaluation on six public benchmark datasets, BPH4Rec demonstrates substantial improvements over state-of-the-art baselines, particularly in promoting long-tail item discovery while maintaining recommendation accuracy. Leqi Zheng, Chaokun Wang, Ziyang Liu 0004, Canzhi Chen, Cheng Wu 0004, Hongwei Li 0032 |
SIGIR | 4 |