VLDB 2026 Research / reviewers in the wild / expert
Jianhui Yang 0001
dblp:35/10764-1
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0004-3547-0472ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Artificial intelligence
1 paper |
Reinforcement learning · 67% Language models and text generation · 33% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance · WWW 2026 |
Machine learning › Reinforcement learning
guided reinforcement learning |
1.0 | 1 | 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance · WWW 2026 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.0 | 1 | 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance · WWW 2026 |
Information retrieval › e-commerce search
e-commerce relevance |
1.0 | 1 | 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance · WWW 2026 |
Information retrieval › ranking
search relevance |
1.0 | 1 | 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance · WWW 2026 |
Information retrieval › evaluation
benchmark |
0.9 | 1 | 2025 | LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation · SIGIR 2025 |
Information retrieval › interactive information retrieval › conversational information seeking
conversational search |
0.9 | 1 | 2025 | LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation · SIGIR 2025 |
Information retrieval
evaluation |
0.9 | 1 | 2025 | LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation · SIGIR 2025 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 2.0reward shaping · 2.0adaptive guided replay · 2.0GRPO · 2.0DPO · 2.0large language model · 0.9LLM-as-a-judge · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search RelevanceabstractQuery-product relevance prediction is fundamental to e-commerce search and has become even more critical in the era of AI-powered shopping, where semantic understanding and complex reasoning directly shape the user experience and business conversion. Large Language Models (LLMs) enable generative, reasoning-based approaches, typically aligned via supervised fine-tuning (SFT) or preference optimization methods like Direct Preference Optimization (DPO). However, the increasing complexity of business rules and user queries exposes the inability of existing methods to endow models with robust reasoning capacity for long-tail and challenging cases. Efforts to address this via reinforcement learning strategies like Group Relative Policy Optimization (GRPO) often suffer from sparse terminal rewards, offering insufficient guidance for multi-step reasoning and slowing convergence. To address these challenges, we propose TaoSR-AGRL, an Adaptive Guided Reinforcement Learning framework for LLM-based relevance prediction in Taobao Search Relevance. TaoSR-AGRL introduces two key innovations: (1) Rule-aware Reward Shaping, which decomposes the final relevance judgment into dense, structured rewards aligned with domain-specific relevance criteria; and (2) Adaptive Guided Replay, which identifies low-accuracy rollouts during training and injects targeted ground-truth guidance to steer the policy away from stagnant, rule-violating reasoning patterns toward compliant trajectories. TaoSR-AGRL was evaluated on large-scale real-world datasets and through online side-by-side human evaluations on Taobao Search. It consistently outperforms DPO and standard GRPO baselines in offline experiments, improving relevance accuracy, rule adherence, and training stability. The model trained with TaoSR-AGRL has been successfully deployed in the main search scenario on Taobao, serving hundreds of millions of users. Jianhui Yang 0001, Pengkun Jiao, Chenhe Dong, Zerui Huang, Shaowei Yao, Xiaojiang Zhou, Dan Ou, Haihong Tang |
WWW | 1 |
| 2025 | HCLeK: Hierarchical Compression of Legal Knowledge for Retrieval-Augmented GenerationabstractPrompt compression for Retrieval-Augmented Generation (RAG) often fails by treating all retrieved information uniformly. This undifferentiated approach neglects the critical distinction between foundational core knowledge and illustrative practical knowledge, a failure especially damaging in hierarchical domains like law where essential principles can be discarded for redundant details, diminishing information gain. Jianhui Yang 0001, Huanghai Liu, Mingruo Yuan, Yiran Hu, Yun Liu 0033, Weixing Shen, Ben Kao |
CIKM | 1 |
| 2025 | LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation ConversationabstractRetrieval-augmented generation (RAG) has proven highly effective in improving large language models (LLMs) across various domains. However, there is no benchmark specifically designed to assess the effectiveness of RAG in the legal domain, which restricts progress in this area. To fill this gap, we propose LexRAG, the first benchmark to evaluate RAG systems for multi-turn legal consultations. LexRAG consists of 1,013 multi-turn dialogue samples and 17,228 candidate legal articles. Each sample is annotated by legal experts and consists of five rounds of progressive questioning. LexRAG includes two key tasks: (1) Conversational knowledge retrieval, requiring accurate retrieval of relevant legal articles based on multi-turn context. (2) Response generation, focusing on producing legally sound answers. To ensure reliable reproducibility, we develop LexiT, a legal RAG toolkit that provides a comprehensive implementation of RAG system components tailored for the legal domain. Additionally, we introduce an LLM-as-a-judge evaluation pipeline to enable detailed and effective assessment. Through experimental analysis of various LLMs and retrieval methods, we reveal the key limitations of existing RAG systems in handling legal consultation conversations. LexRAG establishes a new benchmark for the practical application of RAG systems in the legal domain, with its code and data available at https://github.com/CSHaitao/LexRAG. Haitao Li 0006, Yiran Hu, Qingyao Ai, Jianhui Yang 0001, Yueyue Wu, Zeyang Liu 0004, Yiqun Liu 0001 |
SIGIR | 7 |