VLDB 2026 Research / reviewers in the wild / expert
Changyue Wang 0001
dblp:347/2469-1
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0009-5460-5279ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 82% Trustworthy machine learning · 12% Question answering and dialogue systems · 3% | |
| Databases, data mining, and information retrieval
3 papers |
Information retrieval · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval-augmented generation |
1.7 | 2 | 2025 | JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System · SIGIR 2025 Parametric Retrieval Augmented Generation · SIGIR 2025 |
Natural language and speech › Language models and text generation
hallucination detection |
1.0 | 1 | 2026 | Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model
large reasoning model |
1.0 | 1 | 2026 | Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Knowledge Editing through Chain-of-Thought · EMNLP 2025 |
Natural language and speech › Language models and text generation › knowledge editing
in-context knowledge editing |
0.9 | 1 | 2025 | Knowledge Editing through Chain-of-Thought · EMNLP 2025 |
Natural language and speech › Language models and text generation
knowledge editing |
0.9 | 1 | 2025 | Knowledge Editing through Chain-of-Thought · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge injection |
0.9 | 1 | 2025 | Parametric Retrieval Augmented Generation · SIGIR 2025 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.9 | 1 | 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025 |
Information retrieval › document retrieval › domain-specific retrieval › legal information retrieval
legal case retrieval |
0.9 | 1 | 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025 |
Information retrieval › reranking
neural re-ranking |
0.9 | 1 | 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025 |
Information retrieval › retrieval models
neural retrieval |
0.9 | 1 | 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025 |
Information retrieval › retrieval-augmented generation
parametric retrieval-augmented generation |
0.9 | 1 | 2025 | Parametric Retrieval Augmented Generation · SIGIR 2025 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025 |
Machine learning › Deep learning architectures and training
feedforward neural network |
0.3 | 1 | 2025 | Parametric Retrieval Augmented Generation · SIGIR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System · SIGIR 2025 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.3 | 1 | 2025 | Knowledge Editing through Chain-of-Thought · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 2.6few-shot in-context learning · 2.6automated evaluation framework · 2.6document parameterization · 1.7semantic alignment · 1.0inter-sample consistency · 1.0entropy-based uncertainty · 1.0unsupervised pre-training · 0.9iterative refinement · 0.9contrastive learning · 0.9chain-of-thought prompting · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsabstractLarge Reasoning Models (LRMs) extend large language models with explicit, multi-step reasoning traces to enhance transparency and performance on complex tasks. However, these reasoning traces can be redundant or logically inconsistent, becoming a new and hard-to-detect source of hallucination. Existing hallucination detection methods focus primarily on answer-level uncertainty and often fail to detect hallucinations or logical inconsistencies arising from the model’s reasoning trace. This oversight is particularly problematic for LRMs, where the explicit thinking trace is not only an important support to the model's decision-making process but also a key source of potential hallucination. To this end, we propose RACE (Reasoning and Answer Consistency Evaluation), a novel framework specifically tailored for hallucination detection in LRMs. RACE operates by extracting essential reasoning steps and computing four diagnostic signals: inter-sample consistency of reasoning traces, entropy-based answer uncertainty, semantic alignment between reasoning and answers, and internal coherence of reasoning. This joint analysis enables fine-grained hallucination detection even when the final answer appears correct. Experiments across datasets and different LLMs demonstrate that RACE outperforms existing hallucination detection baselines, offering a robust and generalizable solution for evaluating LRMs. Changyue Wang 0001, Weihang Su, Qingyao Ai, Yiqun Liu 0001 |
AAAI | 1 |
| 2025 | Knowledge Editing through Chain-of-ThoughtabstractKnowledge Editing is a technique that updates large language models (LLMs) with new information to maintain their world knowledge.This approach avoids the need to rebuild the model from scratch, thereby addressing the high costs associated with frequent retraining.Among these, the in-context editing paradigm stands out for its effectiveness in integrating new knowledge while preserving the model's original capabilities.Despite its potential, existing in-context knowledge editing methods are often task-specific, focusing primarily on multi-hop QA tasks using structured knowledge triples.Moreover, their reliance on fewshot prompting for task decomposition makes them unstable and less effective in generalizing across diverse tasks.In response to these limitations, we propose EditCoT, a novel knowledge editing framework that flexibly and efficiently updates LLMs across various tasks without retraining.EditCoT works by generating a chain-of-thought (CoT) for a given input and then iteratively refining this CoT process using a CoT editor based on updated knowledge.We evaluate EditCoT across a diverse range of benchmarks, covering multiple languages and tasks.The results demonstrate that our approach achieves state-of-the-art performance while offering superior generalization, effectiveness, and stability compared to existing methods, marking a significant advancement in the field of knowledge updating 1 . Changyue Wang 0001, Weihang Su, Qingyao Ai, Yichen Tang 0001, Yiqun Liu 0001 |
EMNLP | 1 |
| 2025 | Parametric Retrieval Augmented GenerationabstractRetrieval-augmented generation (RAG) has emerged as a promising solution to enhance the reliability of large language models (LLMs) with external knowledge. Existing RAG methods share a common strategy for knowledge injection: they place the retrieved documents into the input context of the LLM, which we refer to as the in-context knowledge injection method. While this approach is simple and often effective, it has inherent limitations. Firstly, increasing the context length and number of relevant documents can lead to higher computational overhead and degraded performance, especially in complex reasoning tasks. More importantly, in-context knowledge injection operates primarily at the input level, but LLMs store their internal knowledge in their parameters. This gap fundamentally limits the capacity of in-context methods. To this end, we introduce Parametric RAG, a new RAG paradigm that integrates external knowledge directly into the feed-forward networks of an LLM through document parameterization. This approach not only reduces online computational costs by shortening the input context length, but also deepens the integration of external knowledge by enabling LLMs to utilize it in the same way as internal parametric knowledge. Experimental results demonstrate that Parametric RAG substantially enhances the effectiveness and efficiency of knowledge augmentation in LLMs. Also, it can be combined with in-context RAG methods to achieve even better performance. We have open-sourced all the code, data, and models in the following GitHub link: https://github.com/oneal2000/PRAG Weihang Su, Yichen Tang 0001, Qingyao Ai, Junxi Yan, Changyue Wang 0001, Hongning Wang, Ziyi Ye, Yujia Zhou 0002, Yiqun Liu 0001 |
SIGIR | 5 |
| 2025 | JuDGE: Benchmarking Judgment Document Generation for Chinese Legal SystemabstractThis paper introduces JuDGE (Judgment Document Generation Evaluation), a novel benchmark for evaluating the performance of judgment document generation in the Chinese legal system. We define the task as generating a complete legal judgment document from the given factual description of the case. To facilitate this benchmark, we construct a comprehensive dataset consisting of factual descriptions from real legal cases, paired with their corresponding full judgment documents, which serve as the ground truth for evaluating the quality of generated documents. This dataset is further augmented by two external legal corpora that provide additional legal knowledge for the task: one comprising statutes and regulations, and the other consisting of a large collection of past judgment documents. In collaboration with legal professionals, we establish a comprehensive automated evaluation framework to assess the quality of generated judgment documents across various dimensions. We evaluate various baseline approaches, including few-shot in-context learning, fine-tuning, and a multi-source retrieval-augmented generation (RAG) approach, using both general and legal-domain LLMs. The experimental results demonstrate that, while RAG approaches can effectively improve performance in this task, there is still substantial room for further improvement. All the codes and datasets are available at: https://github.com/oneal2000/JuDGE Weihang Su, Baoqing Yue, Qingyao Ai, Yiran Hu, Changyue Wang 0001, Yueyue Wu, Yiqun Liu 0001 |
SIGIR | 6 |
| 2025 | Pre-training for Legal Case Retrieval Based on Inter-Case DistinctionsabstractLegal case retrieval aims to help legal workers find relevant cases related to their cases at hand, which is important for the guarantee of fairness and justice in legal judgments. While recent advances in neural retrieval methods have significantly improved the performance of open-domain retrieval tasks (e.g., Web search), their advantages haven’t been observed in legal case retrieval due to their thirst for annotated data. As annotating large-scale training data in legal domains is prohibitive due to the need for domain expertise, traditional search techniques based on lexical matching such as TF-IDF, BM25, and Query Likelihood are still prevalent in legal case retrieval systems. While previous studies have designed several pre-training methods for IR models in open-domain tasks, these methods are usually suboptimal in legal case retrieval because they cannot understand and capture the key knowledge and data structures in the legal corpus. To this end, we propose a novel pre-training framework named Caseformer that enables the pre-trained models to learn legal knowledge and domain-specific relevance-matching patterns in legal case retrieval without any human-labeled data. This framework is designed to support both dense retrieval models and neural re-ranking models. Through three unsupervised learning tasks, Caseformer is able to capture the special language, document structure, and relevance-matching patterns of legal case documents, making it a strong backbone for downstream legal case retrieval tasks. Experimental results show that our model has achieved state-of-the-art performance in both zero-shot and fine-tuning settings. Also, experiments on both Chinese and English legal datasets demonstrate that the effectiveness of Caseformer is language-independent in legal case retrieval. Weihang Su, Qingyao Ai, Yueyue Wu, Anzhe Xie, Changyue Wang 0001, Haitao Li 0006, Zhijing Wu 0001, Yiqun Liu 0001, Min Zhang 0006 |
ACM Trans. Inf. Syst. | 5 |