Changyue Wang 0001

dblp:347/2469-1 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0009-5460-5279ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 82% Trustworthy machine learning · 12% Question answering and dialogue systems · 3%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
1.722025
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System · SIGIR 2025
Parametric Retrieval Augmented Generation · SIGIR 2025
Natural language and speech › Language models and text generation
hallucination detection
1.012026
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026
Machine learning › Trustworthy machine learning
interpretability
1.012026
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026
Natural language and speech › Language models and text generation › large language model
large reasoning model
1.012026
Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models · AAAI 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Knowledge Editing through Chain-of-Thought · EMNLP 2025
Natural language and speech › Language models and text generation › knowledge editing
in-context knowledge editing
0.912025
Knowledge Editing through Chain-of-Thought · EMNLP 2025
Natural language and speech › Language models and text generation
knowledge editing
0.912025
Knowledge Editing through Chain-of-Thought · EMNLP 2025
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge injection
0.912025
Parametric Retrieval Augmented Generation · SIGIR 2025
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.912025
Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025
Information retrieval › document retrieval › domain-specific retrieval › legal information retrieval
legal case retrieval
0.912025
Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025
Information retrieval › reranking
neural re-ranking
0.912025
Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025
Information retrieval › retrieval models
neural retrieval
0.912025
Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025
Information retrieval › retrieval-augmented generation
parametric retrieval-augmented generation
0.912025
Parametric Retrieval Augmented Generation · SIGIR 2025
Information retrieval
retrieval models
0.912025
Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions · ACM Trans. Inf. Syst. 2025
Machine learning › Deep learning architectures and training
feedforward neural network
0.312025
Parametric Retrieval Augmented Generation · SIGIR 2025
Natural language and speech › Language models and text generation
large language model
0.312025
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System · SIGIR 2025
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.312025
Knowledge Editing through Chain-of-Thought · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

fine-tuning · 2.6few-shot in-context learning · 2.6automated evaluation framework · 2.6document parameterization · 1.7semantic alignment · 1.0inter-sample consistency · 1.0entropy-based uncertainty · 1.0unsupervised pre-training · 0.9iterative refinement · 0.9contrastive learning · 0.9chain-of-thought prompting · 0.9
YearPublicationVenuePosition
2026 Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning Models
abstract
Large Reasoning Models (LRMs) extend large language models with explicit, multi-step reasoning traces to enhance transparency and performance on complex tasks. However, these reasoning traces can be redundant or logically inconsistent, becoming a new and hard-to-detect source of hallucination. Existing hallucination detection methods focus primarily on answer-level uncertainty and often fail to detect hallucinations or logical inconsistencies arising from the model’s reasoning trace. This oversight is particularly problematic for LRMs, where the explicit thinking trace is not only an important support to the model's decision-making process but also a key source of potential hallucination. To this end, we propose RACE (Reasoning and Answer Consistency Evaluation), a novel framework specifically tailored for hallucination detection in LRMs. RACE operates by extracting essential reasoning steps and computing four diagnostic signals: inter-sample consistency of reasoning traces, entropy-based answer uncertainty, semantic alignment between reasoning and answers, and internal coherence of reasoning. This joint analysis enables fine-grained hallucination detection even when the final answer appears correct. Experiments across datasets and different LLMs demonstrate that RACE outperforms existing hallucination detection baselines, offering a robust and generalizable solution for evaluating LRMs.
Changyue Wang 0001, Weihang Su, Qingyao Ai, Yiqun Liu 0001
AAAI1
2025 Knowledge Editing through Chain-of-Thought
abstract
Knowledge Editing is a technique that updates large language models (LLMs) with new information to maintain their world knowledge.This approach avoids the need to rebuild the model from scratch, thereby addressing the high costs associated with frequent retraining.Among these, the in-context editing paradigm stands out for its effectiveness in integrating new knowledge while preserving the model's original capabilities.Despite its potential, existing in-context knowledge editing methods are often task-specific, focusing primarily on multi-hop QA tasks using structured knowledge triples.Moreover, their reliance on fewshot prompting for task decomposition makes them unstable and less effective in generalizing across diverse tasks.In response to these limitations, we propose EditCoT, a novel knowledge editing framework that flexibly and efficiently updates LLMs across various tasks without retraining.EditCoT works by generating a chain-of-thought (CoT) for a given input and then iteratively refining this CoT process using a CoT editor based on updated knowledge.We evaluate EditCoT across a diverse range of benchmarks, covering multiple languages and tasks.The results demonstrate that our approach achieves state-of-the-art performance while offering superior generalization, effectiveness, and stability compared to existing methods, marking a significant advancement in the field of knowledge updating 1 .
Changyue Wang 0001, Weihang Su, Qingyao Ai, Yichen Tang 0001, Yiqun Liu 0001
EMNLP1
2025 Parametric Retrieval Augmented Generation
abstract
Retrieval-augmented generation (RAG) has emerged as a promising solution to enhance the reliability of large language models (LLMs) with external knowledge. Existing RAG methods share a common strategy for knowledge injection: they place the retrieved documents into the input context of the LLM, which we refer to as the in-context knowledge injection method. While this approach is simple and often effective, it has inherent limitations. Firstly, increasing the context length and number of relevant documents can lead to higher computational overhead and degraded performance, especially in complex reasoning tasks. More importantly, in-context knowledge injection operates primarily at the input level, but LLMs store their internal knowledge in their parameters. This gap fundamentally limits the capacity of in-context methods. To this end, we introduce Parametric RAG, a new RAG paradigm that integrates external knowledge directly into the feed-forward networks of an LLM through document parameterization. This approach not only reduces online computational costs by shortening the input context length, but also deepens the integration of external knowledge by enabling LLMs to utilize it in the same way as internal parametric knowledge. Experimental results demonstrate that Parametric RAG substantially enhances the effectiveness and efficiency of knowledge augmentation in LLMs. Also, it can be combined with in-context RAG methods to achieve even better performance. We have open-sourced all the code, data, and models in the following GitHub link: https://github.com/oneal2000/PRAG
Weihang Su, Yichen Tang 0001, Qingyao Ai, Junxi Yan, Changyue Wang 0001, Hongning Wang, Ziyi Ye, Yujia Zhou 0002, Yiqun Liu 0001
SIGIR5
2025 JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
abstract
This paper introduces JuDGE (Judgment Document Generation Evaluation), a novel benchmark for evaluating the performance of judgment document generation in the Chinese legal system. We define the task as generating a complete legal judgment document from the given factual description of the case. To facilitate this benchmark, we construct a comprehensive dataset consisting of factual descriptions from real legal cases, paired with their corresponding full judgment documents, which serve as the ground truth for evaluating the quality of generated documents. This dataset is further augmented by two external legal corpora that provide additional legal knowledge for the task: one comprising statutes and regulations, and the other consisting of a large collection of past judgment documents. In collaboration with legal professionals, we establish a comprehensive automated evaluation framework to assess the quality of generated judgment documents across various dimensions. We evaluate various baseline approaches, including few-shot in-context learning, fine-tuning, and a multi-source retrieval-augmented generation (RAG) approach, using both general and legal-domain LLMs. The experimental results demonstrate that, while RAG approaches can effectively improve performance in this task, there is still substantial room for further improvement. All the codes and datasets are available at: https://github.com/oneal2000/JuDGE
Weihang Su, Baoqing Yue, Qingyao Ai, Yiran Hu, Changyue Wang 0001, Yueyue Wu, Yiqun Liu 0001
SIGIR6
2025 Pre-training for Legal Case Retrieval Based on Inter-Case Distinctions
abstract
Legal case retrieval aims to help legal workers find relevant cases related to their cases at hand, which is important for the guarantee of fairness and justice in legal judgments. While recent advances in neural retrieval methods have significantly improved the performance of open-domain retrieval tasks (e.g., Web search), their advantages haven’t been observed in legal case retrieval due to their thirst for annotated data. As annotating large-scale training data in legal domains is prohibitive due to the need for domain expertise, traditional search techniques based on lexical matching such as TF-IDF, BM25, and Query Likelihood are still prevalent in legal case retrieval systems. While previous studies have designed several pre-training methods for IR models in open-domain tasks, these methods are usually suboptimal in legal case retrieval because they cannot understand and capture the key knowledge and data structures in the legal corpus. To this end, we propose a novel pre-training framework named Caseformer that enables the pre-trained models to learn legal knowledge and domain-specific relevance-matching patterns in legal case retrieval without any human-labeled data. This framework is designed to support both dense retrieval models and neural re-ranking models. Through three unsupervised learning tasks, Caseformer is able to capture the special language, document structure, and relevance-matching patterns of legal case documents, making it a strong backbone for downstream legal case retrieval tasks. Experimental results show that our model has achieved state-of-the-art performance in both zero-shot and fine-tuning settings. Also, experiments on both Chinese and English legal datasets demonstrate that the effectiveness of Caseformer is language-independent in legal case retrieval.
Weihang Su, Qingyao Ai, Yueyue Wu, Anzhe Xie, Changyue Wang 0001, Haitao Li 0006, Zhijing Wu 0001, Yiqun Liu 0001, Min Zhang 0006
ACM Trans. Inf. Syst.5