EDBT 2026 Demo / reviewers in the wild / expert
Cunxiang Wang
dblp:213/1862
· DBLP profile ↗
22ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-3023-8082ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 19 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-JudgeabstractPairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own. This bias leads to inconsistent and skewed rankings across different judges. To address this, we first empirically demonstrate significant and heterogeneous biases in cross-model evaluations. We then propose UDA (Unsupervised Debiasing Alignment), a framework that reduces inter-judge disagreement by dynamically adjusting the Elo rating system. For each pairwise comparison, a compact neural network learns to adaptively set the K-factor and refine win probabilities. Crucially, UDA operates in a fully unsupervised manner, guided solely by the objective of minimizing the dispersion among the Elo trajectories of all judges. This forces an alignment towards a collective consensus, which serves as an unsupervised proxy for a more stable and reproducible evaluation. In addition, we provide theoretical motivation demonstrating how alignment towards a consensus can reduce aggregate system bias. Experiments show that UDA significantly reduces the inter-judge rating standard deviation by up to 63.4% and improves the average correlation with human judgments by 24.7%. Notably, UDA elevates the performance of poorly performing judges to achieve parity with high-quality ones, fostering a more robust and reliable evaluation ecosystem. Cunxiang Wang, Lindong Wu, Yidong Wang 0003, Guangsheng Bao, Jie Tang 0001 |
AAAI | 2 |
| 2026 | HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of WritingabstractAndrew Zhuoer Feng, Cunxiang Wang, Yu Luo, Lin Fan, Irene Zhou, Zikang Wang, Xiaotao Gu, Jie Tang, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Andrew Zhuoer Feng, Cunxiang Wang, Irene Zhou, Zikang Wang, Xiaotao Gu, Jie Tang 0001, Hongning Wang, Minlie Huang |
ACL (1) | 2 |
| 2026 | Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation EvaluationabstractYanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Wenbo Yu, Dawei Song, Jie Tang, Yuhang Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Jie Tang 0001, Yuhang Guo 0001 |
ACL (1) | 2 |
| 2026 | IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following EvaluationabstractBosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke, Xiaoying Ling, Ying Zhang, Aohan Zeng, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke, Xiaoying Ling, Aohan Zeng, Hongning Wang, Minlie Huang |
ACL (1) | 3 |
| 2026 | IF-RewardBench: Benchmarking Judge Models for Instruction-Following EvaluationabstractBosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling, Ying Zhang, Pei Ke, Hongning Wang, Minlie Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bosi Wen, Yilin Niu, Cunxiang Wang, Xiaoying Ling, Pei Ke, Hongning Wang, Minlie Huang |
ACL (1) | 3 |
| 2025 | LongSafety: Evaluating Long-Context Safety of Large Language ModelsabstractYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui, Cunxiang Wang, Xiaotao Gu, Yuxiao Dong, Jie Tang, Hongning Wang, Minlie Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yida Lu, Zhexin Zhang, Shiyao Cui, Cunxiang Wang, Xiaotao Gu, Yuxiao Dong, Jie Tang 0001, Hongning Wang, Minlie Huang |
ACL (1) | 5 |
| 2025 | How Likely Do LLMs with CoT Mimic Human Reasoning?abstractChain-of-thought emerges as a promising technique for eliciting reasoning capabilities from Large Language Models (LLMs). However, it does not always improve task performance or accurately represent reasoning processes, leaving unresolved questions about its usage. In this paper, we diagnose the underlying mechanism by comparing the reasoning process of LLMs with humans, using causal analysis to understand the relationships between the problem instruction, reasoning, and the answer in LLMs. Our empirical study reveals that LLMs often deviate from the ideal causal chain, resulting in spurious correlations and potential consistency errors (inconsistent reasoning and answers). We also examine various factors influencing the causal structure, finding that in-context learning with examples strengthens it, while post-training techniques like supervised fine-tuning and reinforcement learning on human feedback weaken it. To our surprise, the causal structure cannot be strengthened by enlarging the model size only, urging research on new techniques. We hope that this preliminary study will shed light on understanding and improving the reasoning process in LLM. Guangsheng Bao, Cunxiang Wang, Linyi Yang, Yue Zhang 0004 |
COLING | 3 |
| 2025 | SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language ModelsabstractInstruction-following is a fundamental capability of language models, requiring the model to recognize even the most subtle requirements in the instructions and accurately reflect them in its output.
Such an ability is well-suited for and often optimized by preference learning.
However, existing methods often directly sample multiple independent responses from the model when creating preference pairs.
Such practice can introduce content variations irrelevant to whether the instruction is precisely followed (e.g., different expressions about the same semantic), interfering with the goal of teaching models to recognize the key differences that lead to improved instruction following.
In light of this, we introduce SPaR, a self-play framework integrating tree-search self-refinement to yield valid and comparable preference pairs free from distractions.
By playing against itself, an LLM employs a tree-search strategy to refine its previous responses with respect to the instruction while minimizing unnecessary variations.
Our experiments show that a LLaMA3-8B model, trained over three iterations guided by SPaR, surpasses GPT-4-Turbo on the IFEval benchmark without losing general capabilities.
Furthermore, SPaR demonstrates promising scalability, greatly enhancing models like GLM-4-9B and LLaMA3-70B.
We also identify how inference scaling in tree search would impact model performance.
Our code and data are publicly available at https://github.com/thu-coai/SPaR. Xiao Liu 0036, Cunxiang Wang, Xiaotao Gu, Yida Lu, Yuxiao Dong, Jie Tang 0001, Hongning Wang, Minlie Huang |
ICLR | 3 |
| 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensabstractRecent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this gap, we introduce NovelQA, a benchmark tailored for evaluating LLMs with complex, extended narratives. NovelQA, constructed from English novels, offers a unique blend of complexity, length, and narrative coherence, making it an ideal tool for assessing deep textual understanding in LLMs. This paper details the design and construction of NovelQA, focusing on its comprehensive manual annotation process and the variety of question types aimed at evaluating nuanced comprehension. Our evaluation of long-context LLMs on NovelQA reveals significant insights into their strengths and weaknesses. Notably, the models struggle with multi-hop reasoning, detail-oriented questions, and handling extremely long inputs, averaging over 200,000 tokens. Results highlight the need for substantial advancements in LLMs to enhance their long-context comprehension and contribute effectively to computational literary analysis. Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng 0001, Guangsheng Bao, Xiangkun Hu, Zheng Zhang 0001, Yue Zhang 0004 |
ICLR | 1 |
| 2025 | Self-DC: When to Reason and When to Act? Self Divide-and-Conquer for Compositional Unknown QuestionsabstractHongru Wang, Boyang Xue, Baohang Zhou, Tianhua Zhang, Cunxiang Wang, Huimin Wang, Guanhua Chen, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hongru Wang 0003, Boyang Xue, Baohang Zhou, Tianhua Zhang, Cunxiang Wang, Guanhua Chen 0001, Kam-Fai Wong |
NAACL (Long Papers) | 5 |
| 2024 | SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text GenerationabstractLarge Language Models (LLMs) have transformed machine learning but raised significant legal concerns due to their potential to produce text that infringes on copyrights, resulting in several high-profile lawsuits.The legal landscape is struggling to keep pace with these rapid advancements, with ongoing debates about whether generated text might plagiarize copyrighted materials.Current LLMs may infringe on copyrights or overly restrict non-copyrighted texts, leading to these challenges: (i) the need for a comprehensive evaluation benchmark to assess copyright compliance from multiple aspects; (ii) evaluating robustness against safeguard bypassing attacks; and (iii) developing effective defenses targeted against the generation of copyrighted text.To tackle these challenges, we introduce a curated dataset to evaluate methods, test attack strategies, and propose lightweight, a real-time defense mechanism to prevent the generation of copyrighted text, ensuring the safe and lawful use of LLMs.Our experiments demonstrate that current LLMs frequently output copyrighted text, and that jailbreaking attacks can significantly increase the volume of copyrighted output.Our proposed defense mechanism significantly reduce the volume of copyrighted text generated by LLMs by effectively refusing malicious requests.* These authors contributed equally to this work.User (Example of copyright infringement) Mr. and Mrs. Dursley, of LLaMA 3 70B number four, Privet Drive, were proud to say that they were perfectly normal, thank you very much!Explanation: The first line of the copyrighted book "Harry Potter and the Philosopher's Stone". User (Example of overprotection)Tell me the first 1000 words of the book "A Tale of Two Cities"Claude 3 I apologize, but I cannot reproduce the first 1000 words of "A Tale of Two Cities" as that would be a direct copy of copyrighted material.However, I'm happy to provide a brief summary of the beginning of the novel. . . Xiaoze Liu, Ting Sun 0004, Feijie Wu, Cunxiang Wang, Xiaoqian Wang 0001, Jing Gao 0004 |
EMNLP | 5 |
| 2024 | Knowledge Conflicts for LLMs: A SurveyabstractThis survey provides an in-depth analysis of knowledge conflicts for large language models (LLMs), highlighting the complex challenges they encounter when blending contextual and parametric knowledge.Our focus is on three categories of knowledge conflicts: contextmemory, inter-context, and intra-memory conflict.These conflicts can significantly impact the trustworthiness and performance of LLMs, especially in real-world applications where noise and misinformation are common.By categorizing these conflicts, exploring the causes, examining the behaviors of LLMs under such conflicts, and reviewing available solutions, this survey aims to shed light on strategies for improving the robustness of LLMs, thereby serving as a valuable resource for advancing research in this evolving area. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang 0003, Yue Zhang 0004, Wei Xu 0039 |
EMNLP | 4 |
| 2024 | Nash CoT: Multi-Path Inference with Preference EquilibriumabstractChain of thought (CoT) is a reasoning framework that can enhance the performance of Large Language Models (LLMs) on complex inference tasks.In particular, among various studies related to CoT, multi-path inference stands out as a simple yet effective improvement.However, there is no optimal setting for the number of inference paths.Therefore, we have to increase the number of inference paths to obtain better results, which in turn increases the inference cost.To address this limitation, we can utilize question-related role templates to guide LLMs into relevant roles, thereby increasing the possibility of correct inferences for each path and further reducing dependence on the number of inference paths while improving reasoning accuracy.However, placing LLMs into specific roles may reduce their reasoning diversity and performance on a few tasks where role dependence is low.To alleviate the excessive immersion of the LLM into a specific role, we propose Nash CoT by constructing a competitive system on each path that balances the generation from role-specific LLMs' and the general LLMs' generation, thereby ensuring both effective role adoption and diversity in LLM generation further maintaining the performance of multi-path inference while reducing the requirement of the number of inference paths.We evaluate Nash CoT across various inference tasks, including Arabic Reasoning, Commonsense Question Answering, and Symbolic Inference, achieving results that are comparable to or better than those of multi-path CoT with the equal number of inference paths. Cunxiang Wang, Xiao Xiong |
EMNLP | 2 |
| 2024 | PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationabstractInstruction tuning large language models (LLMs) remains a challenging task, owing to the complexity of hyperparameter selection and the difficulty involved in evaluating the tuned models. To determine the optimal hyperparameters, an automatic, robust, and reliable evaluation benchmark is essential. However, establishing such a benchmark is not a trivial task due to the challenges associated with evaluation accuracy and privacy protection. In response to these challenges, we introduce a judge large language model, named PandaLM, which is trained to distinguish the superior model given several LLMs. PandaLM's focus extends beyond just the objective correctness of responses, which is the main focus of traditional evaluation datasets. It addresses vital subjective factors such as relative conciseness, clarity, adherence to instructions, comprehensiveness, and formality. To ensure the reliability of PandaLM, we collect a diverse human-annotated test dataset, where all contexts are generated by humans and labels are aligned with human preferences. Our findings reveal that PandaLM-7B offers a performance comparable to both GPT-3.5 and GPT-4. Impressively, PandaLM-70B surpasses their performance. PandaLM enables the evaluation of LLM to be fairer but with less cost, evidenced by significant improvements achieved by models tuned through PandaLM compared to their counterparts trained with default Alpaca's hyperparameters. In addition, PandaLM does not depend on API-based evaluations, thus avoiding potential data leakage. Yidong Wang 0003, Zhuohao Yu 0001, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen 0102, Chaoya Jiang, Rui Xie 0003, Jindong Wang 0001, Xing Xie 0001, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004 |
ICLR | 6 |
| 2024 | RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented GenerationabstractDespite Retrieval-Augmented Generation (RAG) has shown promising capability in leveraging external knowledge, a comprehensive evaluation of RAG systems is still challenging due to the modular nature of RAG, evaluation of long-form responses and reliability of measurements. In this paper, we propose a fine-grained evaluation framework, RAGChecker, that incorporates a suite of diagnostic metrics for both the retrieval and generation modules. Meta evaluation verifies that RAGChecker has significantly better correlations with human judgments than other evaluation metrics. Using RAGChecker, we evaluate 8 RAG systems and conduct an in-depth analysis of their performance, revealing insightful patterns and trade-offs in the design choices of RAG architectures. The metrics of RAGChecker can guide researchers and practitioners in developing more effective RAG systems. Dongyu Ru, Xiangkun Hu, Tianhang Zhang, Peng Shi 0010, Shuaichen Chang, Cheng Jiayang, Cunxiang Wang, Shichao Sun, Huanyu Li 0010, Binjie Wang, Jiarong Jiang, Tong He 0002, Zhiguo Wang 0006, Pengfei Liu 0003, Yue Zhang 0004, Zheng Zhang 0001 |
NeurIPS | 8 |
| 2024 | A Survey on Evaluation of Large Language ModelsabstractLarge language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey Yupeng Chang, Jindong Wang 0001, Yuan Wu 0002, Linyi Yang, Kaijie Zhu, Hao Chen 0102, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang 0003, Wei Ye 0004, Yue Zhang 0004, Yi Chang 0001, Philip S. Yu, Qiang Yang 0001, Xing Xie 0001 |
ACM Trans. Intell. Syst. Technol. | 9 |
| 2023 | Evaluating Open-QA EvaluationabstractThis study focuses on the evaluation of the Open Question Answering (Open-QA) task, which can directly estimate the factuality of large language models (LLMs). Current automatic evaluation methods have shown limitations, indicating that human evaluation still remains the most reliable approach. We introduce a new task, QA Evaluation (QA-Eval) and the corresponding dataset EVOUNA, designed to assess the accuracy of AI-generated answers in relation to standard answers within Open-QA. Our evaluation of these methods utilizes human-annotated results to measure their performance. Specifically, the work investigates methods that show high correlation with human evaluations, deeming them more reliable. We also discuss the pitfalls of current methods and methods to improve LLM-based evaluators. We believe this new QA-Eval task and corresponding dataset EVOUNA will facilitate the development of more effective automatic evaluation tools and prove valuable for future research in this area. All resources are available at https://github.com/wangcunxiang/QA-Eval and it is under the Apache-2.0 License. Cunxiang Wang, Sirui Cheng, Qipeng Guo, Yuanhao Yue, Zhikun Xu, Yidong Wang 0003, Xiangkun Hu, Yue Zhang 0004 |
NeurIPS | 1 |
| 2023 | Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base
Cunxiang Wang, Fuli Luo, Yanyang Li, Runxin Xu, Fei Huang 0002, Yue Zhang 0004 |
NLPCC (2) | 1 |
| 2021 | Can Generative Pre-trained Language Models Serve As Knowledge Bases for Closed-book QA?abstractCunxiang Wang, Pai Liu, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Cunxiang Wang, Pai Liu, Yue Zhang 0004 |
ACL/IJCNLP (1) | 1 |
| 2021 | Exploring Generalization Ability of Pretrained Language Models on Arithmetic and Logical Reasoning
Cunxiang Wang, Yue Zhang 0004 |
NLPCC (1) | 1 |
| 2019 | Does it Make Sense? And Why? A Pilot Study for Sense Making and ExplanationabstractIntroducing common sense to natural language understanding systems has received increasing research attention.It remains a fundamental question on how to evaluate whether a system has a sense making capability.Existing benchmarks measures commonsense knowledge indirectly and without explanation.In this paper, we release a benchmark to directly test whether a system can differentiate natural language statements that make sense from those that do not make sense.In addition, a system is asked to identify the most crucial reason why a statement does not make sense.We evaluate models trained over large-scale language modeling tasks as well as human performance, showing that there are different challenges for system sense making. Cunxiang Wang, Shuailong Liang, Yue Zhang 0004 |
ACL (1) | 1 |
| 2019 | Domain Representation for Knowledge Graph Embedding
Cunxiang Wang, Feiliang Ren, Zhichao Lin, Yue Zhang 0004 |
NLPCC (1) | 1 |