VLDB 2026 Research / reviewers in the wild / expert
Boxi Cao
dblp:295/9057
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-9916-7406ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does Question Really Matter? The Attribution of Answer Bias in LLM EvaluationabstractMultiple-choices question answering (MCQA) has emerged as one of the most popular task formats for large language models (LLMs) evaluation. Unfortunately, there exist substantial evidence that the evaluation of current MCQA benchmarks suffers from significant answer bias, which severely undermines the reliability of the evaluation conclusions. Specifically, many LLMs achieve performance significantly higher than random selection even when the questions are omitted from input information. To this end, we conduct a systematic investigation of the attribution of answer bias, and demonstrate a strong correlation between the degree of data contamination and the severity of answer bias, while the position of options and the popularity of answers have relatively minor effects. Building on these insights, we further propose OPD, a straightforward yet effective tool for contamination detection and dataset debiasing without requiring access to the model’s internal training data. Our findings and algorithms provide valuable insights for the design of future trustworthy LLM evaluation protocols. Boxi Cao, Ruotong Pan, Xianpei Han, Le Sun 0001 |
AAAI | 1 |
| 2026 | All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAGabstractDan Wang, Guozhao Mo, Yafei Shi, Cheng Zhang, Bo Zheng, Boxi Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Guozhao Mo, Bo Zheng 0007, Boxi Cao, Xuanang Chen, Yaojie Lu 0001, Ben He 0001, Xianpei Han, Le Sun 0001 |
ACL (1) | 6 |
| 2026 | Breaking the Spiral: A Utility-Driven Optimization Framework for Balanced Information Retrieval in the LLM EraabstractThe widespread adoption of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems is reshaping the landscape of information retrieval. However, the long-term effects of LLM-generated texts on retrieval systems remain underexplored, creating challenges for mitigating their impact. The effects are examined in this study, with a particular focus on the “Spiral of Silence” phenomenon, which refers to the marginalization of diverse information as certain types of content dominate, leading to a homogenized information ecosystem. To investigate this, a simulation pipeline is constructed to model the iterative introduction of LLM-generated texts into retrieval systems. Experimental results across multiple iterations reveal that as the presence of LLM-generated texts within the system grows, retrieval systems exhibit a stronger tendency to retrieve these texts. This trend, in turn, reduces the visibility of human-generated content, diminishes diversity, propagates errors, and results in a notable decline in retrieval performance. To address these challenges, we propose a Utility-Driven Multi-Objective Optimization (UMO) framework to effectively mitigate the “Spiral of Silence.” This framework employs a two-phase approach: an optimization phase, leveraging the NSGA-II algorithm to derive optimal preference weights for multiple objectives, and a memorization phase, which directly integrates these weights into the retrieval vector space without requiring additional model retraining. Experimental results demonstrate that this framework maintains stable retrieval effectiveness, improves the retrieval proportion of human-generated content, reduces the excessive influence of LLM-generated texts, and preserves information diversity, effectively mitigating the “Spiral of Silence.” Xiaoyang Chen 0001, Ben He 0001, Xianpei Han, Tianshu Wang 0002, Boxi Cao, Le Sun 0001, Yingfei Sun |
ACM Trans. Inf. Syst. | 6 |
| 2025 | Memorizing is Not Enough: Deep Knowledge Injection Through ReasoningabstractRuoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Yingfei Sun, Xiangang Li, Le Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu 0001, Xianpei Han, Ben He 0001, Yingfei Sun, Xiangang Li, Le Sun 0001 |
ACL (1) | 3 |
| 2024 | Spiral of Silence: How is Large Language Model Killing Information Retrieval? - A Case Study on Open Domain Question AnsweringabstractXiaoyang Chen, Ben He, Hongyu Lin, Xianpei Han, Tianshu Wang, Boxi Cao, Le Sun, Yingfei Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xiaoyang Chen 0001, Ben He 0001, Xianpei Han, Tianshu Wang 0002, Boxi Cao, Le Sun 0001, Yingfei Sun |
ACL (1) | 6 |
| 2024 | Learning or Self-aligning? Rethinking Instruction Fine-tuningabstractMengjie Ren, Boxi Cao, Hongyu Lin, Cao Liu, Xianpei Han, Ke Zeng, Wan Guanglu, Xunliang Cai, Le Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mengjie Ren, Boxi Cao, Cao Liu, Xianpei Han, Guanglu Wan, Le Sun 0001 |
ACL (1) | 2 |
| 2024 | Retentive or Forgetful? Diving into the Knowledge Memorizing Mechanism of Language ModelsabstractMemory is one of the most essential cognitive functions serving as a repository of world knowledge and episodes of activities. In recent years, large-scale pre-trained language models have shown remarkable memorizing ability. On the contrary, vanilla neural networks without pre-training have been long observed suffering from the catastrophic forgetting problem. To investigate such a retentive-forgetful contradiction and understand the memorizing dynamic mechanism of language models, we conduct thorough experiments by controlling the target knowledge types, the learning strategies and the learning schedules. We find that: 1) Vanilla language models without pre-training are forgetful; 2) Pre-training leads to retentive language models; 3) Knowledge relevance and diversification significantly influence the memory formation. These conclusions are useful for understanding the abilities of pre-trained language models and shed light on designing and evaluating new learning and inference algorithms of language models. Boxi Cao, Qiaoyu Tang, Shanshan Jiang 0001, Bin Dong 0003, Xianpei Han, Jiawei Chen 0011, Tianshu Wang 0002, Le Sun 0001 |
LREC/COLING | 1 |
| 2024 | Not All Contexts Are Equal: Teaching LLMs Credibility-aware GenerationabstractThe rapid development of large language models has led to the widespread adoption of Retrieval-Augmented Generation (RAG), which integrates external knowledge to alleviate knowledge bottlenecks and mitigate hallucinations.However, the existing RAG paradigm inevitably suffers from the impact of flawed information introduced during the retrieval phrase, thereby diminishing the reliability and correctness of the generated outcomes.In this paper, we propose Credibility-aware Generation (CAG), a universally applicable framework designed to mitigate the impact of flawed information in RAG.At its core, CAG aims to equip models with the ability to discern and process information based on its credibility.To this end, we propose an innovative data transformation framework that generates data based on credibility, thereby effectively endowing models with the capability of CAG.Furthermore, to accurately evaluate the models' capabilities of CAG, we construct a comprehensive benchmark covering three critical real-world scenarios.Experimental results demonstrate that our model can effectively understand and employ credibility for generation, significantly outperform other models with retrieval augmentation, and exhibit robustness despite the increasing noise in the context. 1 Ruotong Pan, Boxi Cao, Xianpei Han, Jia Zheng 0009, Le Sun 0001 |
EMNLP | 2 |
| 2023 | Learning In-context Learning for Named Entity RecognitionabstractJiawei Chen, Yaojie Lu, Hongyu Lin, Jie Lou, Wei Jia, Dai Dai, Hua Wu, Boxi Cao, Xianpei Han, Le Sun. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiawei Chen 0011, Yaojie Lu 0001, Jie Lou, Dai Dai, Hua Wu 0003, Boxi Cao, Xianpei Han, Le Sun 0001 |
ACL (1) | 8 |
| 2023 | Does the Correctness of Factual Knowledge Matter for Factual Knowledge-Enhanced Pre-trained Language Models?abstractIn recent years, the injection of factual knowledge has been observed to have a significant positive correlation to the downstream task performance of pre-trained language models.However, existing work neither demonstrates that pre-trained models successfully learn the injected factual knowledge nor proves that there is a causal relation between injected factual knowledge and downstream performance improvements.In this paper, we introduce a counterfactual-based analysis framework to explore the causal effects of factual knowledge injection on the performance of language models within pretrain-finetune paradigm.Instead of directly probing the language model or exhaustively enumerating potential confounding factors, we analyze this issue by perturbing the factual knowledge sources at different scales and comparing the performance of pre-trained language models before and after the perturbation.Surprisingly, throughout our experiments, we find that although the knowledge seems to be successfully injected, the correctness of injected knowledge only has a very limited effect on the models' downstream performance.This finding strongly challenges previous assumptions that the injected factual knowledge is the key for language models to achieve performance improvements on downstream tasks in pretrain-finetune paradigm. Boxi Cao, Qiaoyu Tang, Xianpei Han, Le Sun 0001 |
EMNLP | 1 |
| 2022 | Can Prompt Probe Pretrained Language Models? Understanding the Invisible Risks from a Causal ViewabstractPrompt-based probing has been widely used in evaluating the abilities of pretrained language models (PLMs).Unfortunately, recent studies have discovered such an evaluation may be inaccurate, inconsistent and unreliable.Furthermore, the lack of understanding its inner workings, combined with its wide applicability, has the potential to lead to unforeseen risks for evaluating and applying PLMs in real-world applications.To discover, understand and quantify the risks, this paper investigates the promptbased probing from a causal view, highlights three critical biases which could induce biased results and conclusions, and proposes to conduct debiasing via causal intervention.This paper provides valuable insights for the design of unbiased datasets, better probing frameworks and more reliable evaluations of pretrained language models.Furthermore, our conclusions also echo that we need to rethink the criteria for identifying better pretrained language models 1 . Boxi Cao, Xianpei Han, Fangchao Liu, Le Sun 0001 |
ACL (1) | 1 |
| 2022 | Pre-training to Match for Unified Low-shot Relation ExtractionabstractLow-shot relation extraction (RE) aims to recognize novel relations with very few or even no samples, which is critical in real scenario application.Few-shot and zero-shot RE are two representative low-shot RE tasks, which seem to be with similar target but require totally different underlying abilities.In this paper, we propose Multi-Choice Matching Networks to unify low-shot relation extraction.To fill in the gap between zero-shot and few-shot RE, we propose the triplet-paraphrase meta-training, which leverages triplet paraphrase to pre-train zero-shot label matching ability and uses metalearning paradigm to learn few-shot instance summarizing ability.Experimental results on three different low-shot RE tasks show that the proposed method outperforms strong baselines by a large margin, and achieve the best performance on few-shot RE leaderboard 1 . Fangchao Liu, Xianpei Han, Boxi Cao, Le Sun 0001 |
ACL (1) | 4 |
| 2021 | Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge BasesabstractBoxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, Jin Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Boxi Cao, Xianpei Han, Le Sun 0001, Lingyong Yan, Meng Liao, Tong Xue, Jin Xu 0014 |
ACL/IJCNLP (1) | 1 |