VLDB 2026 Research / reviewers in the wild / expert
Zhihui Xie 0002
dblp:31/3570-2
· DBLP profile ↗
16ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-1556-0700ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic WorkflowsabstractQiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong |
ACL (1) | 4 |
| 2025 | VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsabstractVision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io. Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049 |
CVPR | 3 |
| 2025 | Long Chain-of-Thought Fine-tuning via Understanding-to-Reasoning TransitionabstractChenxin An, Zhihui Xie, Xiaonan Li, Ming Zhong, Shansan Gong, Lei Li, Jun Zhang, Jingjing Xu, Lingpeng Kong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chenxin An, Zhihui Xie 0002, Ming Zhong 0005, Shansan Gong, Lei Li 0039, Jun Zhang 0003, Jingjing Xu 0001, Lingpeng Kong |
EMNLP | 2 |
| 2025 | Jailbreaking as a Reward Misspecification ProblemabstractThe widespread adoption of large language models (LLMs) has raised concerns about their safety and reliability, particularly regarding their vulnerability to adversarial attacks. In this paper, we propose a new perspective that attributes this vulnerability to reward misspecification during the alignment process. This misspecification occurs when the reward function fails to accurately capture the intended behavior, leading to misaligned model outputs. We introduce a metric ReGap to quantify the extent of reward misspecification and demonstrate its effectiveness and robustness in detecting harmful backdoor prompts. Building upon these insights, we present ReMiss, a system for automated red teaming that generates adversarial prompts in a reward-misspecified space. ReMiss achieves state-of-the-art attack success rates on the AdvBench benchmark against various target aligned LLMs while preserving the human readability of the generated prompts. Furthermore, these attacks on open-source models demonstrate high transferability to closed-source models like GPT-4o and out-of-distribution tasks from HarmBench. Detailed analysis highlights the unique advantages of the proposed reward misspecification objective compared to previous methods, offering new insights for improving LLM safety and robustness. Zhihui Xie 0002, Jiahui Gao 0002, Lei Li 0039, Zhenguo Li, Qi Liu 0049, Lingpeng Kong |
ICLR | 1 |
| 2025 | Teaching Language Models to Critique via Reinforcement LearningabstractTeaching large language models (LLMs) to critique and refine their outputs is crucial for building systems that can iteratively improve, yet it is fundamentally limited by the ability to provide *accurate judgments* and *actionable suggestions*. In this work, we study LLM critics for code generation and propose $\texttt{CTRL}$, a framework for $\texttt{C}$ritic $\texttt{T}$raining via $\texttt{R}$einforcement $\texttt{L}$earning, which trains a critic model to generate feedback that maximizes correction performance for a fixed generator model without human supervision. Our results demonstrate that critics trained with $\texttt{CTRL}$ significantly enhance pass rates and mitigate compounding errors across both base and stronger generator models.
Furthermore, we show that these critic models act as accurate generative reward models and enable test-time scaling through iterative critique-revision, achieving up to 106.1\% relative improvements across challenging code generation benchmarks. Zhihui Xie 0002, Liyu Chen, Weichao Mao, Jingjing Xu 0001, Lingpeng Kong |
ICML | 1 |
| 2025 | Multi-armed Bandit Algorithms for the Boolean Satisfiability Problem: A SurveyabstractThis paper provides a survey of recent literature on the use of multi-armed bandit algorithms to solve the Boolean satisfiability problem (SAT), a well-known NP-complete problem with broad applications in academia and industry. The application of bandit algorithms in modern SAT solvers has achieved great success in recent years, as evidenced by the excellent performance of SAT solvers using bandit algorithms in SAT competitions. Bandit algorithms are classic randomized optimization algorithms that strike a balance between exploration and exploitation and can aid in designing and improving heuristics in SAT solvers. In this paper, we introduce several aspects of the application of bandit algorithms in modern SAT solvers, ranging from heuristic methods in CDCL and SLS solvers to strategies in parallel SAT solvers. The use of bandit algorithms in SAT solvers still holds great potential. In conclusion of the survey, we summarize the current issues and suggest possible future research directions. Zhihui Xie 0002, Shuai Li 0010 |
SOCS | 1 |
| 2024 | VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models AlignmentabstractLei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049 |
EMNLP | 2 |
| 2024 | Calibrating Reasoning in Language Models with Internal ConsistencyabstractLarge language models (LLMs) have demonstrated impressive capabilities in various reasoning tasks, aided by techniques like chain-of-thought prompting that elicits verbalized reasoning. However, LLMs often generate text with obvious mistakes and contradictions, raising doubts about their ability to robustly process and utilize generated rationales. In this work, we investigate reasoning in LLMs through the lens of internal representations, focusing on how these representations are influenced by generated rationales. Our preliminary analysis reveals that while generated rationales improve answer accuracy, inconsistencies emerge between the model’s internal representations in middle layers and those in final layers, potentially undermining the reliability of their reasoning processes. To address this, we propose internal consistency as a measure of the model’s confidence by examining the agreement of latent predictions decoded from intermediate layers. Extensive empirical studies across different models and datasets demonstrate that internal consistency effectively distinguishes between correct and incorrect reasoning paths. Motivated by this, we propose a new approach to calibrate reasoning by up-weighting reasoning paths with high internal consistency, resulting in a significant boost in reasoning performance. Further analysis uncovers distinct patterns in attention and feed-forward modules across layers, providing insights into the emergence of internal inconsistency. In summary, our results demonstrate the potential of using internal representations for self-evaluation of LLMs. Zhihui Xie 0002, Jizhou Guo, Tong Yu 0001, Shuai Li 0010 |
NeurIPS | 1 |
| 2024 | Learning Versatile Skills with Curriculum MaskingabstractMasked prediction has emerged as a promising pretraining paradigm in offline reinforcement learning (RL) due to its versatile masking schemes, enabling flexible inference across various downstream tasks with a unified model. Despite the versatility of masked prediction, it remains unclear how to balance the learning of skills at different levels of complexity. To address this, we propose CurrMask, a curriculum masking pretraining paradigm for sequential decision making. Motivated by how humans learn by organizing knowledge in a curriculum, CurrMask adjusts its masking scheme during pretraining for learning versatile skills. Through extensive experiments, we show that CurrMask exhibits superior zero-shot performance on skill prompting tasks, goal-conditioned planning tasks, and competitive finetuning performance on offline RL tasks. Additionally, our analysis of training dynamics reveals that CurrMask gradually acquires skills of varying complexity by dynamically adjusting its masking scheme. Zhihui Xie 0002, Zichuan Lin, Deheng Ye, Shuai Li 0010 |
NeurIPS | 2 |
| 2024 | Toward joint utilization of absolute and relative bandit feedback for conversational recommendation
Yu Xia 0007, Zhihui Xie 0002, Tong Yu 0001, Canzhe Zhao, Shuai Li 0010 |
User Model. User Adapt. Interact. | 2 |
| 2023 | Future-conditioned Unsupervised Pretraining for Decision TransformerabstractRecent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promising, return conditioning is limited to training data labeled with rewards and therefore faces challenges in learning from unsupervised data. In this work, we aim to utilize generalized future conditioning to enable efficient unsupervised pretraining from reward-free and sub-optimal offline data. We propose Pretrained Decision Transformer (PDT), a conceptually simple approach for unsupervised RL pretraining. PDT leverages future trajectory information as a privileged context to predict actions during training. The ability to make decisions based on both present and future factors enhances PDT's capability for generalization. Besides, this feature can be easily incorporated into a return-conditioned framework for online finetuning, by assigning return values to possible futures and sampling future embeddings based on their respective values. Empirically, PDT outperforms or performs on par with its supervised pretraining counterpart, especially when dealing with sub-optimal data. Further analysis reveals that PDT can extract diverse behaviors from offline data and controllably sample high-return behaviors by online finetuning. Code is available at here. Zhihui Xie 0002, Zichuan Lin, Deheng Ye, Qiang Fu 0016, Shuai Li 0010 |
ICML | 1 |
| 2022 | Discovering Low-rank Subspaces for Language-agnostic Multilingual RepresentationsabstractLarge pretrained multilingual language models (ML-LMs) have shown remarkable capabilities of zero-shot cross-lingual transfer, without direct cross-lingual supervision.While these results are promising, follow-up works found that, within the multilingual embedding spaces, there exists strong language identity information which hinders the expression of linguistic factors shared across languages.For semantic tasks like cross-lingual sentence retrieval, it is desired to remove such language identity signals to fully leverage semantic information.In this work, we provide a novel view of projecting away language-specific factors from a multilingual embedding space.Specifically, we discover that there exists a low-rank subspace that primarily encodes information irrelevant to semantics (e.g., syntactic information).To identify this subspace, we present a simple but effective unsupervised method based on singular value decomposition with multiple monolingual corpora as input.Once the subspace is found, we can directly project the original embeddings into the null space to boost language agnosticism without finetuning.We systematically evaluate our method on various tasks including the challenging language-agnostic QA retrieval task.Empirical results show that applying our method consistently leads to improvements over commonly used ML-LMs. Zhihui Xie 0002, Handong Zhao, Tong Yu 0001, Shuai Li 0010 |
EMNLP | 1 |
| 2022 | Dynamics-Aware Adaptation for Reinforcement Learning Based Cross-Domain Interactive RecommendationabstractInteractive recommender systems (IRS) have received wide attention in recent years. To capture users' dynamic preferences and maximize their long-term engagement, IRS are usually formulated as reinforcement learning (RL) problems. Despite the promise to solve complex decision-making problems, RL-based methods generally require a large amount of online interaction, restricting their applications due to economic considerations. One possible direction to alleviate this issue is cross-domain recommendation that aims to leverage abundant logged interaction data from a source domain (e.g., adventure genre in movie recommendation) to improve the recommendation quality in the target domain (e.g., crime genre). Nevertheless, prior studies mostly focus on adapting the static representations of users/items. Few have explored how the temporally dynamic user-item interaction patterns transform across domains. Junda Wu, Zhihui Xie 0002, Tong Yu 0001, Handong Zhao, Ruiyi Zhang 0002, Shuai Li 0010 |
SIGIR | 2 |
| 2022 | Knowledge-aware Conversational Preference Elicitation with Bandit FeedbackabstractConversational recommender systems (CRSs) have been proposed recently to mitigate the cold-start problem suffered by the traditional recommender systems. By introducing conversational key-terms, existing conversational recommenders can effectively reduce the need for extensive exploration and elicit the user preferences faster and more accurately. However, existing conversational recommenders leveraging key-terms heavily rely on the availability and quality of the key-terms, and their performances might degrade significantly when the key-terms are incomplete or not well labeled, which usually happens when there are new items being consistently incorporated into the systems and involving lots of human efforts to acquire well-labeled key-terms is costly. Besides, existing CRS methods leverage the feedback to different conversational key-terms separately, without considering the underlying relations between the key-terms. In this case, the learning of the conversational recommenders is sample inefficient, especially when there is a large number of candidate conversational key-terms. Canzhe Zhao, Tong Yu 0001, Zhihui Xie 0002, Shuai Li 0010 |
WWW | 3 |
| 2021 | Comparison-based Conversational Recommender System with Relative Bandit FeedbackabstractWith the recent advances of conversational recommendations, the recommender system is able to actively and dynamically elicit user preference via conversational interactions. To achieve this, the system periodically queries users' preference on attributes and collects their feedback. However, most existing conversational recommender systems only enable the user to provide absolute feedback to the attributes. In practice, the absolute feedback is usually limited, as the users tend to provide biased feedback when expressing the preference. Instead, the user is often more inclined to express comparative preferences, since user preferences are inherently relative. To enable users to provide comparative preferences during conversational interactions, we propose a novel comparison-based conversational recommender system. The relative feedback, though more practical, is not easy to be incorporated since its feedback scale is always mismatched with users' absolute preferences. With effectively collecting and understanding the relative feedback from an interactive manner, we further propose a new bandit algorithm, which we call RelativeConUCB. The experiments on both synthetic and real-world datasets validate the advantage of our proposed method, compared to the existing bandit algorithms in the conversational recommender systems. Zhihui Xie 0002, Tong Yu 0001, Canzhe Zhao, Shuai Li 0010 |
SIGIR | 1 |
| 2020 | Layered Neighborhood Expansion for Incremental Multiple Graph Matching
Zhihui Xie 0002, Junchi Yan, Yinqiang Zheng, Xiaokang Yang 0001 |
ECCV (10) | 2 |