VLDB 2026 Research / reviewers in the wild / expert
Xuefeng Bai 0001
dblp:18/2759-1
· DBLP profile ↗
23ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-7044-0683ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 9 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive ThinkingabstractWeiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen, Yibin Chen, Weili Guan, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weiyang Huang, Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Yibin Chen, Weili Guan, Min Zhang 0005 |
ACL (1) | 2 |
| 2025 | Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine TranslationabstractAndong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai, Muyun Yang, Liqiang Nie, Jie Liu, Tiejun Zhao, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Andong Chen 0001, Kehai Chen, Xuefeng Bai 0001, Muyun Yang, Liqiang Nie, Jie Liu 0001, Tiejun Zhao, Min Zhang 0005 |
ACL (1) | 4 |
| 2025 | Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward ModelingabstractReinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution shift.Current approaches typically address this issue through online sampling from the target policy, which requires significant computational resources.In this paper, we hypothesize that during off-policy training, while the ranking order of output generated by policy changes, their overall distribution remains relatively stable.This stability allows the conversion of the sampling process from the target policy into a computationally efficient reranking of preference data.Building on this hypothesis, we propose a new framework that leverages the model's intrinsic safety judgment capability to extract reward signals, which are then used to calculate label confidence for preference reordering.Extensive experiments and theoretical analysis demonstrate that the proposed method effectively addresses the distribution shift issue, remarkably enhancing the safety performance while avoiding about 300x computational overheads. Qiyuan Deng, Xuefeng Bai 0001, Kehai Chen, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
ACL (1) | 2 |
| 2025 | Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks.Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs.To study the reason behind these limitations, we propose VGCURE, a comprehensive benchmark covering 22 tasks for examining the fundamental graph understanding and reasoning capacities of LVLMs.Extensive evaluations conducted on 14 LVLMs reveal that LVLMs are weak in basic graph understanding and reasoning tasks, particularly those concerning relational or structurally complex information.Based on this observation, we propose a structure-aware fine-tuning framework to enhance LVLMs with structure learning abilities through three self-supervised learning tasks.Experiments validate the effectiveness of our method in improving LVLMs' performance on fundamental and downstream graph learning tasks, as well as enhancing their robustness against complex visual graphs. Yingjie Zhu, Xuefeng Bai 0001, Kehai Chen, Yang Xiang 0003, Jun Yu 0002, Min Zhang 0005 |
ACL (1) | 2 |
| 2025 | Benchmarking LLMs for Translating Classical Chinese Poetry: Evaluating Adequacy, Fluency, and EleganceabstractLarge language models (LLMs) have shown remarkable performance in general translation tasks.However, the increasing demand for high-quality translations that are not only adequate but also fluent and elegant.To assess the extent to which current LLMs can meet these demands, we introduce a suitable benchmark (PoetMT) for translating classical Chinese poetry into English.This task requires not only adequacy in translating culturally and historically significant content but also a strict adherence to linguistic fluency and poetic elegance.Our study reveals that existing LLMs fall short of this task.To address these issues, we propose RAT, a Retrieval-Augmented machine Translation method that enhances the translation process by incorporating knowledge related to classical poetry.Additionally, we propose an automatic evaluation metric based on GPT-4, which better assesses translation quality in terms of adequacy, fluency, and elegance, overcoming the limitations of traditional metrics.Our dataset and code will be made available 1 . Andong Chen 0001, Lianzhang Lou, Kehai Chen, Xuefeng Bai 0001, Yang Xiang 0003, Muyun Yang, Tiejun Zhao, Min Zhang 0005 |
EMNLP | 4 |
| 2025 | Generator-Assistant Stepwise Rollback Framework for Large Language Model AgentabstractLarge language model (LLM) agents typically adopt a step-by-step reasoning framework, in which they interleave the processes of thinking and acting to accomplish the given task.However, this paradigm faces a deeprooted one-pass issue whereby each generated intermediate thought is plugged into the trajectory regardless of its correctness, which can cause irreversible error propagation.To address the issue, this paper proposes a novel framework called Generator-Assistant Stepwise Rollback (GA-Rollback) to induce better decision-making for LLM agents.Particularly, GA-Rollback utilizes a generator to interact with the environment and an assistant to examine each action produced by the generator, where the assistant triggers a rollback operation upon detection of incorrect actions.Moreover, we introduce two additional strategies tailored for the rollback scenario to further improve its effectiveness.Extensive experiments show that GA-Rollback achieves significant improvements over several strong baselines on three widely used benchmarks.Our analysis further reveals that GA-Rollback can function as a robust plug-and-play module, integrating seamlessly with other methods. 1 Xingzuo Li, Kehai Chen, Xuefeng Bai 0001, Yong Xu 0001, Min Zhang 0005 |
EMNLP | 4 |
| 2025 | Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated MarginabstractAdapting vision-language models (VLMs) to downstream tasks with pseudolabels has gained increasing attention. A major obstacle is that the pseudolabels generated by VLMs tend to be imbalanced, leading to inferior performance. While existing methods have explored various strategies to address this, the underlying causes of imbalance remain insufficiently investigated. To fill this gap, we delve into imbalanced pseudolabels and identify two primary contributing factors: concept mismatch and concept confusion. To mitigate these two issues, we propose a novel framework incorporating concept alignment and confusion-aware calibrated margin mechanisms. The core of our approach lies in enhancing underperforming classes and promoting balanced predictions across categories, thus mitigating imbalance. Extensive experiments on six benchmark datasets with three learning paradigms demonstrate that the proposed method effectively enhances the accuracy and balance of pseudolabels, achieving a relative improvement of 6.29% over the SoTA method. Our code is avaliable at https://github.com/Noahwangyuchen/CAP Xuefeng Bai 0001, Xiucheng Li, Weili Guan, Liqiang Nie, Xinyang Chen 0001 |
ICML | 2 |
| 2025 | A Survey on the Feedback Mechanism of LLM-based AI AgentsabstractLarge language models (LLMs) are increasingly being adopted to develop general-purpose AI agents. However, it remains challenging for these LLM-based AI agents to efficiently learn from feedback and iteratively optimize their strategies. To address this challenge, tremendous efforts have been dedicated to designing diverse feedback mechanisms for LLM-based AI agents. To provide a comprehensive overview of this rapidly evolving field, this paper presents a systematic review of these studies, offering a holistic perspective on the feedback mechanisms in LLM-based AI agents. We begin by discussing the construction of LLM-based AI agents, introducing a generalized framework that encapsulates much of the existing work. Next, we delve into the exploration of feedback mechanisms, categorizing them into four distinct types: internal feedback, external feedback, multi-agent feedback, and human feedback. Additionally, we provide an overview of evaluation protocols and benchmarks specifically tailored for LLM-based AI agents. Finally, we highlight the significant challenges and identify potential directions for future studies. The relevant papers are summarized and will be consistently updated at https://github.com/kevinson7515/Agents-Feedback-Mechanisms. Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Xiucheng Li, Yang Xiang 0003, Jin Liu 0012, Hong-Dong Li, Yaowei Wang 0001, Liqiang Nie, Min Zhang 0005 |
IJCAI | 2 |
| 2025 | XIFBench: Evaluating Large Language Models on Multilingual Instruction FollowingabstractLarge Language Models (LLMs) have demonstrated remarkable instruction-following capabilities across various applications. However, their performance in multilingual settings lacks systematic investigation, with existing evaluations lacking fine-grained constraint analysis across diverse linguistic contexts. We introduce XIFBench, a comprehensive constraint-based benchmark for evaluating multilingual instruction-following abilities of LLMs, comprising 558 instructions with 0-5 additional constraints across five categories (Content, Style, Situation, Format, and Numerical) in six languages spanning different resource levels. To support reliable and consistent cross-lingual evaluation, we implement three methodological innovations: cultural accessibility annotation, constraint-level translation validation, and requirement-based evaluation using English requirements as semantic anchors across languages. Extensive experiments with various LLMs not only quantify performance disparities across resource levels but also provide detailed insights into how language resources, constraint categories, instruction complexity, and cultural specificity influence multilingual instruction-following. Our code and data are available at https://github.com/zhenyuli801/XIFBench. Kehai Chen, Xuefeng Bai 0001, Yaoyin Zhang, Xuchen Wei, Juntao Li 0005, Min Zhang 0005 |
NeurIPS | 4 |
| 2025 | Exploring the Translation Mechanism of Large Language ModelsabstractWhile large language models (LLMs) demonstrate remarkable success in multilingual translation, their internal core translation mechanisms, even at the fundamental word level, remain insufficiently understood.
To address this critical gap, this work introduces a systematic framework for interpreting the mechanism behind LLM translation from the perspective of computational components.
This paper first proposes subspace-intervened path patching for precise, fine-grained causal analysis, enabling the detection of components crucial to translation tasks and subsequently characterizing their behavioral patterns in human-interpretable terms.
Comprehensive experiments reveal that translation is predominantly driven by a sparse subset of components: specialized attention heads serve critical roles in extracting source language, translation indicators, and positional features, which are then integrated and processed by specific multi-layer perceptrons (MLPs) into intermediary English-centric latent representations before ultimately yielding the final translation.
The significance of these findings is underscored by the empirical demonstration that targeted fine-tuning a minimal parameter subset (<5%) enhances translation performance while preserving general capabilities. This result further indicates that these crucial components generalize effectively to sentence-level translation and are instrumental in elucidating more intricate translation tasks. Kehai Chen, Xuefeng Bai 0001, Xiucheng Li, Yang Xiang 0003, Min Zhang 0005 |
NeurIPS | 3 |
| 2025 | Adaptive Inner Speech Text Alignment for LLM-Based Speech Translation
Henglyu Liu, Andong Chen 0001, Kehai Chen, Xuefeng Bai 0001, Meizhi Zhong, Yuan Qiu 0001, Min Zhang 0005 |
NLPCC (3) | 4 |
| 2025 | TianWen: A Comprehensive Benchmark for Evaluating LLMs in Chinese Classical Poetry Understanding and Reasoning
Zhenwu Pei, Rongbo Chen, Xuefeng Bai 0001, Kehai Chen, Yingjie Zhu, Andong Chen 0001, Min Zhang 0005 |
NLPCC (1) | 3 |
| 2025 | TF-Attack: Transferable and fast adversarial attacks on large language models
Kehai Chen, Lemao Liu, Xuefeng Bai 0001, Yang Xiang 0003, Min Zhang 0005 |
Knowl. Based Syst. | 4 |
| 2023 | Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved AnnotationabstractYulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai, Yueguan Wang, Ming Zhong, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yulong Chen 0001, Xuefeng Bai 0001, Yueguan Wang, Ming Zhong 0005, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang 0004 |
ACL (1) | 4 |
| 2022 | Graph Pre-training for AMR Parsing and Generationabstractmeaning representation (AMR) highlights the core semantic information of text in a graph structure.Recently, pre-trained language models (PLMs) have advanced tasks of AMR parsing and AMR-to-text generation, respectively.However, PLMs are typically pretrained on textual data, thus are sub-optimal for modeling structural knowledge.To this end, we investigate graph self-supervised training to improve the structure awareness of PLMs over AMR graphs.In particular, we introduce two graph auto-encoding strategies for graphto-graph pre-training and four tasks to integrate text and graph information during pre-training.We further design a unified framework to bridge the gap between pre-training and fine-tuning tasks.Experiments on both AMR parsing and AMR-to-text generation show the superiority of our model.To our knowledge, we are the first to consider pre-training on semantic graphs. Xuefeng Bai 0001, Yulong Chen 0001, Yue Zhang 0004 |
ACL (1) | 1 |
| 2022 | Semantic-based Pre-training for Dialogue UnderstandingabstractPre-trained language models have made great progress on dialogue tasks. However, these models are typically trained on surface dialogue text, thus are proven to be weak in understanding the main semantic meaning of a dialogue context. We investigate Abstract Meaning Representation (AMR) as explicit semantic knowledge for pre-training models to capture the core semantic information in dialogues during pre-training. In particular, we propose a semantic-based pre-training framework that extends the standard pre-training framework (Devlin et al.,2019) by three tasks for learning 1) core semantic units, 2) semantic relations and 3) the overall semantic representation according to AMR graphs. Experiments on the understanding of both chit-chats and task-oriented dialogues show the superiority of our model. To our knowledge, we are the first to leverage a deep semantic representation for dialogue pre-training. Xuefeng Bai 0001, Linfeng Song, Yue Zhang 0004 |
COLING | 1 |
| 2022 | Cross-domain Generalization for AMR ParsingabstractMeaning Representation (AMR) parsing aims to predict an AMR graph from textual input.Recently, there has been notable growth in AMR parsing performance.However, most existing work focuses on improving the performance in the specific domain, ignoring the potential domain dependence of AMR parsing systems.To address this, we extensively evaluate five representative AMR parsers on five domains and analyze challenges to cross-domain AMR parsing.We observe that challenges to cross-domain AMR parsing mainly arise from the distribution shift of words and AMR concepts.Based on our observation, we investigate two approaches to reduce the domain distribution divergence of text and AMR features, respectively.Experimental results on two out-of-domain test sets show the superiority of our method. Xuefeng Bai 0001, Sen Yang 0005, Leyang Cui, Linfeng Song, Yue Zhang 0004 |
EMNLP | 1 |
| 2021 | Semantic Representation for Dialogue ModelingabstractXuefeng Bai, Yulong Chen, Linfeng Song, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xuefeng Bai 0001, Yulong Chen 0001, Linfeng Song, Yue Zhang 0004 |
ACL/IJCNLP (1) | 1 |
| 2021 | Sentence-State LSTMs For Sequence-to-Sequence Learning
Xuefeng Bai 0001, Yafu Li, Zhirui Zhang, Mingzhou Xu, Boxing Chen, Weihua Luo, Derek F. Wong, Yue Zhang 0004 |
NLPCC (1) | 1 |
| 2021 | Investigating Typed Syntactic Dependencies for Targeted Sentiment Classification Using Graph Attention Neural NetworkabstractTargeted sentiment classification predicts the sentiment polarity on given target mentions in input texts. Dominant methods employ neural networks for encoding the input sentence and extracting relations between target mentions and their contexts. Recently, graph neural network has been investigated for integrating dependency syntax for the task, achieving the state-of-the-art results. However, existing methods do not consider dependency label information, which can be intuitively useful. To solve the problem, we investigate a novel relational graph attention network that integrates typed syntactic dependency information. Results on standard benchmarks show that our method can effectively leverage label information for improving targeted sentiment classification performances. Our final model significantly outperforms state-of-the-art syntax-based approaches. Xuefeng Bai 0001, Pengbo Liu 0006, Yue Zhang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Online Back-Parsing for AMR-to-Text GenerationabstractAMR-to-text generation aims to recover a text containing the same meaning as an input AMR graph.Current research develops increasingly powerful graph encoders to better represent AMR graphs, with decoders based on standard language modeling being used to generate outputs.We propose a decoder that back predicts projected AMR graphs on the target sentence during text generation.As the result, our outputs can better preserve the input meaning than standard decoders.Experiments on two AMR benchmarks show the superiority of our model over the previous state-of-the-art system based on graph Transformer. Xuefeng Bai 0001, Linfeng Song, Yue Zhang 0004 |
EMNLP (1) | 1 |
| 2019 | A Bilingual Adversarial Autoencoder for Unsupervised Bilingual Lexicon InductionabstractUnsupervised bilingual lexicon induction aims to generate bilingual lexicons without any cross-lingual signals. Successfully solving this problem would benefit many downstream tasks, such as unsupervised machine translation and transfer learning. In this work, we propose an unsupervised framework, named bilingual adversarial autoencoder, which automatically generates bilingual lexicon for a pair of languages from their monolingual word embeddings. In contrast to existing frameworks which learn a direct cross-lingual mapping of word embeddings from the source language to the target language, we train two autoencoders jointly to transform the source and the target monolingual word embeddings into a shared embedding space, where a word and its translation are close to each other. In this way, we capture the cross-lingual features of word embeddings from different languages and use them to induce bilingual lexicons. By conducting extensive experiments across eight language pairs, we demonstrate that the proposed method significantly outperforms the existing adversarial methods and even achieves best-published results across most language pairs. Xuefeng Bai 0001, Hailong Cao, Kehai Chen, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Improving Vector Space Word Representations Via Kernel Canonical Correlation AnalysisabstractCross-lingual word embeddings are representations for vocabularies of two or more languages in one common continuous vector space and are widely used in various natural language processing tasks. A state-of-the-art way to generate cross-lingual word embeddings is to learn a linear mapping, with an assumption that the vector representations of similar words in different languages are related by a linear relationship. However, this assumption does not always hold true, especially for substantially different languages. We therefore propose to use kernel canonical correlation analysis to capture a non-linear relationship between word embeddings of two languages. By extensively evaluating the learned word embeddings on three tasks (word similarity, cross-lingual dictionary induction, and cross-lingual document classification) across five language pairs, we demonstrate that our proposed approach achieves essentially better performances than previous linear methods on all of the three tasks, especially for language pairs with substantial typological difference. Xuefeng Bai 0001, Hailong Cao, Tiejun Zhao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |