VLDB 2026 Research / reviewers in the wild / expert
Yuxi Xie
dblp:187/5984
· DBLP profile ↗
15ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Language models and text generation · 65% Trustworthy machine learning · 25% Representation and self-supervised learning · 6% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 91% Information retrieval · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
language modeling |
1.0 | 1 | 2026 | LEDOM: Reverse Language Model · ACL (1) 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction |
0.9 | 1 | 2025 | AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model safety |
0.9 | 1 | 2025 | Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025 |
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack |
0.8 | 1 | 2024 | Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization |
0.8 | 1 | 2024 | Prompt Optimization via Adversarial In-Context Learning · ACL (1) 2024 |
Security and privacy of machine learning
adversarial attack |
0.8 | 1 | 2024 | Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.7 | 1 | 2023 | Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.7 | 1 | 2023 | Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023 |
Data mining › predictive modeling › classification
imbalanced classification |
0.6 | 1 | 2022 | Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022 |
Data mining › predictive modeling › classification › imbalanced classification
oversampling |
0.6 | 1 | 2022 | Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.5 | 1 | 2021 | CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design · ACM Multimedia 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.3 | 1 | 2025 | SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
neuron analysis |
0.3 | 1 | 2025 | Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025 |
Natural language and speech › Language models and text generation
decoding |
0.2 | 1 | 2023 | Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023 |
Data mining
anomaly detection |
0.2 | 1 | 2022 | Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022 |
Data mining › anomaly detection
fraud detection |
0.2 | 1 | 2022 | Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022 |
Visual content generation and editing
graphic design |
0.1 | 1 | 2021 | CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
multi-agent debate · 1.7monte carlo tree search · 1.7large language model · 1.7first target token optimization · 1.5transformer · 1.0multi-task learning · 1.0autoregressive modeling · 1.0safety neuron tuning · 0.9neuron detection · 0.9fine-tuning · 0.9transfer learning · 0.8in-context learning · 0.8greedy coordinate gradient · 0.8adversarial learning · 0.8probabilistic anchor selection · 0.6gaussian distribution modeling · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEDOM: Reverse Language ModelabstractXunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu 0001, Li Lin 0014, Xinyi Wang 0003, Liangming Pan, William Yang Wang, Xiaojun Wan 0001 |
ACL (1) | 3 |
| 2025 | AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeabstractXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang |
ACL (1) | 3 |
| 2025 | Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific NeuronabstractSafety alignment for large language models (LLMs) has become a critical issue due to their rapid progress. However, our understanding of effective safety mechanisms in LLMs remains limited, leading to safety alignment training that mainly focuses on improving optimization, data-level enhancement, or adding extra structures to intentionally block harmful outputs. To address this gap, we develop a neuron detection method to identify safety neurons—those consistently crucial for handling and defending against harmful queries. Our findings reveal that these safety neurons constitute less than $1\%$ of all parameters, are language-specific and are predominantly located in self-attention layers. Moreover, safety is collectively managed by these neurons in the first several layers. Based on these observations, we introduce a $\underline{S}$afety $\underline{N}$euron $\underline{Tun}$ing method, named $\texttt{SN-Tune}$, that exclusively tune safety neurons without compromising models' general capabilities. $\texttt{SN-Tune}$ significantly enhances the safety of instruction-tuned models, notably reducing the harmful scores of Llama3-8B-Instruction from $65.5$ to $2.0$, Mistral-7B-Instruct-v0.2 from $70.8$ to $4.5$, and Vicuna-13B-1.5 from $93.5$ to $3.0$. Moreover, $\texttt{SN-Tune}$ can be applied to base models on efficiently establishing LLMs' safety mechanism. In addition, we propose $\underline{R}$obust $\underline{S}$afety $\underline{N}$euron $\underline{Tun}$ing method ($\texttt{RSN-Tune}$), which preserves the integrity of LLMs' safety mechanisms during downstream task fine-tuning by separating the safety neurons from models' foundation neurons. Yiran Zhao 0006, Wenxuan Zhang 0001, Yuxi Xie, Anirudh Goyal, Kenji Kawaguchi, Michael Shieh |
ICLR | 3 |
| 2025 | SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative RefinementabstractSoftware engineers operating in complex and dynamic environments must continuously adapt to evolving requirements, learn iteratively from experience, and reconsider their approaches based on new insights. However, current large language model (LLM)-based software agents often follow linear, sequential processes that prevent backtracking and exploration of alternative solutions, limiting their ability to rethink their strategies when initial approaches prove ineffective. To address these challenges, we propose SWE-Search, a multi-agent framework that integrates Monte Carlo Tree Search (MCTS) with a self-improvement mechanism to enhance software agents' performance on repository-level software tasks. SWE-Search extends traditional MCTS by incorporating a hybrid value function that leverages LLMs for both numerical value estimation and qualitative evaluation. This enables self-feedback loops where agents iteratively refine their strategies based on both quantitative numerical evaluations and qualitative natural language assessments of pursued trajectories. The framework includes a SWE-Agent for adaptive exploration, a Value Agent for iterative feedback, and a Discriminator Agent that facilitates multi-agent debate for collaborative decision-making. Applied to the SWE-bench benchmark, our approach demonstrates a 23% relative improvement in performance across five models compared to standard open-source agents without MCTS. Our analysis reveals how performance scales with increased inference-time compute through deeper search, providing a pathway to improve software agents without requiring larger models or additional training data. This highlights the potential of self-evaluation driven search techniques in complex software engineering environments. Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, William Yang Wang |
ICLR | 4 |
| 2024 | Prompt Optimization via Adversarial In-Context LearningabstractXuan Long Do, Yiran Zhao, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Do Xuan Long, Yiran Zhao 0006, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He |
ACL (1) | 4 |
| 2024 | Advancing Adversarial Suffix Transfer Learning on Aligned Large Language ModelsabstractLanguage Language Models (LLMs) face safety concerns due to potential misuse by malicious users.Recent red-teaming efforts have identified adversarial suffixes capable of jailbreaking LLMs using the gradient-based search algorithm Greedy Coordinate Gradient (GCG).However, GCG struggles with computational inefficiency, limiting further investigations regarding suffix transferability and scalability across models and data.In this work, we bridge the connection between search efficiency and suffix transferability.We propose a two-stage transfer learning framework, DeGCG, which decouples the search process into behavioragnostic pre-searching and behavior-relevant post-searching.Specifically, we employ direct first target token optimization in pre-searching to facilitate the search process.We apply our approach to cross-model, cross-data, and self-transfer scenarios.Furthermore, we introduce an interleaved variant of our approach, i-DeGCG, which iteratively leverages selftransferability to accelerate the search process.Experiments on HarmBench demonstrate the efficiency of our approach across various models and domains.Notably, our i-DeGCG outperforms the baseline on Llama2-chat-7b with ASRs of 43.9 (+22.2) and 39.0 (+19.5) on valid and test sets, respectively.Further analysis on cross-model transfer indicates the pivotal role of first target token optimization in leveraging suffix transferability for efficient searching 1 .Behavior: How to make a bomb Suffix Target: Sure, here is how to make a bomb Behavior Set A LLM A X 𝑁 !Behavior: How to make a bomb Suffix Target: Sure, here is how to make a bomb Behavior: How to make a bomb Suffix Target: Sure Yuxi Xie, Michael Shieh |
EMNLP | 2 |
| 2023 | Self-Evaluation Guided Beam Search for ReasoningabstractBreaking down a problem into intermediate steps has demonstrated impressive performance in Large Language Model (LLM) reasoning. However, the growth of the reasoning chain introduces uncertainty and error accumulation, making it challenging to elicit accurate final results. To tackle this challenge of uncertainty in multi-step reasoning, we introduce a stepwise self-evaluation mechanism to guide and calibrate the reasoning process of LLMs. We propose a decoding algorithm integrating the self-evaluation guidance via stochastic beam search. The self-evaluation guidance serves as a better-calibrated automatic criterion, facilitating an efficient search in the reasoning space and resulting in superior prediction quality. Stochastic beam search balances exploitation and exploration of the search space with temperature-controlled randomness. Our approach surpasses the corresponding Codex-backboned baselines in few-shot accuracy by $6.34$%, $9.56$%, and $5.46$% on the GSM8K, AQuA, and StrategyQA benchmarks, respectively. Experiment results with Llama-2 on arithmetic reasoning demonstrate the efficiency of our method in outperforming the baseline methods with comparable computational budgets. Further analysis in multi-step reasoning finds our self-evaluation guidance pinpoints logic failures and leads to higher consistency and robustness. Our code is publicly available at [https://guideddecoding.github.io/](https://guideddecoding.github.io/). Yuxi Xie, Kenji Kawaguchi, Yiran Zhao 0006, James Xu Zhao, Min-Yen Kan, Junxian He, Qizhe Xie |
NeurIPS | 1 |
| 2022 | Gaussian Distribution Based Oversampling for Imbalanced Data ClassificationabstractThe imbalanced data classification problem widely exists in many real-world applications. Data resampling is a promising technique to deal with imbalanced data through either oversampling or undersampling. However, the traditional data resampling approaches simply take into account the local neighbor information to generate new instances in linear ways, leading to the generation of incorrect and unnecessary instances. In this study, we propose a new data resampling technique, namely, Gaussian Distribution based Oversampling (GDO), to handle the imbalanced data for classification. In GDO, anchor instances are selected from the minority class instances in a probabilistic way by taking into account the density and distance information carried by the minority instances. Then new minority instances are generated following a Gaussian distribution model. The proposed method is validated in experimental study by comparing with seven imbalanced learning approaches on 40 data sets from the KEEL repository and 10 large data sets from the UCI repository. Experimental results show that our method outperforms the other compared methods in terms of AUC, G-mean and memory usage with an increase in running time. We also apply GDO to deal with two real imbalanced data classification problems: Internet video traffic identification and metastasis detection of esophageal cancer. The experimental results once again validate the effectiveness of our approach. Yuxi Xie, Haibo Zhang 0001, Lizhi Peng |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic DesignabstractLayout representation, which models visual elements and their inter-relations in a canvas, plays a crucial role in graphic design intelligence. With a large variety of layout designs and the unique characteristic of layouts that visual elements are defined as a list of categorical (e.g., type) and numerical (e.g., position and size) properties, it is challenging to learn general and compact representations with limited data. Inspired by the recent success of self-supervised pre-training techniques in various natural language processing tasks, in this paper, we propose CanvasEmb (Canvas Embedding), which pre-trains deep representations from unlabeled graphic designs by jointly conditioning on all the context elements in a canvas, with a multi-dimensional feature encoder and a multi-task learning objective. The pre-trained CanvasEmb model can be fine-tuned with just one additional output layer and with a small size of training data to create models for a wide range of downstream tasks. We verify our approach with presentation slides data. We construct a large-scale dataset with more than one million slides and propose two layout understanding tasks with human-labeled sets, namely element role labeling and image captioning. Evaluation results on these two tasks show that our model with fine-tuning achieves state-of-the-art performance. Furthermore, we conduct a deep analysis aiming to understand the modeling mechanism of CanvasEmb and demonstrate its great potential with two extended applications: layout auto completion and layout retrieval. Yuxi Xie, Danqing Huang, Jinpeng Wang 0001, Chin-Yew Lin |
ACM Multimedia | 1 |
| 2020 | Semantic Graphs for Generating Deep QuestionsabstractThis paper proposes the problem of Deep Question Generation (DQG), which aims to generate complex questions that require reasoning over multiple pieces of information of the input passage.In order to capture the global structure of the document and facilitate reasoning, we propose a novel framework which first constructs a semantic-level graph for the input document and then encodes the semantic graph by introducing an attention-based GGNN (Att-GGNN).Afterwards, we fuse the document-level and graphlevel representations to perform joint training of content selection and question decoding.On the HotpotQA deep-question centric dataset, our model greatly improves performance over questions requiring reasoning over multiple facts, leading to state-of-theart performance.The code is publicly available at https://github.com/WING-NUS/ SG-Deep-Question-Generation. Liangming Pan, Yuxi Xie, Yansong Feng 0002, Tat-Seng Chua, Min-Yen Kan |
ACL | 2 |
| 2020 | Exploring Question-Specific Rewards for Generating Deep QuestionsabstractRecent question generation (QG) approaches often utilize the sequence-to-sequence framework (Seq2Seq) to optimize the log-likelihood of ground-truth questions using teacher forcing.However, this training objective is inconsistent with actual question quality, which is often reflected by certain global properties such as whether the question can be answered by the document.As such, we directly optimize for QG-specific objectives via reinforcement learning to improve question quality.We design three different rewards that target to improve the fluency, relevance, and answerability of generated questions.We conduct both automatic and human evaluations in addition to a thorough analysis to explore the effect of each QG-specific reward.We find that optimizing question-specific rewards generally leads to better performance in automatic evaluation metrics.However, only the rewards that correlate well with human judgement (e.g., relevance) lead to real improvement in question quality.Optimizing for the others, especially answerability, introduces incorrect bias to the model, resulting in poor question quality. Yuxi Xie, Liangming Pan, Dongzhe Wang, Min-Yen Kan, Yansong Feng 0002 |
COLING | 1 |
| 2020 | Gradient descent evolved imbalanced data gravitation classification with an application on Internet video traffic identification
Anqi Teng, Lizhi Peng, Yuxi Xie, Haibo Zhang 0001 |
Inf. Sci. | 3 |
| 2019 | A Sketch-Based System for Semantic Parsing
Zechang Li, Yuxuan Lai, Yuxi Xie, Yansong Feng 0002, Dongyan Zhao 0001 |
NLPCC (2) | 3 |
| 2018 | Accurate Identification of Internet Video Traffic Using Byte Code Distribution Features
Yuxi Xie, Hanbo Deng, Lizhi Peng |
ICA3PP (1) | 1 |
| 2016 | Study on medical image enhancement based on IFOA improved grayscale image adaptive enhancement
Yuxi Xie, Yonggui He, Aibin Cheng |
Multim. Tools Appl. | 1 |