Yuxi Xie

dblp:187/5984 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 65% Trustworthy machine learning · 25% Representation and self-supervised learning · 6%
Databases, data mining, and information retrieval
2 papers
Data mining · 91% Information retrieval · 9%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
language modeling
1.012026
LEDOM: Reverse Language Model · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction
0.912025
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack
0.812024
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models · EMNLP 2024
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization
0.812024
Prompt Optimization via Adversarial In-Context Learning · ACL (1) 2024
Security and privacy of machine learning
adversarial attack
0.812024
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models · EMNLP 2024
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.712023
Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model reasoning
0.712023
Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023
Data mining › predictive modeling › classification
imbalanced classification
0.612022
Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022
Data mining › predictive modeling › classification › imbalanced classification
oversampling
0.612022
Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.512021
CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design · ACM Multimedia 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.312025
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
neuron analysis
0.312025
Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron · ICLR 2025
Natural language and speech › Language models and text generation
decoding
0.212023
Self-Evaluation Guided Beam Search for Reasoning · NeurIPS 2023
Data mining
anomaly detection
0.212022
Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022
Data mining › anomaly detection
fraud detection
0.212022
Gaussian Distribution Based Oversampling for Imbalanced Data Classification · IEEE Trans. Knowl. Data Eng. 2022
Visual content generation and editing
graphic design
0.112021
CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design · ACM Multimedia 2021

Methods — techniques the papers use, named apart from their topics

multi-agent debate · 1.7monte carlo tree search · 1.7large language model · 1.7first target token optimization · 1.5transformer · 1.0multi-task learning · 1.0autoregressive modeling · 1.0safety neuron tuning · 0.9neuron detection · 0.9fine-tuning · 0.9transfer learning · 0.8in-context learning · 0.8greedy coordinate gradient · 0.8adversarial learning · 0.8probabilistic anchor selection · 0.6gaussian distribution modeling · 0.6
YearPublicationVenuePosition
2026 LEDOM: Reverse Language Model
abstract
Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu 0001, Li Lin 0014, Xinyi Wang 0003, Liangming Pan, William Yang Wang, Xiaojun Wan 0001
ACL (1)3
2025 AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
abstract
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao, Yubo Ma, Mingzhe Du, Rui Mao, Anh Tuan Luu, William Yang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou, Shuai Zhao 0007, Yubo Ma, Mingzhe Du, Rui Mao 0010, Anh Tuan Luu, William Yang Wang
ACL (1)3
2025 Understanding and Enhancing Safety Mechanisms of LLMs via Safety-Specific Neuron
abstract
Safety alignment for large language models (LLMs) has become a critical issue due to their rapid progress. However, our understanding of effective safety mechanisms in LLMs remains limited, leading to safety alignment training that mainly focuses on improving optimization, data-level enhancement, or adding extra structures to intentionally block harmful outputs. To address this gap, we develop a neuron detection method to identify safety neurons—those consistently crucial for handling and defending against harmful queries. Our findings reveal that these safety neurons constitute less than $1\%$ of all parameters, are language-specific and are predominantly located in self-attention layers. Moreover, safety is collectively managed by these neurons in the first several layers. Based on these observations, we introduce a $\underline{S}$afety $\underline{N}$euron $\underline{Tun}$ing method, named $\texttt{SN-Tune}$, that exclusively tune safety neurons without compromising models' general capabilities. $\texttt{SN-Tune}$ significantly enhances the safety of instruction-tuned models, notably reducing the harmful scores of Llama3-8B-Instruction from $65.5$ to $2.0$, Mistral-7B-Instruct-v0.2 from $70.8$ to $4.5$, and Vicuna-13B-1.5 from $93.5$ to $3.0$. Moreover, $\texttt{SN-Tune}$ can be applied to base models on efficiently establishing LLMs' safety mechanism. In addition, we propose $\underline{R}$obust $\underline{S}$afety $\underline{N}$euron $\underline{Tun}$ing method ($\texttt{RSN-Tune}$), which preserves the integrity of LLMs' safety mechanisms during downstream task fine-tuning by separating the safety neurons from models' foundation neurons.
Yiran Zhao 0006, Wenxuan Zhang 0001, Yuxi Xie, Anirudh Goyal, Kenji Kawaguchi, Michael Shieh
ICLR3
2025 SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
abstract
Software engineers operating in complex and dynamic environments must continuously adapt to evolving requirements, learn iteratively from experience, and reconsider their approaches based on new insights. However, current large language model (LLM)-based software agents often follow linear, sequential processes that prevent backtracking and exploration of alternative solutions, limiting their ability to rethink their strategies when initial approaches prove ineffective. To address these challenges, we propose SWE-Search, a multi-agent framework that integrates Monte Carlo Tree Search (MCTS) with a self-improvement mechanism to enhance software agents' performance on repository-level software tasks. SWE-Search extends traditional MCTS by incorporating a hybrid value function that leverages LLMs for both numerical value estimation and qualitative evaluation. This enables self-feedback loops where agents iteratively refine their strategies based on both quantitative numerical evaluations and qualitative natural language assessments of pursued trajectories. The framework includes a SWE-Agent for adaptive exploration, a Value Agent for iterative feedback, and a Discriminator Agent that facilitates multi-agent debate for collaborative decision-making. Applied to the SWE-bench benchmark, our approach demonstrates a 23% relative improvement in performance across five models compared to standard open-source agents without MCTS. Our analysis reveals how performance scales with increased inference-time compute through deeper search, providing a pathway to improve software agents without requiring larger models or additional training data. This highlights the potential of self-evaluation driven search techniques in complex software engineering environments.
Antonis Antoniades, Albert Örwall, Kexun Zhang, Yuxi Xie, Anirudh Goyal, William Yang Wang
ICLR4
2024 Prompt Optimization via Adversarial In-Context Learning
abstract
Xuan Long Do, Yiran Zhao, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Do Xuan Long, Yiran Zhao 0006, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He
ACL (1)4
2024 Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
abstract
Language Language Models (LLMs) face safety concerns due to potential misuse by malicious users.Recent red-teaming efforts have identified adversarial suffixes capable of jailbreaking LLMs using the gradient-based search algorithm Greedy Coordinate Gradient (GCG).However, GCG struggles with computational inefficiency, limiting further investigations regarding suffix transferability and scalability across models and data.In this work, we bridge the connection between search efficiency and suffix transferability.We propose a two-stage transfer learning framework, DeGCG, which decouples the search process into behavioragnostic pre-searching and behavior-relevant post-searching.Specifically, we employ direct first target token optimization in pre-searching to facilitate the search process.We apply our approach to cross-model, cross-data, and self-transfer scenarios.Furthermore, we introduce an interleaved variant of our approach, i-DeGCG, which iteratively leverages selftransferability to accelerate the search process.Experiments on HarmBench demonstrate the efficiency of our approach across various models and domains.Notably, our i-DeGCG outperforms the baseline on Llama2-chat-7b with ASRs of 43.9 (+22.2) and 39.0 (+19.5) on valid and test sets, respectively.Further analysis on cross-model transfer indicates the pivotal role of first target token optimization in leveraging suffix transferability for efficient searching 1 .Behavior: How to make a bomb Suffix Target: Sure, here is how to make a bomb Behavior Set A LLM A X 𝑁 !Behavior: How to make a bomb Suffix Target: Sure, here is how to make a bomb Behavior: How to make a bomb Suffix Target: Sure
Yuxi Xie, Michael Shieh
EMNLP2
2023 Self-Evaluation Guided Beam Search for Reasoning
abstract
Breaking down a problem into intermediate steps has demonstrated impressive performance in Large Language Model (LLM) reasoning. However, the growth of the reasoning chain introduces uncertainty and error accumulation, making it challenging to elicit accurate final results. To tackle this challenge of uncertainty in multi-step reasoning, we introduce a stepwise self-evaluation mechanism to guide and calibrate the reasoning process of LLMs. We propose a decoding algorithm integrating the self-evaluation guidance via stochastic beam search. The self-evaluation guidance serves as a better-calibrated automatic criterion, facilitating an efficient search in the reasoning space and resulting in superior prediction quality. Stochastic beam search balances exploitation and exploration of the search space with temperature-controlled randomness. Our approach surpasses the corresponding Codex-backboned baselines in few-shot accuracy by $6.34$%, $9.56$%, and $5.46$% on the GSM8K, AQuA, and StrategyQA benchmarks, respectively. Experiment results with Llama-2 on arithmetic reasoning demonstrate the efficiency of our method in outperforming the baseline methods with comparable computational budgets. Further analysis in multi-step reasoning finds our self-evaluation guidance pinpoints logic failures and leads to higher consistency and robustness. Our code is publicly available at [https://guideddecoding.github.io/](https://guideddecoding.github.io/).
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao 0006, James Xu Zhao, Min-Yen Kan, Junxian He, Qizhe Xie
NeurIPS1
2022 Gaussian Distribution Based Oversampling for Imbalanced Data Classification
abstract
The imbalanced data classification problem widely exists in many real-world applications. Data resampling is a promising technique to deal with imbalanced data through either oversampling or undersampling. However, the traditional data resampling approaches simply take into account the local neighbor information to generate new instances in linear ways, leading to the generation of incorrect and unnecessary instances. In this study, we propose a new data resampling technique, namely, Gaussian Distribution based Oversampling (GDO), to handle the imbalanced data for classification. In GDO, anchor instances are selected from the minority class instances in a probabilistic way by taking into account the density and distance information carried by the minority instances. Then new minority instances are generated following a Gaussian distribution model. The proposed method is validated in experimental study by comparing with seven imbalanced learning approaches on 40 data sets from the KEEL repository and 10 large data sets from the UCI repository. Experimental results show that our method outperforms the other compared methods in terms of AUC, G-mean and memory usage with an increase in running time. We also apply GDO to deal with two real imbalanced data classification problems: Internet video traffic identification and metastasis detection of esophageal cancer. The experimental results once again validate the effectiveness of our approach.
Yuxi Xie, Haibo Zhang 0001, Lizhi Peng
IEEE Trans. Knowl. Data Eng.1
2021 CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design
abstract
Layout representation, which models visual elements and their inter-relations in a canvas, plays a crucial role in graphic design intelligence. With a large variety of layout designs and the unique characteristic of layouts that visual elements are defined as a list of categorical (e.g., type) and numerical (e.g., position and size) properties, it is challenging to learn general and compact representations with limited data. Inspired by the recent success of self-supervised pre-training techniques in various natural language processing tasks, in this paper, we propose CanvasEmb (Canvas Embedding), which pre-trains deep representations from unlabeled graphic designs by jointly conditioning on all the context elements in a canvas, with a multi-dimensional feature encoder and a multi-task learning objective. The pre-trained CanvasEmb model can be fine-tuned with just one additional output layer and with a small size of training data to create models for a wide range of downstream tasks. We verify our approach with presentation slides data. We construct a large-scale dataset with more than one million slides and propose two layout understanding tasks with human-labeled sets, namely element role labeling and image captioning. Evaluation results on these two tasks show that our model with fine-tuning achieves state-of-the-art performance. Furthermore, we conduct a deep analysis aiming to understand the modeling mechanism of CanvasEmb and demonstrate its great potential with two extended applications: layout auto completion and layout retrieval.
Yuxi Xie, Danqing Huang, Jinpeng Wang 0001, Chin-Yew Lin
ACM Multimedia1
2020 Semantic Graphs for Generating Deep Questions
abstract
This paper proposes the problem of Deep Question Generation (DQG), which aims to generate complex questions that require reasoning over multiple pieces of information of the input passage.In order to capture the global structure of the document and facilitate reasoning, we propose a novel framework which first constructs a semantic-level graph for the input document and then encodes the semantic graph by introducing an attention-based GGNN (Att-GGNN).Afterwards, we fuse the document-level and graphlevel representations to perform joint training of content selection and question decoding.On the HotpotQA deep-question centric dataset, our model greatly improves performance over questions requiring reasoning over multiple facts, leading to state-of-theart performance.The code is publicly available at https://github.com/WING-NUS/ SG-Deep-Question-Generation.
Liangming Pan, Yuxi Xie, Yansong Feng 0002, Tat-Seng Chua, Min-Yen Kan
ACL2
2020 Exploring Question-Specific Rewards for Generating Deep Questions
abstract
Recent question generation (QG) approaches often utilize the sequence-to-sequence framework (Seq2Seq) to optimize the log-likelihood of ground-truth questions using teacher forcing.However, this training objective is inconsistent with actual question quality, which is often reflected by certain global properties such as whether the question can be answered by the document.As such, we directly optimize for QG-specific objectives via reinforcement learning to improve question quality.We design three different rewards that target to improve the fluency, relevance, and answerability of generated questions.We conduct both automatic and human evaluations in addition to a thorough analysis to explore the effect of each QG-specific reward.We find that optimizing question-specific rewards generally leads to better performance in automatic evaluation metrics.However, only the rewards that correlate well with human judgement (e.g., relevance) lead to real improvement in question quality.Optimizing for the others, especially answerability, introduces incorrect bias to the model, resulting in poor question quality.
Yuxi Xie, Liangming Pan, Dongzhe Wang, Min-Yen Kan, Yansong Feng 0002
COLING1
2020 Gradient descent evolved imbalanced data gravitation classification with an application on Internet video traffic identification
Anqi Teng, Lizhi Peng, Yuxi Xie, Haibo Zhang 0001
Inf. Sci.3
2019 A Sketch-Based System for Semantic Parsing
Zechang Li, Yuxuan Lai, Yuxi Xie, Yansong Feng 0002, Dongyan Zhao 0001
NLPCC (2)3
2018 Accurate Identification of Internet Video Traffic Using Byte Code Distribution Features
Yuxi Xie, Hanbo Deng, Lizhi Peng
ICA3PP (1)1
2016 Study on medical image enhancement based on IFOA improved grayscale image adaptive enhancement
Yuxi Xie, Yonggui He, Aibin Cheng
Multim. Tools Appl.1