Xiaohu Du

dblp:273/5574 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-4455-3128ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SEAttack: A self-evolving jailbreak attack to induce toxic responses for non-toxic queries in large language models
Huijun Liu 0003, Shasha Li 0001, Bin Ji 0002, Xiaohu Du, Xiaopeng Li 0006, Jun Ma 0015, Jie Yu 0008
Inf. Process. Manag.4
2025 Multi-Turn Jailbreaking Large Language Models via Attention Shifting
abstract
Large Language Models (LLMs) have achieved significant performance in various natural language processing tasks but also pose safety and ethical threats, thus requiring red teaming and alignment processes to bolster their safety. To effectively exploit these aligned LLMs, recent studies have introduced jailbreak attacks based on multi-turn dialogues. These attacks aim to prompt LLMs to generate harmful or biased content by guiding them through contextual content. However, the underlying reasons for the effectiveness of multi-turn jailbreaks remain unclear. Existing attacks often focus on optimizing queries and escalating toxicity to construct dialogues, lacking a thorough analysis of the inherent vulnerabilities of LLMs. In this paper, we first conduct an in-depth analysis of the differences between single-turn and multi-turn jailbreaks and find that successful multi-turn jailbreaks can effectively disperse the attention of LLMs on keywords associated with harmful behaviors, especially in historical responses. Based on this, we propose ASJA, a new multi-turn jailbreak approach by shifting the attention of LLMs, specifically by iteratively fabricating the dialogue history through a genetic algorithm to induce LLMs to generate harmful content. Extensive experiments on three LLMs and two datasets show that our approach surpasses existing approaches in jailbreak effectiveness, the stealth of jailbreak prompts, and attack efficiency. Our work emphasizes the importance of enhancing the robustness of LLMs' attention mechanism in multi-turn dialogue scenarios for a better defense strategy.
Xiaohu Du, Fan Mo 0004, Ming Wen 0001, Tu Gu, Huadi Zheng, Hai Jin 0001, Jie Shi 0005
AAAI1
2025 Refusal-Aware Red Teaming: Exposing Inconsistency in Safety Evaluations
abstract
The responsible deployment of Large Language Models (LLMs) necessitates rigorous safety evaluations.However, a critical challenge arises from inconsistencies between an LLM's internal refusal decisions and external safety assessments, hindering effective validation.This paper introduces the concept of the 'refusal gap' to formally define these discrepancies.We then present a novel, refusal-aware red teaming framework designed to automatically generate test cases that expose such gaps.Our framework employs 'refusal probes', which leverage the target model's hidden states, to detect internal model refusals.These are subsequently contrasted with judgments from an external safety evaluator.The identified discrepancy serves as a signal to guide a red-teaming model in crafting test cases that maximize this refusal gap.To further enhance test case diversity and address challenges related to sparse rewards, we introduce a hierarchical, curiositydriven mechanism that incentivizes both refusal gap maximization and broad topic exploration.Empirical results demonstrate that our method significantly outperforms existing reinforcement learning-based approaches in generating diverse test cases and achieves a substantially higher discovery rate of refusal gaps.
Xiaohu Du, Xiaotian Zou, Chongyang Zhao 0004, Xiaohui Kuang
EMNLP2
2023 An Extensive Study on Adversarial Attack against Pre-trained Models of Code
abstract
Transformer-based pre-trained models of code (PTMC) have been widely utilized and have achieved state-of-the-art performance in many mission-critical applications. However, they can be vulnerable to adversarial attacks through identifier substitution or coding style transformation, which can significantly degrade accuracy and may further incur security concerns. Although several approaches have been proposed to generate adversarial examples for PTMC, the effectiveness and efficiency of such approaches, especially on different code intelligence tasks, has not been well understood. To bridge this gap, this study systematically analyzes five state-of-the-art adversarial attack approaches from three perspectives: effectiveness, efficiency, and the quality of generated examples. The results show that none of the five approaches balances all these perspectives. Particularly, approaches with a high attack success rate tend to be time-consuming; the adversarial code they generate often lack naturalness, and vice versa. To address this limitation, we explore the impact of perturbing identifiers under different contexts and find that identifier substitution within for and if statements is the most effective. Based on these findings, we propose a new approach that prioritizes different types of statements for various tasks and further utilizes beam search to generate adversarial examples. Evaluation results show that it outperforms the state-of-the-art ALERT in terms of both effectiveness and efficiency while preserving the naturalness of the generated adversarial examples.
Xiaohu Du, Ming Wen 0001, Zichao Wei, Shangwen Wang, Hai Jin 0001
ESEC/SIGSOFT FSE1
2023 Understanding the Bug Characteristics and Fix Strategies of Federated Learning Systems
abstract
Federated learning (FL) is an emerging machine learning paradigm that aims to address the problem of isolated data islands. To preserve privacy, FL allows machine learning models and deep neural networks to be trained from decentralized data kept privately at individual devices. FL has been increasingly adopted in missioncritical fields such as finance and healthcare. However, bugs in FL systems are inevitable and may result in catastrophic consequences such as financial loss, inappropriate medical decision, and violation of data privacy ordinance. While many recent studies were conducted to understand the bugs in machine learning systems, there is no existing study to characterize the bugs arising from the unique nature of FL systems. To fill the gap, we collected 395 real bugs from six popular FL frameworks (Tensorflow Federated, PySyft, FATE, Flower, PaddleFL, and Fedlearner) in GitHub and StackOverflow, and then manually analyzed their symptoms and impacts, prone stages, root causes, and fix strategies. Furthermore, we report a series of findings and actionable implications that can potentially facilitate the detection of FL bugs.
Xiaohu Du, Xiao Chen 0026, Jialun Cao, Ming Wen 0001, Shing-Chi Cheung, Hai Jin 0001
ESEC/SIGSOFT FSE1
2021 Combating Word-level Adversarial Text with Robust Adversarial Training
abstract
NLP models perform well on many tasks, but they are also easy to be fooled by adversarial examples. A small perturbation can change the output of the deep neural network model. This kind of perturbation is hard to be perceived by humans, especially adversarial examples generated by word-level adversarial attack. Character-level adversarial attack can be defended by grammar detection and word recognition. The existing word-level textual adversarial attacks are based on synonym replacement, so adversarial texts usually have correct grammar and semantics. The defense of word-level adversarial attack is more challenging. In this paper, we propose a framework which is called Robust Adversarial Training (RAT) to defend against word-level adversarial attacks. RAT enhances the model by combining adversarial training and data perturbation during training. Our experiments on two datasets show that the model based on our framework can effectively defend against word-level adversarial attacks. Compared with the existing defense methods, the model trained under RAT has a higher defense success rate on 1000 adversarial examples. In addition, the accuracy of our model on the standard testing set is also better than the existing defense methods, and the accuracy is very close to or even higher than that of the standard model.
Xiaohu Du, Jie Yu 0008, Shasha Li 0001, Zibo Yi, Jun Ma 0015
IJCNN1