Zeyu Qin

dblp:271/5778 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-1733-7892ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Preserving Diversity in Supervised Fine-Tuning of Large Language Models
abstract
Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However, CE maximizes the likelihood of observed data without accounting for alternative possibilities. As such, CE usually leads to reduced diversity in the model's outputs, which hinders further development that requires sampling to explore better responses. To address this limitation, this paper introduces a new game-theoretic formulation for SFT. In this framework, an auxiliary variable is introduced to regulate the learning process. We prove that the proposed game-theoretic approach connects to the problem of reverse KL minimization with entropy regularization. This regularization prevents over-memorization of training data and promotes output diversity. To implement this framework, we develop GEM, a new training algorithm that is computationally efficient as CE by leveraging some unique properties of LLMs. Empirical studies of pre-trained models from 3B to 70B parameters show that GEM achieves comparable downstream performance to CE while significantly enhancing output diversity. This increased diversity translates to performance gains in test-time compute scaling for chat and code generation tasks. Moreover, we observe that preserving output diversity has the added benefit of mitigating forgetting, as maintaining diverse outputs encourages models to retain pre-trained knowledge throughout the training process.
Ziniu Li, Congliang Chen, Tian Xu 0003, Zeyu Qin, Jiancong Xiao, Zhi-Quan Luo, Ruoyu Sun 0001
ICLR4
2025 Safety Reasoning with Guidelines
abstract
Training safe LLMs remains a critical challenge. The most widely used method, Refusal Training (RT), struggles to generalize against various Out-of-Distribution (OOD) jailbreaking attacks. Although various advanced methods have been proposed to address this issue, we instead question whether OOD attacks inherently surpass the capability of vanilla RT. Evaluations using Best-of-N (BoN) reveal significant safety improvements as N increases, indicating models possess adequate latent safety knowledge but RT fails to consistently elicit it under OOD scenarios. Further domain adaptation analysis reveals that direct RT causes reliance on superficial shortcuts, resulting in non-generalizable representation mappings. Inspired by our findings, we propose training model to perform safety reasoning for each query. Specifically, we synthesize reasoning supervision aligned with specified guidelines that reflect diverse perspectives on safety knowledge. This encourages model to engage in deeper reasoning, explicitly eliciting and utilizing latent safety knowledge for each query. Extensive experiments show that our method significantly improves model generalization against OOD attacks.
Haoyu Wang 0018, Zeyu Qin, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao, Minhao Cheng
ICML2
2025 Lifelong Safety Alignment for Language Models
abstract
LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for *unseen* attacks that may arise during deployment. To address this, we propose a **lifelong safety alignment** framework that enables LLMs to continuously adapt to new and evolving jailbreaking strategies. Our framework introduces a competitive setup between two components: a **Meta-Attacker**, trained to actively discover novel jailbreaking strategies, and a **Defender**, trained to resist them. To effectively warm up the Meta-Attacker, we first leverage the GPT-4o API to extract key insights from a large collection of jailbreak-related research papers. Through iterative training, the first iteration Meta-Attacker achieves a 73% attack success rate (ASR) on RR and a 57% transfer ASR on LAT using only *single-turn* attacks. Meanwhile, the Defender progressively improves its robustness and ultimately reduces the Meta-Attacker's success rate to just 7%, enabling safer and more reliable deployment of LLMs in open-ended environments.
Haoyu Wang 0018, Zeyu Qin, Xueqian Wang 0001, Tianyu Pang
NeurIPS3
2025 RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion Models
abstract
Preserving intellectual property (IP) within a pre-trained diffusion model is critical for protecting the model's copyright and preventing unauthorized model deployment. In this regard, model watermarking is a common practice for IP protection that embeds traceable information within models and allows for further verification. Nevertheless, existing watermarking schemes often face challenges due to their vulnerability to fine-tuning, limiting their practical application in general pre-training and fine-tuning paradigms. Inspired by using mode connectivity to analyze model performance between a pair of connected models, we investigate watermark vulnerability by leveraging Linear Mode Connectivity (LMC) as a proxy to analyze the fine-tuning dynamics of watermark performance. Our results show that existing watermarked models tend to converge to sharp minima in the loss landscape, thus making them vulnerable to fine-tuning. To tackle this challenge, we propose **RoMa**, a **Ro**bust **M**odel w**a**termarking scheme that improves the robustness of watermarks against fine-tuning. Specifically, RoMa decomposes watermarking into two components, including *Embedding Functionality*, which preserves reliable watermark detection capability, and *Path-specific Smoothness*, which enhances the smoothness along the watermark-connected path to improve robustness. Extensive experiments on benchmark datasets MS-COCO-2017 and CUB-200-2011 demonstrate that RoMa significantly improves watermark robustness against fine-tuning while maintaining generation quality, outperforming baselines. The code is available at [https://github.com/xiekks/RoMa](https://github.com/xiekks/RoMa).
Yingsha Xie, Zeyu Qin, Fei Ma 0006, Li Shen 0008, F. Richard Yu, Xiaochun Cao
NeurIPS3
2024 Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
abstract
Backdoor attacks pose a significant threat to Deep Neural Networks (DNNs) as they allow attackers to manipulate model predictions with backdoor triggers. To address these security vulnerabilities, various backdoor purification methods have been proposed to purify compromised models. Typically, these purified models exhibit low Attack Success Rates (ASR), rendering them resistant to backdoored inputs. However, \textit{Does achieving a low ASR through current safety purification methods truly eliminate learned backdoor features from the pretraining phase?} In this paper, we provide an affirmative answer to this question by thoroughly investigating the \textit{Post-Purification Robustness} of current backdoor purification methods. We find that current safety purification methods are vulnerable to the rapid re-learning of backdoor behavior, even when further fine-tuning of purified models is performed using a very small number of poisoned samples. Based on this, we further propose the practical Query-based Reactivation Attack (QRA) which could effectively reactivate the backdoor by merely querying purified models. We find the failure to achieve satisfactory post-purification robustness stems from the insufficient deviation of purified models from the backdoored model along the backdoor-connected path. To improve the post-purification robustness, we propose a straightforward tuning defense, Path-Aware Minimization (PAM), which promotes deviation along backdoor-connected paths with extra model updates. Extensive experiments demonstrate that PAM significantly improves post-purification robustness while maintaining a good clean accuracy and low ASR. Our work provides a new perspective on understanding the effectiveness of backdoor safety tuning and highlights the importance of faithfully assessing the model's safety.
Zeyu Qin, Nevin Lianwen Zhang, Li Shen 0008, Minhao Cheng
NeurIPS2
2023 Beyond Factuality: A Comprehensive Evaluation of Large Language Models as Knowledge Generators
abstract
Large language models (LLMs) outperform information retrieval techniques for downstream knowledge-intensive tasks when being prompted to generate world knowledge.Yet, community concerns abound regarding the factuality and potential implications of using this uncensored knowledge.In light of this, we introduce CONNER, a COmpreheNsive kNowledge Evaluation fRamework, designed to systematically and automatically evaluate generated knowledge from six important perspectives -Factuality, Relevance, Coherence, Informativeness, Helpfulness and Validity.We conduct an extensive empirical analysis of the generated knowledge from three different types of LLMs on two widely-studied knowledge-intensive tasks, i.e., open-domain question answering and knowledge-grounded dialogue.Surprisingly, our study reveals that the factuality of generated knowledge, even if lower, does not significantly hinder downstream tasks.Instead, the relevance and coherence of the outputs are more important than small factual mistakes.Further, we show how to use CONNER to improve knowledge-intensive tasks by designing two strategies: Prompt Engineering and Knowledge Selection.Our evaluation code and LLMgenerated knowledge with human annotations will be released 1 to facilitate future research.
Liang Chen 0001, Yang Deng 0002, Yatao Bian, Zeyu Qin, Bingzhe Wu, Tat-Seng Chua, Kam-Fai Wong
EMNLP4
2023 Revisiting Personalized Federated Learning: Robustness Against Backdoor Attacks
abstract
In this work, besides improving prediction accuracy, we study whether personalization could bring robustness benefits to backdoor attacks. We conduct the first study of backdoor attacks in the pFL framework, testing 4 widely used backdoor attacks against 6 pFL methods on benchmark datasets FEMNIST and CIFAR-10, a total of 600 experiments. The study shows that pFL methods with partial model-sharing can significantly boost robustness against backdoor attacks. In contrast, pFL methods with full model-sharing do not show robustness. To analyze the reasons for varying robustness performances, we provide comprehensive ablation studies on different pFL methods. Based on our findings, we further propose a lightweight defense method, Simple-Tuning, which empirically improves defense performance against backdoor attacks. We believe that our work could provide both guidance for pFL application in terms of its robustness and offer valuable insights to design more robust FL methods in the future. We open-source our code to establish the first benchmark for black-box backdoor attacks in pFL: https://github.com/alibaba/FederatedScope/tree/backdoor-bench.
Zeyu Qin, Liuyi Yao, Daoyuan Chen, Yaliang Li, Bolin Ding, Minhao Cheng
KDD1
2023 Imitation Learning from Imperfection: Theoretical Justifications and Algorithms
abstract
Imitation learning (IL) algorithms excel in acquiring high-quality policies from expert data for sequential decision-making tasks. But, their effectiveness is hampered when faced with limited expert data. To tackle this challenge, a novel framework called (offline) IL with supplementary data has been proposed, which enhances learning by incorporating an additional yet imperfect dataset obtained inexpensively from sub-optimal policies. Nonetheless, learning becomes challenging due to the potential inclusion of out-of-expert-distribution samples. In this work, we propose a mathematical formalization of this framework, uncovering its limitations. Our theoretical analysis reveals that a naive approach—applying the behavioral cloning (BC) algorithm concept to the combined set of expert and supplementary data—may fall short of vanilla BC, which solely relies on expert data. This deficiency arises due to the distribution shift between the two data sources. To address this issue, we propose a new importance-sampling-based technique for selecting data within the expert distribution. We prove that the proposed method eliminates the gap of the naive approach, highlighting its efficacy when handling imperfect data. Empirical studies demonstrate that our method outperforms previous state-of-the-art methods in tasks including robotic locomotion control, Atari video games, and image classification. Overall, our work underscores the potential of improving IL by leveraging diverse data sources through effective data selection.
Ziniu Li, Tian Xu 0003, Zeyu Qin, Yang Yu 0001, Zhi-Quan Luo
NeurIPS3
2023 Towards Stable Backdoor Purification through Feature Shift Tuning
abstract
It has been widely observed that deep neural networks (DNN) are vulnerable to backdoor attacks where attackers could manipulate the model behavior maliciously by tampering with a small set of training samples. Although a line of defense methods is proposed to mitigate this threat, they either require complicated modifications to the training process or heavily rely on the specific model architecture, which makes them hard to deploy into real-world applications. Therefore, in this paper, we instead start with fine-tuning, one of the most common and easy-to-deploy backdoor defenses, through comprehensive evaluations against diverse attack scenarios. Observations made through initial experiments show that in contrast to the promising defensive results on high poisoning rates, vanilla tuning methods completely fail at low poisoning rate scenarios. Our analysis shows that with the low poisoning rate, the entanglement between backdoor and clean features undermines the effect of tuning-based defenses. Therefore, it is necessary to disentangle the backdoor and clean features in order to improve backdoor purification. To address this, we introduce Feature Shift Tuning (FST), a method for tuning-based backdoor purification. Specifically, FST encourages feature shifts by actively deviating the classifier weights from the originally compromised weights. Extensive experiments demonstrate that our FST provides consistently stable performance under different attack settings. Without complex parameter adjustments, FST also achieves much lower tuning costs, only $10$ epochs. Our codes are available at https://github.com/AISafety-HKUST/stable_backdoor_purification.
Zeyu Qin, Li Shen 0008, Minhao Cheng
NeurIPS2
2023 Multi-Agent Reinforcement Learning Aided Computation Offloading in Aerial Computing for the Internet-of-Things
abstract
LEO satellite networks have become a necessary supplement to terrestrial networks aiming to provide worldwide, ubiquitous connectivity, especially in complicated areas (e.g., mountains, oceans, and disaster areas) where terrestrial network infrastructures are typically sparingly distributed or unavailable. However, the increasing computation-intensive Internet-of-Things (IoT) applications (e.g., real-time remote monitoring, intelligent transportation) require not only efficient and reliable communication but also massive computing capabilities. Constrained by the battery and computing resources, the computing tasks and data of applications have to be transmitted to remote cloud servers. This bandwidth limitation and high transmission delay in LEO networks will reduce the quality-of-service (QoS) of IoT applications. Recently, the combination of LEO networks and edge computing (i.e., Satellite Mobile Edge Computing, SMEC) offers significant opportunities to address these problems. The IoT devices can directly get the computing resources directly from satellites rather than remote servers, thus avoiding long-distance transmission. Considering the resource constraints on satellites, offloading policy plays a crucial role in whole system performance. In this paper, we design a hybrid offloading architecture, which applies a centralized training and distributed execution framework. Also, we propose a multi-agent actor-critic reinforcement learning algorithm, where a centralized “critic” is augmented with the global network state to ease the training procedure of distributed user equipments (UE) by evaluating the benefits of their decisions, while the UEs can adjust their policies according to the critic’s evaluation and choose their own decisions relying on their observations.
Zeyu Qin, Haipeng Yao, Tianle Mai, Di Wu 0001, Song Guo 0001
IEEE Trans. Serv. Comput.1
2022 Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation
abstract
Deep neural networks (DNNs) have been shown to be vulnerable to adversarial examples, which can produce erroneous predictions by injecting imperceptible perturbations. In this work, we study the transferability of adversarial examples, which is significant due to its threat to real-world applications where model architecture or parameters are usually unknown. Many existing works reveal that the adversarial examples are likely to overfit the surrogate model that they are generated from, limiting its transfer attack performance against different target models. To mitigate the overfitting of the surrogate model, we propose a novel attack method, dubbed reverse adversarial perturbation (RAP). Specifically, instead of minimizing the loss of a single adversarial point, we advocate seeking adversarial example located at a region with unified low loss value, by injecting the worst-case perturbation (the reverse adversarial perturbation) for each step of the optimization procedure. The adversarial attack with RAP is formulated as a min-max bi-level optimization problem. By integrating RAP into the iterative process for attacks, our method can find more stable adversarial examples which are less sensitive to the changes of decision boundary, mitigating the overfitting of the surrogate model. Comprehensive experimental comparisons demonstrate that RAP can significantly boost adversarial transferability. Furthermore, RAP can be naturally combined with many existing black-box attack techniques, to further boost the transferability. When attacking a real-world image recognition system, Google Cloud Vision API, we obtain 22% performance improvement of targeted attacks over the compared method. Our codes are available at https://github.com/SCLBD/TransferattackRAP.
Zeyu Qin, Yanbo Fan, Li Shen 0008, Yong Zhang 0034, Jue Wang 0001, Baoyuan Wu
NeurIPS1
2022 A multidomain virtual network embedding algorithm based on multiobjective optimization for Internet of Drones architecture in Industry 4.0
abstract
Summary Unmanned aerial vehicle (UAV) has a broad application prospect in the future, especially in the Industry 4.0. The development of Internet of Drones (IoD) makes UAV operation more autonomous. Network virtualization technology is a promising technology to support IoD, so the allocation of virtual resources becomes a crucial issue in IoD. How to rationally allocate potential material resources has become an urgent problem to be solved. The main work of this paper is presented as follows: (a) In order to improve the optimization performance and reduce the computation time, we propose a multidomain virtual network embedding algorithm (MP‐VNE) adopting the centralized hierarchical multidomain architecture. The proposed algorithm can avoid the local optimum through incorporating the genetic variation factor into the traditional particle swarm optimization process. (b) In order to simplify the multiobjective optimization problem, we transform the multiobjective problem into a single‐objective problem through weighted summation method. The results prove that the proposed algorithm can rapidly converge to the optimal solution. (c) In order to reduce the mapping cost, we propose an algorithm for selecting candidate nodes based on the estimated mapping cost. Each physical domain calculates the estimated mapping cost of all nodes according to the formula of the estimated mapping cost, and chooses the node with the lowest estimated mapping cost as the candidate node. The simulation results show that the proposed MP‐VNE algorithm has better performance than MC‐VNM, LID‐VNE, and other algorithms in terms of delay, cost and comprehensive indicators.
Peiying Zhang 0001, Chao Wang 0093, Zeyu Qin, Haotong Cao
Softw. Pract. Exp.3
2021 Collaborate Q-learning Aided Load Balance in Satellites Communications
abstract
In recent years, satellite communications have played an increasingly important role in daily life. With the explosive growth of new businesses, the expectations for the performance and reliability of satellite communications are greater than ever. However, due to the unique characteristics of satellite node (e.g., fast transmission speed, saturation of resources), it brings unprecedented challenges for load balance in multiple satellite paths. In this paper, to overcome this issue, we proposed a multi-agent reinforcement learning aided load balance architecture. We formulate the load balance in satellites communications as a partially observable Markov decision process (POMDP). Besides, we adopt a multi-agent reinforcement algorithm named Collaborate Q-learning (CollaQ) in our architecture. In addition, some stimulation are performed to evaluate the correctness of our architecture and algorithm.
Haipeng Yao, Zeyu Qin, Tianle Mai
IWCMC3
2021 Random Noise Defense Against Query-Based Black-Box Attacks
abstract
The query-based black-box attacks have raised serious threats to machine learning models in many real applications. In this work, we study a lightweight defense method, dubbed Random Noise Defense (RND), which adds proper Gaussian noise to each query. We conduct the theoretical analysis about the effectiveness of RND against query-based black-box attacks and the corresponding adaptive attacks. Our theoretical results reveal that the defense performance of RND is determined by the magnitude ratio between the noise induced by RND and the noise added by the attackers for gradient estimation or local search. The large magnitude ratio leads to the stronger defense performance of RND, and it's also critical for mitigating adaptive attacks. Based on our analysis, we further propose to combine RND with a plausible Gaussian augmentation Fine-tuning (RND-GF). It enables RND to add larger noise to each query while maintaining the clean accuracy to obtain a better trade-off between clean accuracy and defense performance. Additionally, RND can be flexibly combined with the existing defense methods to further boost the adversarial robustness, such as adversarial training (AT). Extensive experiments on CIFAR-10 and ImageNet verify our theoretical findings and the effectiveness of RND and RND-GF.
Zeyu Qin, Yanbo Fan, Hongyuan Zha, Baoyuan Wu
NeurIPS1
2020 Traffic Optimization in Satellites Communications: A Multi-agent Reinforcement Learning Approach
abstract
Past few years have witnessed the compelling applications of the satellite communications and networking in our daily life. Due to the extremely high moving speeds and limited networking resources of LEO satellites, how to optimize inter-satellite traffic has received amount of attention from both academia and industry. In this paper, we proposed a hybrid satellites network traffic control paradigm. In our architecture, the centralized platform collect the global state and the joint action from each agent during the training phase to ease the training, and during execution, the each agent can return the action to the local state through the trained policy. Besides, we adopt a multiagent actor-critic algorithms named Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments(MADDPG) to our architecture. In addition, some simulation results are presented to evaluate the correctness of our architecture and algorithm.
Zeyu Qin, Haipeng Yao, Tianle Mai
IWCMC1