EDBT 2026 Demo / reviewers in the wild / expert
Chaoxiang He
dblp:306/1330
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0008-8936-9336ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SoK: Robustness in Large Language Models against Jailbreak Attacks
Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang, Hanqing Hu, Zhengyan Zhou, Bin B. Zhu, Shifeng Sun 0001, Dawu Gu, Shuo Wang 0012 |
SP | 3 |
| 2025 | Enhancing Adversarial Transferability with Checkpoints of a Single Model's TrainingabstractAdversarial attacks threaten the integrity of deep neural networks (DNNs), particularly in high-stakes applications. In this paper, we present a novel black-box adversarial attack that leverages the diverse checkpoints generated during a single model’s training trajectory. Unlike conventional ensemble attacks that require multiple surrogate models with diverse architectures, our approach exploits the intrinsic diversity captured over different training stages of a single surrogate model. By decomposing the learned representations into task-intrinsic and task-irrelevant components, we employ an accuracy gap-based selection strategy to identify checkpoints that predominantly capture transferable, task-intrinsic knowledge. Extensive experiments on ImageNet and CIFAR-10 demonstrate that our method consistently outperforms traditional ensemble attacks in terms of transferability, even under resource-constrained and practical settings. This work offers a resource-efficient solution for crafting highly transferable adversarial examples and provides new insights into the dynamics of adversarial vulnerability. Shixin Li 0001, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Shuo Wang 0012, Hongsheng Hu, Dongmei Zhang 0001, Linchen Yu |
CVPR | 2 |
| 2025 | RESF: Regularized-Entropy-Sensitive Fingerprinting for Black-Box Tamper Detection of Large Language ModelsabstractThe proliferation of Machine Learning as a Service (MLaaS) has enabled widespread deployment of large language models (LLMs) via cloud APIs, but also raises critical concerns about model integrity and security.Existing black-box tamper detection methods, such as watermarking and fingerprinting, rely on the stability of model outputs-a property that does not hold for inherently stochastic LLMs.We address this challenge by formulating blackbox tamper detection for LLMs as a hypothesistesting problem.To enable efficient and sensitive fingerprinting, we derive a first-order surrogate for KL divergence-the entropy-gradient norm-to identify prompts most responsive to parameter perturbations.Building on this, we propose Regularized Entropy-Sensitive Fingerprinting (RESF), which enhances sensitivity while regularizing entropy to improve output stability and control false positives.To further distinguish tampering from benign randomness, such as temperature shifts, RESF employs a lightweight two-tier sequential test combining support-based and distributional checks with rigorous false-alarm control.Comprehensive analysis and experiments across multiple LLMs show that RESF achieves up to 98.80% detection accuracy under challenging conditions, such as minimal LoRA fine-tuning with five optimized fingerprints.RESF consistently demonstrates strong sensitivity and robustness, providing an effective and scalable solution for black-box tamper detection in cloud-deployed LLMs. Pingyi Hu, Xiaofan Bai, Xiaojing Ma 0002, Chaoxiang He, Dongmei Zhang 0001, Bin B. Zhu |
EMNLP | 4 |
| 2025 | Fine-Grained and Efficient Self-Unlearning with Layered IterationabstractAs machine learning models become widely deployed in data-driven applications, ensuring compliance with the 'right to be forgotten' as required by many privacy regulations is vital for safeguarding user privacy. To forget the given data, existing re-labeling based unlearning methods employ a single-step adjustment scheme that revises the decision boundaries in one re-labeling phase. However, such single-step approaches lead to coarse-grained changes in decision boundaries among the remaining classes and impose adverse effects on the model utility. To address these limitations, we propose 'Self-Unlearning with Layered Iteration (SULI),' a novel unlearning approach that introduces a layered iteration strategy to re-label the forgetting data iteratively and refine the decision boundaries progressively. We further develop a 'Selective Probability Adjustment (SPA)' technique, which uses a soft-label mechanism to promote smoother decision-boundary transitions. Comprehensive experiments on three benchmark datasets demonstrate that SULI achieves superior performance in effectiveness, efficiency, and privacy compared to the state-of-the-art baselines in both class-wise and instance-wise unlearning scenarios. The source code is released at https://github.com/Hongyi-Lyu-MQ/SULI. Hongyi Lyu, Xuyun Zhang, Hongsheng Hu, Shuo Wang 0026, Chaoxiang He, Lianyong Qi |
IJCAI | 5 |
| 2025 | BadFU: Backdoor Federated Learning through Adversarial Machine UnlearningabstractFederated learning (FL) has been widely adopted as a decentralized training paradigm that enables multiple clients to collaboratively learn a shared model without exposing their local data. As concerns over data privacy and regulatory compliance grow, machine unlearning, which aims to remove the influence of specific data from trained models, has become increasingly important in the federated setting to meet legal, ethical, or user-driven demands. However, integrating unlearning into FL introduces new challenges and raises largely unexplored security risks. In particular, adversaries may exploit the unlearning process to compromise the integrity of the global model. In this paper, we present the first backdoor attack in the context of federated unlearning, demonstrating that an adversary can inject backdoors into the global model through seemingly legitimate unlearning requests. Specifically, we propose BadFU, an attack strategy where a malicious client uses both backdoor and camouflage samples to train the global model normally during the federated training process. Once the client requests unlearning of the camouflage samples, the global model transitions into a backdoored state. Extensive experiments under various FL frameworks and unlearning strategies validate the effectiveness of BadFU, revealing a critical vulnerability in current federated unlearning practices and underscoring the urgent need for more secure and robust federated unlearning mechanisms. Bingguang Lu, Hongsheng Hu, Yuantian Miao, Shaleeza Sohail, Chaoxiang He, Shuo Wang 0012, Xiao Chen 0002 |
RAID | 5 |
| 2025 | Artificial intelligence security and privacy: a surveyabstractAbstract Artificial intelligence (AI) is revolutionizing both industries and reshaping the global economy. However, the rapid advancement of AI technologies brings significant security and privacy challenges. Recent incidents highlight vulnerabilities in AI systems, such as data leakage and malicious code injection, leading to severe financial losses and privacy breaches. Although existing studies have discussed specific security threats, they often lack detailed granularity and cover a limited scope. In this survey, we fill this gap by systematically categorizing and analyzing the threats and countermeasures in AI systems, which span both the training and inference stages, encompass centralized and distributed settings, and address both conventional and foundation AI models. By reviewing existing literature, we aim to provide AI researchers and practitioners with a thorough understanding of system vulnerabilities and current countermeasures. We hope to inspire further research into robust solutions, ultimately contributing to the development of resilient AI technologies. Xinlei He 0001, Guowen Xu, Xingshuo Han, Qian Wang 0002, Lingchen Zhao, Chao Shen 0001, Chenhao Lin, Zhengyu Zhao 0001, Qian Li 0024, Le Yang 0007, Shouling Ji, Shaofeng Li 0001, Haojin Zhu, Zhibo Wang 0001, Tianqing Zhu, Qi Li 0002, Chaoxiang He, Hongsheng Hu, Shuo Wang 0012, Shifeng Sun 0001, Hongwei Yao, Qinyu Zhang 0001, Kai Chen 0012, Yue Zhao 0027, Hongwei Li 0001, Xinyi Huang 0001, Dengguo Feng |
Sci. China Inf. Sci. | 18 |
| 2024 | MysticMask: Adversarial Mask for Impersonation Attack Against Face Recognition SystemsabstractIn our increasingly interconnected digital world, face recognition serves as a vital security layer for identity verification. However, impersonation attacks pose a significant threat to face recognition systems. While adversarial attacks have proven effective in impersonation, current methods primarily target face recognition, overlooking other crucial processes in real-world applications, such as action-based liveness detection.To address this gap, we introduce MysticMask, a novel adversarial mask attack designed to penetrate the entire face recognition pipeline for impersonation. MysticMask operates in the 3D domain, leveraging foldable medical face masks that automatically align with facial landmarks, enabling accurate physical simulations. MysticMask attacks all processes in a face recognition system simultaneously, including face detection, action-based liveness detection, and face recognition. Additionally, a landmark alignment technique is proposed to pass action-based liveness detection. Extensive experiments conducted in both simulated 3D and real-world scenarios demonstrate MysticMask’s superiority over state-of-the-art methods. Chaoxiang He, Yimiao Zeng, Xiaojing Ma 0002, Bin B. Zhu, Shixin Li 0001, Hai Jin 0001 |
ICME | 1 |
| 2024 | Intersecting-Boundary-Sensitive Fingerprinting for Tampering Detection of DNN ModelsabstractCloud-based AI services offer numerous benefits but also introduce vulnerabilities, allowing for tampering with deployed DNN models, ranging from injecting malicious behaviors to reducing computing resources. Fingerprint samples are generated to query models to detect such tampering. In this paper, we present Intersecting-Boundary-Sensitive Fingerprinting (IBSF), a novel method for black-box integrity verification of DNN models using only top-1 labels. Recognizing that tampering with a model alters its decision boundary, IBSF crafts fingerprint samples from normal samples by maximizing the partial Shannon entropy of a selected subset of categories to position the fingerprint samples near decision boundaries where the categories in the subset intersect. These fingerprint samples are almost indistinguishable from their source samples. We theoretically establish and confirm experimentally that these fingerprint samples’ expected sensitivity to tampering increases with the cardinality of the subset. Extensive evaluation demonstrates that IBSF surpasses existing state-of-the-art fingerprinting methods, particularly with larger subset cardinality, establishing its state-of-the-art performance in black-box tampering detection using only top-1 labels. The IBSF code is available at https://github.com/CGCL-codes/IBSF. Xiaofan Bai, Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Hai Jin 0001 |
ICML | 2 |
| 2024 | Towards Stricter Black-box Integrity Verification of Deep Neural Network ModelsabstractCloud-based machine learning services offer significant advantages but also introduce the risk of tampering with cloud-deployed deep neural network (DNN) models. Black-box integrity verification (BIV) allows model owners and end-users to determine if a cloud-deployed DNN model has been tampered with by examining only the top-1 label responses. Fingerprinting generates fingerprint samples to query the model, achieving BIV with no impact on the model's accuracy. In this paper, we present BIVBench, the first comprehensive benchmark for BIV of DNN models. BIVBench covers 16 types of model modifications, providing extensive coverage of practical modification scenarios. Our analysis reveals that existing fingerprinting methods, which are typically focused on significant tampering, lack the sensitivity needed to effectively detect subtle yet common and potentially severe modifications. To address this limitation, we propose MiSentry (Model Integrity Sentry), a novel fingerprinting method that leverages meta-learning. MiSentry strategically incorporates a few subtly modified models into the meta-learning model zoo and maximizes the divergence of output predictions between the target model and the modified models in the model zoo to generate highly sensitive, generalizable, and effective fingerprint samples. Extensive evaluations using BIVBench demonstrate that MiSentry outperforms existing state-of-the-art methods overall and significantly surpasses them in detecting subtle modifications. The BIVBench and supplementary materials are available at: https://github.com/CGCL-codes/BIVBench. Chaoxiang He, Xiaofan Bai, Xiaojing Ma 0002, Bin B. Zhu, Pingyi Hu, Jiayun Fu, Hai Jin 0001, Dongmei Zhang 0001 |
ACM Multimedia | 1 |
| 2024 | DorPatch: Distributed and Occlusion-Robust Adversarial Patch to Evade Certifiable Defenses
Chaoxiang He, Xiaojing Ma 0002, Bin B. Zhu, Yimiao Zeng, Hanqing Hu, Xiaofan Bai, Hai Jin 0001, Dongmei Zhang 0001 |
NDSS | 1 |
| 2021 | Feature-Indistinguishable Attack to Circumvent Trapdoor-Enabled DefenseabstractDeep neural networks (DNNs) are vulnerable to adversarial attacks. A great effort has been directed to developing effective defenses against adversarial attacks and finding vulnerabilities of proposed defenses. A recently proposed defense called Trapdoor-enabled Detection (TeD) deliberately injects trapdoors into DNN models to trap and detect adversarial examples targeting categories protected by TeD. TeD can effectively detect existing state-of-the-art adversarial attacks. In this paper, we propose a novel black-box adversarial attack on TeD, called Feature-Indistinguishable Attack (FIA). It circumvents TeD by crafting adversarial examples indistinguishable in the feature (i.e., neuron-activation) space from benign examples in the target category. To achieve this goal, FIA jointly minimizes the distance to the expectation of feature representations of benign samples in the target category and maximizes the distances to positive adversarial examples generated to query TeD in the preparation phase. A constraint is used to ensure that the feature vector of a generated adversarial example is within the distribution of feature vectors of benign examples in the target category. Our extensive empirical evaluation with different configurations and variants of TeD indicates that our proposed FIA can effectively circumvent TeD. FIA opens a door for developing much more powerful adversarial attacks. The FIA code is available at: https://github.com/CGCL-codes/FeatureIndistinguishableAttack. Chaoxiang He, Bin B. Zhu, Xiaojing Ma 0002, Hai Jin 0001, Shengshan Hu |
CCS | 1 |