Jiakai Wang

dblp:202/9216 · DBLP profile ↗
← Back
60ranked-venue papers
7as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 24 since 2021Security and privacy · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Activation Manipulation Attack: Penetrating and Harmful Jailbreak Attack Against Large Vision-Language Models
abstract
Recently, Large Vision-Language Models (LVLMs) have been demonstrated to be vulnerable to jailbreak attacks, highlighting the urgent need for further research to comprehensively identify and mitigate these threats. Unfortunately, existing jailbreak studies primarily focus on coarse-grained input manipulation to elicit specific responses, overlooking the exploitation of internal representations, i.e., intermediate activations, which constrains their ability to penetrate alignment safeguards and generate harmful responses. To tackle this issue, we propose the Activation Manipulation (ActMan) Attack framework, which performs fine-grained activation manipulations inspired by the perception and cognition stages of human decision-making, enhancing both the penetration capability and harmfulness of attacks. To improve penetration capability, we introduce a Deceptive Visual Camouflage module inspired by the masking effect in human perception. This module uses a benign activation-guided attention redirection strategy to conceal abnormal activation patterns, thereby suppressing LVLM's defense detection during early-stage decoding. To enhance harmfulness, we design a Malicious Semantic Induction module drawing from the framing effect in human cognition, which reconstructs jailbreak instructions using malicious activation guidance to change LVLM’s risk assessment during late-stage decoding, thereby amplifying the harmfulness of model responses. Extensive experiments on six mainstream LVLMs demonstrate that our method remarkably outperforms state-of-the-art baselines, achieving an average relative ASR improvement of 12.06%.
Haojie Hao, Jiakai Wang, Aishan Liu, Yuqing Ma, Haotong Qin, Yuanfang Guo, Xianglong Liu 0001
AAAI2
2026 Dormant Backdoor: Weaponizing Model Finetuning for Feasible Backdoor Attacks Against Pretrained Models
abstract
As the pretraining-finetuning paradigm becomes dominant in modern AI, the security of model supply chains faces new risks from backdoor attacks. Existing work primarily studies backdoors injected during pretraining and treats subsequent finetuning with clean data as a defense, while recent finetuning-activated attacks assume white-box access to the downstream data distribution, which is rarely realistic in practice. We introduce Dormant Backdoor, a finetuning-activated attack that requires no prior knowledge of downstream tasks. Instead of binding the backdoor to static input patterns, Dormant Backdoor exploits the universal dynamics of gradient-based optimization as a process-as-trigger mechanism. We formulate the attack as a bilevel optimization problem that simulates the victim's finetuning trajectory on proxy data, and jointly optimizes the poisoned model and trigger under lethality, utility, and stealth objectives. Before finetuning, the poisoned model remains behaviorally close to a clean model and can evade existing backdoor detectors; after finetuning, the same adaptation process reliably amplifies the backdoor on diverse downstream datasets and finetuning strategies. Our results reveal a previously underexplored class of process-as-trigger vulnerabilities and highlight the need for defenses that explicitly secure the model adaptation process.
Ruitao Li, Jiakai Wang, Hairong Chen, Huihu Ding, Jinghan Zhou, Renshuai Tao
AAAI2
2026 Query-Routed Activation Editing with Truth-hierarchical Preference Optimization
abstract
Hallucination has emerged as a pivotal challenge of Large Language Models (LLMs) that generate plausible yet non‑factual content, significantly impeding the trustworthy AI applications in real-world scenarios like medical diagnosis and autonomous driving. Editing the internal activations of LLMs during inference has shown promising effectiveness in mitigating hallucinations with minimal cost. However, previous editing approaches neglect the query‑specific inference pathways that require tailored truthful steering vectors, resulting in suboptimal hallucination mitigation. To address these issues, we propose the Query-Routed Activation Editing (QRAE) framework, which comprises Divergence-sensitive Head Routing (DHR) and Truth-hierarchical Preference Steering (TPS), to fully leverage query-specific semantics for adaptive activation editing. Specifically, DHR is proposed to establish a query-aware head selection criterion, thereby dynamically routing to truth-critical attention heads. Subsequently, TPS introduces a query-specific steering vector calibration policy with the guidance of progressive truth-preferred optimization, enabling precise and adaptive editing for each distinct query. Extensive experiments on the widely recognized TruthfulQA benchmark demonstrate that QRAE outperforms SOTA methods by up to 13.2% in MC1. Meanwhile, QRAE demonstrates strong generalization to out-of-distribution TriviaQA and Natural Questions benchmarks.
Kewei Liao, Yuqing Ma, Zhange Zhang, Zhicheng Geng, Jiakai Wang, Xianglong Liu 0001
AAAI7
2026 Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles
abstract
The generation of safety-critical scenarios in simulation has become increasingly crucial for safety evaluation in autonomous vehicles (AV) prior to road deployment in society. However, current approaches largely rely on predefined threat patterns or rule-based strategies, which limit their ability to expose diverse and unforeseen failure modes. To overcome these, we propose ScenGE, a framework that can generate plentiful safety-critical scenarios by reasoning novel adversarial cases and then amplifying them with complex traffic flows. Given a simple prompt of a benign scene, it first performs Meta-Scenario Generation, where a large language model (LLM), grounded in structured driving knowledge (e.g., traffic regulations, real-world accident records), infers an adversarial agent whose behavior poses a threat that is both plausible and deliberately challenging. This meta-scenario is then specified in executable code for precise in-simulator control. Subsequently, Complex Scenario Evolution uses background vehicles to amplify the core threat introduced by Meta-Scenario. It builds an adversarial collaborator graph to identify key agent trajectories for optimization. These perturbations are designed to simultaneously reduce the ego vehicle's maneuvering space and create critical occlusions. Extensive experiments conducted on multiple reinforcement learning (RL) based AV models show that ScenGE uncovers more severe collision cases (+31.96%) on average than SoTA baselines. Additionally, our ScenGE can be applied to large model based AV systems and deployed on different simulators; we further observe that adversarial training on our scenarios improves the model robustness. We hope our paper can build up a critical step towards building public trust and ensuring their safe deployment.
Jiangfan Liu 0001, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang 0004, Zonglei Jing, Siyuan Liang 0004, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu 0001
AAAI7
2026 First-Order Error Matters: Accurate Compensation for Quantized Large Language Models
abstract
Post-training quantization (PTQ) offers an efficient approach to compressing large language models (LLMs), significantly reducing memory access and computational costs. Existing compensation-based weight calibration methods often rely on a second-order Taylor expansion to model quantization error, under the assumption that the first-order term is negligible in well-trained full-precision models. However, we reveal that the progressive compensation process introduces accumulated first-order deviations between latent weights and their full-precision counterparts, making this assumption fundamentally flawed. To address this, we propose FOEM, a novel PTQ method that explicitly incorporates first-order gradient terms to improve quantization error compensation. FOEM approximates gradients by performing a first-order Taylor expansion around the pre-quantization weights. This yields an approximation based on the difference between latent and full-precision weights as well as the Hessian matrix. When substituted into the theoretical solution, the formulation eliminates the need to explicitly compute the Hessian, thereby avoiding the high computational cost and limited generalization of backpropagation-based gradient methods. This design introduces only minimal additional computational overhead. Extensive experiments across a wide range of models and benchmarks demonstrate that FOEM consistently outperforms the classical GPTQ method. In 3-bit weight-only quantization, FOEM reduces the perplexity of Llama3-8B by 17.3% and increases the 5-shot MMLU accuracy from 53.8% achieved by GPTAQ to 56.1%. Moreover, FOEM can be seamlessly combined with advanced techniques such as SpinQuant, delivering additional gains under the challenging W4A4KV4 setting and further narrowing the performance gap with full-precision baselines, surpassing existing state-of-the-art methods.
Xingyu Zheng, Haotong Qin, Yuye Li, Haoran Chu, Jiakai Wang, Jinyang Guo 0002, Michele Magno, Xianglong Liu 0001
AAAI5
2026 CASE: Conflict-assessed Knowledge-sensitive Neuron Tuning for Lifelong Model Editing
abstract
Large Language Models (LLMs) inevitably encounter factual hallucinations and knowledge obsolescence, necessitating lifelong knowledge editing to sustain reliability and factual advancement. While mainstream lifelong editing paradigms aim to alleviate knowledge forgetting through allocating and updating isolated parameter subspaces, they often overlook conflict assessment among distinct editing processes, leading to unjustified subspace allocation and indiscriminate neuron tuning. To address these issues, we propose the Conflict-Assessed Sensitive Editing (CASE) framework, which integrates a Conflict-Assessed Editing Allocation (CAA) module and a Knowledge-sensitive Neuron Tuning (KNT) strategy. The CAA module quantitatively assesses editing conflicts to enable justified subspace allocation, thereby reducing globally significant conflicts and routing errors. The KNT strategy adaptively identifies and tunes knowledge-sensitive neurons through a calibrated sensitivity threshold, effectively eliminating local conflicts and enhancing subspace stability. Extensive experiments on standard lifelong editing benchmarks demonstrate that CASE achieves state-of-the-art performance, improving average editing accuracy by nearly 10% after 1,000 sequential edits. Overall, CASE substantially mitigates editing conflicts and enhances knowledge retention, offering a scalable and conflict-resilient solution for lifelong model editing.
Zhange Zhang, Yuqing Ma, Jiakai Wang, Xianglong Liu 0001
WWW5
2026 Milmer: a framework for multiple instance learning based multimodal emotion recognition
Zaitian Wang, Yu Liang 0003, Xiyuan Hu, Tianhao Peng 0002, Jiakai Wang, Weili Zhang, Shuang Niu, Xiaoyang Xie
Neurocomputing7
2026 MoMIL: Multi-order enhanced multiple instance learning for computational pathology
Jiakai Wang, Baoyu Liang, Yuancheng Yang, Chao Tong 0001
Image Vis. Comput.3
2026 DynamicPAE: Generating Scene-Aware Physical Adversarial Examples in Real-Time
abstract
Physical adversarial examples (PAEs) are regarded as “whistle-blowers” of real-world risks in deep-learning applications, thus worth further investigation. However, current PAE generation studies show limited adaptive attacking ability to diverse and varying scenes, revealing the urgent requirement of dynamic PAEs that are generated in real time and conditioned on the observation from the attacker. The key challenge in generating dynamic PAEs is learning the sparse relation between PAEs and the observation of attackers under the noisy feedback of attack training. To address the challenge, we present DynamicPAE, the first generative framework that enables scene-aware real-time physical attacks. Specifically, to address the noisy feedback problem that obfuscates the exploration of scene-related PAEs, we introduce the residual-guided adversarial pattern exploration technique. We first introduce the limited feedback information restriction to model the training degeneracy problem under noisy feedback. Then, residual-guided training, which relaxes the attack training with a reconstruction task, is proposed to enrich the feedback information, thereby achieving a more comprehensive exploration of PAEs. To address the alignment problem between the trained generator, which represents the learned relation, and the real-world scenario, we introduce the distribution-matched attack scenario alignment, consisting of the conditional-uncertainty-aligned data module and the skewness-aligned objective re-weighting module. The former aligns the training environment with the incomplete observation of the real-world attacker. The latter facilitates consistent stealth control across different attack targets by balancing the objectives with the skewness indicator. Extensive digital and physical evaluations demonstrate the superior attack performance of DynamicPAE, attaining a 2.07× boost ( 58.8% average AP drop under attack) on representative object detectors (e.g., DETR) over state-of-the-art static PAE generating methods. Overall, our work opens the door to end-to-end modeling of dynamic PAEs.
Xianglong Liu 0001, Jiakai Wang, Xianqi Yang, Haotong Qin, Yuqing Ma, Ke Xu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Prioritized scanning: Combining spatial information multiple instance learning for computational pathology
Jiakai Wang, Baoyu Liang, Yuancheng Yang, Siyang Wu, Chao Tong 0001
Pattern Recognit.2
2026 Leveraging Robustness-Aware Channel Activation for Privacy Protection and Tracing Forensics
abstract
Sharing personal photos on social media exposes users to unauthorized identity recognition and unconsented model training, raising severe privacy and copyright concerns. Existing methods typically focus on either privacy protection, which misleads recognition models to prevent unauthorized automated recognition, or tracing forensics, which embeds traceable patterns for ownership verification. However, they fail to achieve both simultaneously. The core challenge is to jointly achieve privacy protection and tracing forensics within a single perturbation, since the two objectives rely on different feature behaviors and naive combinations are ineffective in practice. In this work, we propose ATP (Adversarial Tracing Perturbation), a novel perturbation generation method that activates robustness-aware feature channels to balance privacy and traceability. ATP leverages non-robust channel activation to mislead recognition models for privacy protection, while robust channel activation embeds traceable patterns for reliable tracing forensics. Extensive experiments on image classification and face recognition show that ATP achieves strong dual protection, improving overall dual-protection performance by bm 3.54× over the baselines while remaining effective under adaptive attacks, thereby demonstrating strong robustness and practical applicability. © 2026 IEEE.
Haodi Wang, Kai Dong 0001, Jiakai Wang, Xianglong Liu 0001, Guangdong Bai
IEEE Trans. Dependable Secur. Comput.4
2026 NOAE: Noise-Optimized Adversarial Examples for Multivariate Time Series Anomaly Detection of the Industrial Internet of Things
Yusheng Kong, Jiakai Wang, Shixiang Li, Lei Ren 0001
IEEE Trans. Inf. Forensics Secur.3
2026 QuEST: Quantization-Conditioned Efficient Stealthy Trojan
abstract
Quantization-conditioned backdoor attacks, which exploit model quantization states to trigger malicious behavior, pose a hidden threat to deep learning security. However, current studies ignore the feasibility of attacks,i.e., detection visibility, and computational overhead, leading to significant constraints on their practical deployment in real-world adversarial scenarios. To address these limitations, we propose Quantization-conditioned Efficient Stealthy Trojan (QuEST), a novel framework that enhances both the stealth and efficiency of backdoor attacks under quantization constraints. For enhancing stealthiness, we design a stealth-optimized training scheme that benefits from the parametric backdoor injection and trigger scaling augmentation to maintain the consistency of model behavior during attack. In this way, the defender will fail to capture the suspicious behavior differences for detection due to the made efforts in both model-side and data-side. To improve efficiency, we introduce information-guided parameter sharing, which utilizes parameter redundancy analysis and Fisher divergence metrics to identify a minimal amount of quantization-preserved parameters for back-door injection. These parameters are strategically shared between the malicious and benign models, enabling concurrent training and substantially reducing overall training time. Extensive experiments demonstrate that QuEST maintains competitive attack success rates while improving stealth performance by 18.75% and reducing computational costs by 26.66% on average compared to state-of-the-art methods, highlighting QuEST’s potential for more practical adversarial deployments in real-world scenarios. Our code is available here.
Shuchao Pang, Jiakai Wang, Yunhuai Liu, Xianglong Liu 0001, Yongbin Zhou
IEEE Trans. Inf. Forensics Secur.3
2026 Practical and Flexible Backdoor Attack Against Deep Learning Models via Shell Code Injection
abstract
Recently, backdoor attack, which aims to implant malicious logic into deep learning models (DLMs), has attracted so extensive research attention. Among them, the non-poisoning-based backdoor attack appears considerable development prospects owing to the posed threats against the DLMs-based artificial intelligence applications in cyberspace. However, previous non-poisoning- based backdoor attacks for DLMs are limited to the impractical attacking forms, resulting in certain weaknesses in both attacking complexity and attacking adaptability. To tackle the mentioned issues, this paper proposes a novel backdoor attack framework, namely the shell code injection (SCI), to perform backdoor attacks against DLMs with lower complexity and higher adaptability. Specifically, for alleviating the attacking complexity, we elaborate the logic-driven stealthy backdoor shell motivated by the biological behavior in nature,e.g., the camouflage and attack strategy of crabs. By introducing the trigger consistency verification and short-circuit code packaging strategies, the SCI misleads the victim models to output wrong predictions without training requirements according to the preset poisonous decision logic. For enhancing the attacking adaptability, we design the LLM-assisted adaptive attacking target code generation that consists of the model concept detection module and the attack target adjusting module. Since the attacking goals could be generated dynamically according to the aware victim model information and appointed attacker preset instructions, the SCI could achieve more flexible attacking performance. Extensive experiments are conducted to demonstrate that the proposed backdoor attack framework appears awesome attacking ability (almost 100% ASR) under various settings. Additionally, we provide a case study on combining the cyber attack with SCI, which also exhibits certain space for imagination of new-type backdoor attacks.
Jiakai Wang, Renshuai Tao, Xianglong Liu 0001, Yao Zhao 0001
IEEE Trans. Inf. Forensics Secur.1
2026 Toward Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-Critical Scenarios
abstract
Autonomous driving has made significant progress in both academia and industry, including performance improvements in perception tasks and the development of end-to-end autonomous driving systems. However, the safety and robustness assessment of autonomous driving has not received sufficient attention. Current evaluations of autonomous driving are typically conducted in natural driving scenarios. However, accidents often occur in edge cases, also known as safety-critical scenarios. These safety-critical scenarios are difficult to collect, and there is currently no clear definition of what constitutes a safety-critical scenario. In this work, we explore the safety and robustness of autonomous driving in safety-critical scenarios. First, we provide a definition of safety-critical scenarios, including static traffic scenarios such as adversarial attack scenarios and natural distribution shifts, as well as dynamic traffic scenarios such as accident scenarios. Then, we develop an autonomous driving test framework to comprehensively evaluate autonomous driving systems, encompassing not only the assessment of perception modules but also system-level evaluations. Our work systematically constructs a safety verification process for autonomous driving, providing technical support for the industry to establish standardized test framework.
Jingzheng Li, Xianglong Liu 0001, Shikui Wei, Yufei Ge, Bing Li 0001, Qing Guo 0005, Xianqi Yang, Yanjun Pu, Qianren Mao, Jiakai Wang
IEEE Trans. Image Process.11
2025 Lexical Diversity-aware Relevance Assessment for Retrieval-Augmented Generation
abstract
Zhange Zhang, Yuqing Ma, Yulong Wang, Shan He, Tianbo Wang, Siqi He, Jiakai Wang, Xianglong Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhange Zhang, Yuqing Ma, Siqi He, Jiakai Wang, Xianglong Liu 0001
ACL (1)7
2025 Harnessing Global-Local Collaborative Adversarial Perturbation for Anti-Customization
abstract
Though achieving significant success in personalized image synthesis, Latent Diffusion Models (LDMs) pose substantial social risks caused by unauthorized misuse (e.g., face theft). To counter these threats, the Anti-Customization (AC) method that exploits adversarial perturbations was proposed. Unfortunately, existing AC methods show insufficient defense ability due to the ignorance of hierarchical characteristics, i.e., global feature correlations and local facial attributes, leading to weak resistance to concept transfer and semantic theft in customization methods. To address these limitations, we are motivated to propose a Global-Local Collaborated Anti-Customization (GoodAC) framework to generate powerful adversarial perturbations by disturbing both feature correlations and facial attributes. To enhance the ability to resist concept transfer, we disrupt the spatial correlation of perceptual features that form the basis of model generation at a global level, thereby creating highly concept-transfer-resistant adversarial camouflage. To improve the ability to resist semantic theft, leveraging the fact that facial attributes are personalized, we designed a personalized and precise facial attribute distortion strategy locally, focusing the attack on the individual’s image structure to generate strong camouflage. Extensive experiments on various customization methods, including Dreambooth and LoRA, have strongly demonstrated that our GoodAC outperforms other state-of-the-art approaches by large margins, e.g., over 50% improvements on ISM.1
Jiakai Wang, Haojie Hao, Haotong Qin, Jiejie Zhao, Xianglong Liu 0001
CVPR2
2025 Token-Aware Editing of Internal Activations for Large Language Model Alignment
abstract
Intervening the internal activations of large language models (LLMs) provides an effective inference-time alignment approach to mitigate undesirable behaviors, such as generating erroneous or harmful content, thereby ensuring safe and reliable applications of LLMs.However, previous methods neglect the misalignment discrepancy among varied tokens, resulting in deviant alignment direction and inflexible editing strength.To address these issues, we propose a token-aware editing (TAE) approach to fully utilize token-level alignment information in the activation space, therefore realizing superior post-intervention performance.Specifically, a Mutual Information-guided Graph Aggregation (MIG) module first develops an MI-guided graph to exploit the tokens' informative interaction for activation enrichment, thus improving alignment probing and facilitating intervention.Subsequently, Misalignment-aware Adaptive Intervention (MAI) comprehensively perceives the token-level misalignment degree from token representation and prediction to guide the adaptive adjustment of editing strength, thereby enhancing final alignment performance.Extensive experiments on three alignment capabilities demonstrate the efficacy of TAE, notably surpassing baseline by 25.8% on the primary metric of truthfulness with minimal cost. 1 MHSA
Yuqing Ma, Kewei Liao, Chengzhao Yang, Zhange Zhang, Jiakai Wang, Xianglong Liu 0001
EMNLP6
2025 Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring
abstract
There is a growing concern about adversarial attacks against automatic speech recognition (ASR) systems. Although research into targeted universal adversarial examples (AEs) has progressed, current methods are constrained by inefficient exploitation of audio features, demonstrating insufficient attack ability and robustness in the physical world. To solve this problem, we propose a phoneme-tailored attack (PTA) to improve the quality of the generated AEs. Specifically, to improve attack ability, we propose a Diverse Audio Composition Enrichment method, which enhances the utilization of audio features through phoneme-level slicing and recombination. To adapt AEs to complex environments, we propose a Natural Noise Pattern Guidance method to align AEs with natural noise patterns to improve their robustness. Experiments show that our method achieves an average accuracy of more than 72.34% and 98% with and without a norm constraint, and also demonstrates excellent performance in terms of generalization across datasets and resilience to MP3 compression.
Yanqu Chen, Jiakai Wang, Renshuai Tao, Xianglong Liu 0001
ICASSP3
2025 MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
abstract
Large Language Models (LLMs) have displayed massive improvements in reason- ing and decision-making skills and can hold natural conversations with users. Recently, many tool-use benchmark datasets have been proposed. However, existing datasets have the following limitations: (1). Insufficient evaluation scenarios (e.g., only cover limited tool-use scenes). (2). Extensive evaluation costs (e.g., GPT API costs). To address these limitations, in this work, we propose a multi-granularity tool-use benchmark for large language models called MTU-Bench. For the "multi-granularity" property, our MTU-Bench covers five tool usage scenes (i.e., single-turn and single-tool, single-turn and multiple-tool, multiple-turn and single-tool, multiple-turn and multiple-tool, and out-of-distribution tasks). Besides, all evaluation metrics of our MTU-Bench are based on the prediction results and the ground truth without using any GPT or human evaluation metrics. Moreover, our MTU-Bench is collected by transforming existing high-quality datasets to simulate real-world tool usage scenarios, and we also propose an instruction dataset called MTU-Instruct data to enhance the tool-use abilities of existing LLMs. Comprehensive experimental results demonstrate the effectiveness of our MTU-Bench.
Noah Wang, Xiaoshuai Song, Z. Y. Peng, Ken Deng, Jiakai Wang, Junran Peng, Ge Zhang 0009, Hangyu Guo, Zhaoxiang Zhang 0001, Wenbo Su, Bo Zheng 0007
ICLR9
2025 BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
abstract
With the advancement of diffusion models (DMs) and the substantially increased computational requirements, quantization emerges as a practical solution to obtain compact and efficient low-bit DMs. However, the highly discrete representation leads to severe accuracy degradation, hindering the quantization of diffusion models to ultra-low bit-widths. This paper proposes a novel weight binarization approach for DMs, namely BinaryDM, pushing binarized DMs to be accurate and efficient by improving the representation and optimization. From the representation perspective, we present an Evolvable-Basis Binarizer (EBB) to enable a smooth evolution of DMs from full-precision to accurately binarized. EBB enhances information representation in the initial stage through the flexible combination of multiple binary bases and applies regularization to evolve into efficient single-basis binarization. The evolution only occurs in the head and tail of the DM architecture to retain the stability of training. From the optimization perspective, a Low-rank Representation Mimicking (LRM) is applied to assist the optimization of binarized DMs. The LRM mimics the representations of full-precision DMs in low-rank space, alleviating the direction ambiguity of the optimization process caused by fine-grained alignment. Comprehensive experiments demonstrate that BinaryDM achieves significant accuracy and efficiency gains compared to SOTA quantization methods of DMs under ultra-low bit-widths. With 1-bit weight and 4-bit activation (W1A4), BinaryDM achieves as low as 7.74 FID and saves the performance from collapse (baseline FID 10.87). As the first binarization method for diffusion models, W1A4 BinaryDM achieves impressive 15.2x OPs and 29.2x model size savings, showcasing its substantial potential for edge deployment.
Xingyu Zheng, Xianglong Liu 0001, Haotong Qin, Xudong Ma, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo 0002, Michele Magno
ICLR7
2025 Unlocking the Potential of Lightweight Quantized Models for Deepfake Detection
abstract
Deepfake detection is increasingly crucial due to the rapid rise of AI-generated content. Existing methods achieve high performance relying on computationally intensive large models, making real-time detection on resource-constrained edge devices challenging. Given that deepfake detection is a binary classification task, there is potential for model compression and acceleration. In this paper, we propose a low-bit quantization framework for lightweight and efficient deepfake detection. The Connected Quantized Block extracts common forgery features via the quantized path and retains method-specific textures through the shortcut connections. Additionally, the Shifted Logarithmic Redistribution Quantizer mitigates information loss in near-zero domains by unfolding the unbalanced activations, enabling finer quantization granularity. Comprehensive experiments demonstrate that this new framework significantly reduces 10.8x computational costs and 12.4x storage requirements while maintaining high detection performance, even surpassing SOTA methods using less than 5% FLOPs, paving the way for efficient deepfake detection in resource-limited scenarios.
Renshuai Tao, Ziheng Qin, Chuangchuang Tan, Jiakai Wang, Wei Wang 0108
IJCAI5
2025 BadMDA: Towards Backdoor Injection during Domain Adaptation to Collapse Multi-Agent Perception
abstract
Domain adaptation, which bridges the domain gap between heterogeneous agents, has emerged as an effective solution to improve the perception capabilities of multi-agent systems. However, it may introduce backdoor vulnerabilities, as adversaries could exploit the collaborative process to propagate malicious features across agents, yet these threats remain largely unexplored. In this paper, we take the first step to study the backdoor attacks in this safety-critical scenario, with the 3D object detection task as the representative case. To this end, we propose BadMDA, the first backdoor attack tailored for the domain adaptation process to collapse multi-agent perception. Specifically, we first propose a gradient-suppression trigger optimization module to mitigate trigger distortion during the domain adaptation. By utilizing the optimizable additive triggers and minimizing gradient variations of triggered features induced by the domain adaptation, we reduce the transformation magnitude of triggered features, thereby maintaining the trigger effectiveness. Then, we propose a dual-gradient guided poisoning module to achieve clean-label poisoning in 3D object detection tasks. This module aligns training gradients with poisoned ones to learn malicious features, while enforcing the orthogonality between training and benign gradients. Consequently, the learned malicious features mislead the victim's finetuning updates, causing detection failures upon receiving triggered features while only slightly affecting the victim agent's model utility. Extensive experiments on various dominant domain adaptation methods show the superior attacking effectiveness and universality of BadMDA, underscoring the need for a more advanced defense.
Bowen Du 0001, Jiejie Zhao, Hanyang Xia, Jiakai Wang
ACM Multimedia6
2025 MetAdv: A Unified and Interactive Adversarial Testing Platform for Autonomous Driving
abstract
Evaluating and ensuring the adversarial robustness of autonomous driving (AD) systems is a critical and unresolved challenge. This paper introduces MetAdv, a novel adversarial testing platform that enables realistic, dynamic, and interactive evaluation by tightly integrating virtual simulation with physical vehicle feedback. At its core, MetAdv establishes a hybrid virtual-physical sandbox, within which we design a three-layer closed-loop testing environment with dynamic adversarial test evolution. This architecture facilitates end-to-end adversarial evaluation, ranging from high-level unified adversarial generation, through mid-level simulation-based interaction, to low-level execution on physical vehicles. Additionally, MetAdv supports a broad spectrum of AD tasks, algorithmic paradigms (e.g., modular deep learning pipelines, end-to-end learning, vision-language models). It supports flexible 3D vehicle modeling and seamless transitions between simulated and physical environments, with built-in compatibility for commercial platforms such as Apollo and Tesla. A key feature of MetAdv is its human-in-the-loop capability: besides flexible environmental configuration for more customized evaluation, it enables real-time capture of physiological signals and behavioral feedback from drivers, offering new insights into human-machine trust under adversarial conditions. We believe MetAdv can offer a scalable and unified framework for adversarial assessment, paving the way for safer AD. Our demo can be found at https://sites.google.com/view/metadv-demo-video.
Aishan Liu, Jiakai Wang, Tianyuan Zhang 0004, Hainan Li, Jiangfan Liu 0001, Siyuan Liang 0004, Yilong Ren, Xianglong Liu 0001, Dacheng Tao
ACM Multimedia2
2025 Exploring Semantic-constrained Adversarial Example with Instruction Uncertainty Reduction
abstract
Recently, semantically constrained adversarial examples (SemanticAE), which are directly generated from natural language instructions, have become a promising avenue for future research due to their flexible attacking forms, but have not been thoroughly explored yet. To generate SemanticAEs, current methods fall short of satisfactory attacking ability as the key underlying factors of semantic uncertainty in human instructions, such as $\textit{referring diversity}$, $\textit{descriptive incompleteness}$, and $\textit{boundary ambiguity}$, have not been fully investigated. To tackle the issues, this paper develops a multi-dimensional $\textbf{ins}$truction $\textbf{u}$ncertainty $\textbf{r}$eduction ($\textbf{InSUR}$) framework to generate more satisfactory SemanticAE, $\textit{i.e.}$, transferable, adaptive, and effective. Specifically, in the dimension of the sampling method, we propose the residual-driven attacking direction stabilization to alleviate the unstable adversarial optimization caused by the diversity of language references. By coarsely predicting the language-guided sampling process, the optimization process will be stabilized by the designed ResAdv-DDIM sampler‌, therefore releasing the transferable and robust adversarial capability of multi-step diffusion models. In task modeling, we propose the context-encoded attacking scenario constraint to supplement the missing knowledge from incomplete human instructions. Guidance masking and renderer integration are proposed to regulate the constraints of 2D/3D SemanticAE, activating stronger scenario-adapted attacks. Moreover, in the dimension of generator evaluation, we propose the semantic-abstracted attacking evaluation enhancement by clarifying the evaluation boundary based on the label taxonomy, facilitating the development of more effective SemanticAE generators. Extensive experiments demonstrate the superiority of the transfer attack performance of InSUR. Besides, it is worth highlighting that we realize the reference-free generation of semantically constrained 3D adversarial examples by utilizing language-guided 3D generation models for the first time.
Jiakai Wang, Linna Jing, Haotong Qin, Aishan Liu, Ke Xu 0001, Xianglong Liu 0001
NeurIPS2
2025 Dual Intention Escape: Penetrating and Toxic Jailbreak Attack against Large Language Models
abstract
Recently, the jailbreak attack, which generates adversarial prompts to bypass safety measures and mislead large language models (LLMs) to output harmful answers, has attracted extensive interest due to its potential to reveal the vulnerabilities of LLMs. However, ignoring the exploitation of the characteristics in intention understanding, existing studies could only generate prompts with weak attacking ability, failing to evade defenses (e.g., sensitive word detect) and causing malice(e.g., harmful outputs). Motivated by the mechanism in the psychology of human misjudgment, we propose a dual intention escape (DIE) jailbreak attack framework to generate more stealthy and toxic prompts to deceive LLMs to output harmful content. For stealthiness, inspired by the anchoring effect, we designed the Intention-anchored Malicious Concealment(IMC) module that hides the harmful intention behind a generated anchor intention by the recursive decomposition block and contrary intention nesting block. Since the anchor intention will be received first, the LLMs might pay less attention to the harmful intention and enter response status. For toxicity, we propose the Intention-reinforced Malicious Inducement (IMI) module based on the availability bias mechanism in a progressive malicious prompting approach. Due to the ongoing emergence of statements correlated to harmful intentions, the output content of LLMs will be closer to these more accessible intentions, i.e., more toxic. We conducted extensive experiments under black-box settings, supporting that DIE could achieve 100% ASR-R and 92.9% ASR-G against GPT3.5-turbo.
Yanni Xue, Jiakai Wang, Zixin Yin, Yuqing Ma, Haotong Qin, Renshuai Tao, Xianglong Liu 0001
WWW2
2025 Pre-trained Trojan Attacks for Visual Recognition
Aishan Liu, Xianglong Liu 0001, Xinwei Zhang 0010, Yisong Xiao, Yuguang Zhou, Siyuan Liang 0004, Jiakai Wang, Xiaochun Cao, Dacheng Tao
Int. J. Comput. Vis.7
2025 Attacking cooperative multi-agent reinforcement learning by adversarial minority influence
Jun Guo 0009, Jingqiao Xiu, Yuwei Zheng, Pu Feng, Xin Yu 0009, Jiakai Wang, Aishan Liu, Yaodong Yang 0001, Bo An 0001, Wenjun Wu 0001, Xianglong Liu 0001
Neural Networks7
2025 LS-PRISM: A layer-selective pruning method via low-rank approximation and sparsification for efficient large language model compression
Renshuai Tao, Hairong Chen, Yuzhe Guo, Jiakai Wang, Boying Wang, Yao Zhao 0001
Neural Networks4
2025 Dual Dependency Disentangling for Defending Model Inversion Attacks in Split Federated Learning
abstract
Recent studies have revealed that Split Federated Learning (SFL) is vulnerable to Model Inversion (MI) attacks, where the attacker can reconstruct clients’ raw data by exploiting collected features. Though achieving results, current defenses are unsatisfactory due to the limited ability to suppress the sensitive information while preserving task-conducive information within features. Since such limited ability can be attributed to insufficient disentanglement of data-feature and feature-task dependencies, we propose a Dual Dependency Disentangling framework for SFL (D3SFL) to strengthen defense ability against MI attacks while maintaining the utility. Specifically, we first propose a variable-structure data-feature dependency decoupling module, which produces privacy-preserving features by learning input-specific sub-networks, therefore enhancing the disentanglement of data-feature dependencies to hide sensitive information. Then, we propose a stochastic feature-task dependency separating module that adopts sparse binary masks to preserve the target-task-critical features and reduce sensitive information, resulting in effective disentanglement of feature-task dependencies for lower privacy leakage and better utility maintenance. Extensive experiments on image-classification datasets (CIFAR-100 and FaceScrub) and the time-series dataset (METR-LA) show that D3SFL outperforms the comparisons, achieving remarkable defense ability against MI attacks (with up to 54×, 17×, and 18× reconstruction MSE on average, respectively) while maintaining better utility (with only 0.13% and 0.06% Accuracy drops over the standard SFL on CIFAR-100 and FaceScrub, respectively, and only a 0.03 MAE increase on METR-LA over CNFGNN). Our code is available at https://github.com/Shawn-CT/D3SFL.
Jiakai Wang, Jiejie Zhao, Bowen Du 0001, Xiaoshan Bai, Zheng Lin 0005, Xianglong Liu 0001
IEEE Trans. Inf. Forensics Secur.2
2025 Compromising LLM Driven Embodied Agents With Contextual Backdoor Attacks
Aishan Liu, Yuguang Zhou, Xianglong Liu 0001, Tianyuan Zhang 0004, Siyuan Liang 0004, Jiakai Wang, Yanjun Pu, Tianlin Li, Wenbo Zhou 0004, Qing Guo 0005, Dacheng Tao
IEEE Trans. Inf. Forensics Secur.6
2025 SAGNet: Decoupling Semantic-Agnostic Artifacts From Limited Training Data for Robust Generalization in Deepfake Detection
abstract
Deepfake detection presents a significant challenge, particularly when the available training data is constrained to a limited set of semantic categories—a common and realistic scenario. In deepfake detection, the training labels typically indicate whether an image is real or fake, without specifying the semantic content, such as object classes. Moreover, we cannot know in advance the object categories present in an image to be detected. Ideally, a deepfake detection model should perform consistently across different semantic categories during inference, irrespective of the content. However, existing methods often exhibit significant performance bias between seen and unseen classes, struggling to generalize effectively. To address this issue, we propose Semantic-AGnostic artifact Network (SAGNet), an innovative and efficient approach designed to decouple semantic-agnostic artifacts from content-specific distributions in the training data. Our method eliminates semantic-specific biases, ensuring that the model focuses on universal artifacts related to image authenticity rather than content-dependent features. By employing this decoupling strategy, SAGNet greatly enhances the model’s generalization capacity, even when trained on limited data. Remarkably, through experiments, we demonstrate that SAGNet achieves performance comparable to models trained with 10 times more data, despite being trained on only 2 classes (comparing SAGNet trained on 2 classes in Table I with Ojha [1] trained on 20 categories in Table IV). Furthermore, through extensive experiments, we show that SAGNet’s improvements are not only evident across different semantic categories but also extend to various generative methods, including multiple GAN-based and diffusion-based models. This cross-method generalization emphasizes SAGNet’s versatility and effectiveness in diverse generative scenarios. Overall, our method represents a significant advancement in deepfake detection, particularly in realistic situations where the training data is limited. The code is released at https://github.com/rstao-bjtu/SAGNet/.
Renshuai Tao, Chuangchuang Tan, Huan Liu 0030, Jiakai Wang, Haotong Qin, Yakun Chang, Wei Wang 0108, Yao Zhao 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Causality-Inspired Debiasing Learning for Open World Object Detection
abstract
Open world object detection (OWOD) aims to identify both known instances of trained classes and unknown ones. Despite recent advancements, existing methods exhibit a detection bias towards known classes, as detectors are exclusively trained under the supervision of known classes. To address this problem, we construct a causal graph to scrutinize OWOD from a causal perspective, revealing that the bias problem primarily arises due to the confounding effect of known classes, and the causality between unknown objects and their predictions learned by the detector is weak. Therefore, we propose a causality-inspired debiasing framework for OWOD, aiming to bolster the performance of OWOD models by eliminating confounders and encouraging appropriate features. Specifically, a semantic causal intervention module is proposed to remove the confounding effect from known classes to unknown features, which introduces the known semantics to interact fairly with all unknown features through backdoor adjustment. Moreover, an unknown causality enhancement module is employed to enhance the causality of unknown objects and their predictions acquired by the model, which imposes constraints for different unknown classes in feature space with the contrastive learning paradigm from the perspective of intervention effect. Extensive experiments conducted on the commonly-used OWOD benchmarks demonstrate that our framework consistently yields superior results on unknown classes compared with state-of-the-art methods by a large margin (+25.0% UD-Pre, +10.2% Recall on unknown classes) and even better on known classes (+1.4% mAP on known classes).
Yuqing Ma, Chengtao Lv, Jiakai Wang, Xianglong Liu 0001
IEEE Trans. Multim.4
2024 NAPGuard: Towards Detecting Naturalistic Adversarial Patches
abstract
Recently, the emergence of naturalistic adversarial patch (NAP), which possesses a deceptive appearance and various representations, underscores the necessity of developing robust detection strategies. However, existing approaches fail to differentiate the deep-seated natures in adversarial patches, i.e., aggressiveness and naturalness, leading to unsatisfactory precision and generalization against NAPs. To tackle this issue, we propose NAP-Guard to provide strong detection capability against NAPs via the elaborated critical feature modulation framework. For improving precision, we propose the aggressive feature aligned learning to enhance the model's capability in capturing accurate aggressive patterns. Considering the challenge of inaccurate model learning caused by deceptive appearance, we align the aggressive features by the proposed pattern alignment loss during training. Since the model could learn more accurate aggressive patterns, it is able to detect deceptive patches more precisely. To enhance generalization, we design the natural feature suppressed inference to universally mitigate the disturbance from different NAPs. Since various representations arise in diverse disturbing forms to hinder generalization, we suppress the natural features in a unified approach via the feature shield module. Therefore, the models could recognize NAPs within less disturbance and activate the generalized detection ability. Extensive experiments show that our method surpasses state-of-the-art methods by large margins in detecting NAPs (improve 60.24% [email protected] on average).11Our code is available at https://github.com/wsynuiag/NAPGaurd.
Siyang Wu, Jiakai Wang, Jiejie Zhao, Yazhe Wang, Xianglong Liu 0001
CVPR2
2024 Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian Game
abstract
In this study, we explore the robustness of cooperative multi-agent reinforcement learning (c-MARL) against Byzantine failures, where any agent can enact arbitrary, worst-case actions due to malfunction or adversarial attack. To address the uncertainty that any agent can be adversarial, we propose a Bayesian Adversarial Robust Dec-POMDP (BARDec-POMDP) framework, which views Byzantine adversaries as nature-dictated types, represented by a separate transition. This allows agents to learn policies grounded on their posterior beliefs about the type of other agents, fostering collaboration with identified allies and minimizing vulnerability to adversarial manipulation. We define the optimal solution to the BARDec-POMDP as an ex interim robust Markov perfect Bayesian equilibrium, which we proof to exist and the corresponding policy weakly dominates previous approaches as time goes to infinity. To realize this equilibrium, we put forward a two-timescale actor-critic algorithm with almost sure convergence under specific conditions. Experiments on matrix game, Level-based Foraging and StarCraft II indicate that, our method successfully acquires intricate micromanagement skills and adaptively aligns with allies under worst-case perturbations, showing resilience against non-oblivious adversaries, random allies, observation-based attacks, and transfer-based attacks.
Jun Guo 0009, Jingqiao Xiu, Ruixiao Xu, Xin Yu 0009, Jiakai Wang, Aishan Liu, Yaodong Yang 0001, Xianglong Liu 0001
ICLR6
2024 Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation
Yanni Xue, Haojie Hao, Jiakai Wang, Qiang Sheng 0001, Renshuai Tao, Pu Feng, Xianglong Liu 0001
IJCAI3
2024 RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts
abstract
Believable agents can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication. Recently, generative agents have been proposed to simulate believable human behavior by using Large Language Models. However, the existing method heavily relies on human-annotated agent profiles (e.g., name, age, personality, relationships with others, and so on) for the initialization of each agent, which cannot be scaled up easily. In this paper, we propose a scalable RoleAgent framework to generate high-quality role-playing agents from raw scripts, which includes building and interacting stages. Specifically, in the building stage, we use a hierarchical memory system to extract and summarize the structure and high-level information of each agent for the raw script. In the interacting stage, we propose a novel innovative mechanism with four steps to achieve a high-quality interaction between agents. Finally, we introduce a systematic and comprehensive evaluation benchmark called RoleAgentBench to evaluate the effectiveness of our RoleAgent, which includes 100 and 28 roles for 20 English and 5 Chinese scripts, respectively. Extensive experimental results on RoleAgentBench demonstrate the effectiveness of RoleAgent.
Zehao Ni, Haoran Que, Tao Sun 0016, Noah Wang, Jian Yang 0030, Jiakai Wang, Hongcheng Guo, Zhongyuan Peng, Ge Zhang 0009, Xingyuan Bu, Ke Xu 0001, Wenge Rong, Junran Peng, Zhaoxiang Zhang 0001
NeurIPS7
2024 DDK: Distilling Domain Knowledge for Efficient Large Language Models
abstract
Despite the advanced intelligence abilities of large language models (LLMs) in various applications, they still face significant computational and storage demands. Knowledge Distillation (KD) has emerged as an effective strategy to improve the performance of a smaller LLM (i.e., the student model) by transferring knowledge from a high-performing LLM (i.e., the teacher model). Prevailing techniques in LLM distillation typically use a black-box model API to generate high-quality pretrained and aligned datasets, or utilize white-box distillation by altering the loss function to better transfer knowledge from the teacher LLM. However, these methods ignore the knowledge differences between the student and teacher LLMs across domains. This results in excessive focus on domains with minimal performance gaps and insufficient attention to domains with large gaps, reducing overall performance. In this paper, we introduce a new LLM distillation framework called DDK, which dynamically adjusts the composition of the distillation dataset in a smooth manner according to the domain performance differences between the teacher and student models, making the distillation process more stable and effective. Extensive evaluations show that DDK significantly improves the performance of student models, outperforming both continuously pretrained baselines and existing knowledge distillation methods by a large margin.
Yuanxing Zhang, Haoran Que, Ken Deng, Zhiqi Bai, Jie Liu 0047, Ge Zhang 0009, Jiakai Wang, Congnan Liu, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS10
2024 D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
abstract
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1\% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law.
Haoran Que, Ge Zhang 0009, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan 0003, Jie Fu 0001, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007
NeurIPS9
2024 BiDM: Pushing the Limit of Quantization for Diffusion Models
abstract
Diffusion models (DMs) have been significantly developed and widely used in various applications due to their excellent generative qualities. However, the expensive computation and massive parameters of DMs hinder their practical use in resource-constrained scenarios. As one of the effective compression approaches, quantization allows DMs to achieve storage saving and inference acceleration by reducing bit-width while maintaining generation performance. However, as the most extreme quantization form, 1-bit binarization causes the generation performance of DMs to face severe degradation or even collapse. This paper proposes a novel method, namely BiDM, for fully binarizing weights and activations of DMs, pushing quantization to the 1-bit limit. From a temporal perspective, we introduce the Timestep-friendly Binary Structure (TBS), which uses learnable activation binarizers and cross-timestep feature connections to address the highly timestep-correlated activation features of DMs. From a spatial perspective, we propose Space Patched Distillation (SPD) to address the difficulty of matching binary features during distillation, focusing on the spatial locality of image generation tasks and noise estimation networks. As the first work to fully binarize DMs, the W1A1 BiDM on the LDM-4 model for LSUN-Bedrooms 256$\times$256 achieves a remarkable FID of 22.74, significantly outperforming the current state-of-the-art general binarization methods with an FID of 59.44 and invalid generative samples, and achieves up to excellent 28.0 times storage and 52.7 times OPs savings.
Xingyu Zheng, Xianglong Liu 0001, Yichen Bian, Xudong Ma, Yulun Zhang 0001, Jiakai Wang, Jinyang Guo 0002, Haotong Qin
NeurIPS6
2024 Transferable Multimodal Attack on Vision-Language Pre-training Models
abstract
Vision-Language Pre-training (VLP) models have achieved remarkable success in practice, while easily being misled by adversarial attack. Though harmful, adversarial attacks are valuable in revealing the blind-spots of VLP models and promoting their robustness. However, existing adversarial attacking studies pay insufficient attention to the key roles of different modality-correlated features, leading to unsatisfactory transferable attacking performance. To tackle this issue, we propose the Transferable MultiModal (TMM) attack framework, which tailors both the modality consistency and modality discrepancy features. To promote transferability, we propose the attention-directed feature perturbation to disturb the modality-consistency features in critical attention regions. In light of the commonly employed cross-attention can represent the consistent features among diverse models, it is more possible to mislead the similar model perception for activating stronger transferability. For improving attacking ability, we proposed the orthogonal-guided feature heterogenization to guide the adversarial perturbation to contain more modality-discrepancy features in the encoded embeddings. Since VLP models rely more on aligned features among different modalities during decision-making, increasing the modality-discrepant could confuse the learned representation for better attacking ability. Extensive experiments under diverse settings demonstrate that the proposed TMM outperforms the comparisons by large margins, i.e., 20.47% improvements in transferable attacking ability on average. Moreover, we highlight that our TMM also shows outstanding attacking performance on large models, such as MiniGPT-4, Otter, etc.
Haodi Wang, Kai Dong 0001, Zhilei Zhu, Haotong Qin, Aishan Liu, Xiaolin Fang 0001, Jiakai Wang, Xianglong Liu 0001
SP7
2024 Generate Transferable Adversarial Physical Camouflages via Triplet Attention Suppression
Jiakai Wang, Xianglong Liu 0001, Zixin Yin, Jun Guo 0009, Haotong Qin, Qingtao Wu, Aishan Liu
Int. J. Comput. Vis.1
2024 Adversarial Examples Against WiFi Fingerprint-Based Localization in the Physical World
abstract
WiFi Fingerprint-based Localization (WFL) has recently achieved promising results in the bloom of deep learning techniques. Unfortunately, current studies reveal the great risks of deep-learning models when facing adversarial attacks, raising broader concerns about Deep-learning-based WiFi Fingerprint Localization Models (DFLMs). However, real-world adversarial attacks targeting DFLMs are not fully investigated, making it unclear how to counter this potential threat. In this paper, we take the first step to introduce adversarial examples into the physical world against DFLMs. Specifically, we propose a general attack method named Phy-Adv, consisting of a physical attenuation loss and a differentiable simulation module, the generated adversarial noise could be feasibly produced in the real world and make effects on DFLMs, i.e., misleading the DFLMs from the signal source end. Furthermore, aiming at countering this typical adversarial threat, we propose a Relaxant Multiple Batch Normalization (RMBN) approach, which alleviates the weak robustness of DFLMs by the data-end adaptive training-set segmenting and model-end multiple batch normalization designing. To demonstrate the de facto effectiveness of the proposed physical adversarial examples and the adversarial defense strategy, we conducted extensive experiments on 2 datasets, i.e., BHD and TUT, and multiple deep models, e.g., AlexNet, VGG, and ResNet. The experimental results strongly support that our Phy-Adv shows satisfactory adversarial attacking ability in the physical world, meanwhile, the RMBN enjoys considerable defense ability against the adversarial attacks.
Jiakai Wang, Ye Tao 0003, Wanting Liu, Yusheng Kong, Shaolin Tan, Rongen Yan, Xianglong Liu 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Hierarchical Perceptual Noise Injection for Social Media Fingerprint Privacy Protection
abstract
Billions of people share images from their daily lives on social media every day. However, their biometric information (e.g., fingerprints) could be easily stolen from these images. The threat of fingerprint leakage from social media has created a strong desire to anonymize shared images while maintaining image quality, since fingerprints act as a lifelong individual biometric password. To guard the fingerprint leakage, adversarial attack that involves adding imperceptible perturbations to fingerprint images have emerged as a feasible solution. However, existing works of this kind are either weak in black-box transferability or cause the images to have an unnatural appearance. Motivated by the visual perception hierarchy (i.e., high-level perception exploits model-shared semantics that transfer well across models while low-level perception extracts primitive stimuli that result in high visual sensitivity when a suspicious stimulus is provided), we propose FingerSafe, a hierarchical perceptual protective noise injection framework to address the above mentioned problems. For black-box transferability, we inject protective noises into the fingerprint orientation field to perturb the model-shared high-level semantics (i.e., fingerprint ridges). Considering visual naturalness, we suppress the low-level local contrast stimulus by regularizing the response of the Lateral Geniculate Nucleus. Our proposed FingerSafe is the first to provide feasible fingerprint protection in both digital (up to 94.12%) and realistic scenarios (Twitter and Facebook, up to 68.75%). Our code can be found at https://github.com/nlsde-safety-team/FingerSafe.
Huangxinxin Xu, Jiakai Wang, Ruixiao Xu, Aishan Liu, Fazhi He, Xianglong Liu 0001, Dacheng Tao
IEEE Trans. Image Process.3
2024 Improving Deepfake Detection Generalization by Invariant Risk Minimization
abstract
The abuse of deepfake techniques has raised serious concerns about social security and ethical problems, which motivates the development of deepfake detection. However, without fully addressing the domain gap issue, existing deepfake detection methods still show weak generalization ability among datasets belonging to different domains with domain-specific characteristics like identities and generation methods, limiting their practical applications. In this paper, we propose theInvariant Domain-oriented Deepfake Detection method (ID$_{3}$), which improves the generalization of deepfake detection on multiple domains through invariant risk minimization, a novel learning paradigm that addresses the domain gap problem by jointly training a purified invariant predictor and learning an aligned invariant representation. To train a purified invariant predictor, we design theDomain Refinement Data Augmentationstrategy with self-face-swapping and region-erasing approaches, which suppresses domain-specific features and encourages the models to focus on critical domain-invariant characteristics. To learn an aligned invariant representation, we propose theDomain Calibration Batch Normalizationapproach with multiple BN branches, which normalizes input features from different domains into aligned representations during both training and testing. Extensive experiments on multiple datasets demonstrate that our framework can boost the deepfake detection generalization ability and outperform other baselines by large margins. Our codes can be found here
Zixin Yin, Jiakai Wang, Yisong Xiao, Tianlin Li, Wenbo Zhou 0004, Aishan Liu, Xianglong Liu 0001
IEEE Trans. Multim.2
2024 BiFSMNv2: Pushing Binary Neural Networks for Keyword Spotting to Real-Network Performance
abstract
Deep neural networks, such as the deep-FSMN, have been widely studied for keyword spotting (KWS) applications while suffering expensive computation and storage. Therefore, network compression technologies such as binarization are studied to deploy KWS models on edge. In this article, we present a strong yet efficient binary neural network for KWS, namely, BiFSMNv2, pushing it to the real-network accuracy performance. First, we present a dual-scale thinnable 1-bit-architecture (DTA) to recover the representation capability of the binarized computation units by dual-scale activation binarization and liberate the speedup potential from an overall architecture perspective. Second, we also construct a frequency-independent distillation (FID) scheme for KWS binarization-aware training, which distills the high- and low-frequency components independently to mitigate the information mismatch between full-precision and binarized representations. Moreover, we propose the learning propagation binarizer (LPB), a general and efficient binarizer that enables the forward and backward propagation of binary KWS networks to be continuously improved through learning. We implement and deploy BiFSMNv2 on ARMv8 real-world hardware with a novel fast bitwise computation kernel (FBCK), which is proposed to fully use registers and increase instruction throughput. Comprehensive experiments show our BiFSMNv2 outperforms the existing binary networks for KWS by convincing margins across different datasets and achieves comparable accuracy with the full-precision networks (only a tiny 1.51% drop on Speech Commands V1-12). We highlight that benefiting from the compact architecture and optimized hardware kernel, BiFSMNv2 can achieve an impressive 25.1× speedup and 20.2× storage-saving on edge hardware.
Haotong Qin, Xudong Ma, Yifu Ding 0001, Yang Zhang 0088, Zejun Ma 0001, Jiakai Wang, Jie Luo 0004, Xianglong Liu 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Towards Benchmarking and Assessing Visual Naturalness of Physical World Adversarial Attacks
abstract
Physical world adversarial attack is a highly practical and threatening attack, which fools real world deep learning systems by generating conspicuous and maliciously crafted real world artifacts. In physical world attacks, evaluating naturalness is highly emphasized since human can easily detect and remove unnatural attacks. However, current studies evaluate naturalness in a case-by-case fashion, which suffers from errors, bias and inconsistencies. In this paper, we take the first step to benchmark and assess visual naturalness of physical world attacks, taking autonomous driving scenario as the first attempt. First, to benchmark attack naturalness, we contribute the first Physical Attack Naturalness (PAN) dataset with human rating and gaze. PAN verifies several insights for the first time: naturalness is (disparately) affected by contextual features (i.e., environmental and semantic variations) and correlates with behavioral feature (i.e., gaze signal). Second, to automatically assess attack naturalness that aligns with human ratings, we further introduce Dual Prior Alignment (DPA) network, which aims to embed human knowledge into model reasoning process. Specifically, DPA imitates human reasoning in naturalness assessment by rating prior alignment and mimics human gaze behavior by attentive prior alignment. We hope our work fosters researches to improve and automatically assess naturalness of physical world attacks. Our code and dataset can be found at https://github.com/zhangsn-19/PAN.
Gujun Chen, Pu Feng, Jiakai Wang, Aishan Liu, Xin Yi 0001, Xianglong Liu 0001
CVPR6
2023 X-Adv: Physical Adversarial Object Attacks against X-ray Prohibited Item Detection
Aishan Liu, Jun Guo 0009, Jiakai Wang, Siyuan Liang 0004, Renshuai Tao, Wenbo Zhou 0004, Cong Liu 0006, Xianglong Liu 0001, Dacheng Tao
USENIX Security Symposium3
2023 Diverse Sample Generation: Pushing the Limit of Generative Data-Free Quantization
abstract
Generative data-free quantization emerges as a practical compression approach that quantizes deep neural networks to low bit-width without accessing the real data. This approach generates data utilizing batch normalization (BN) statistics of the full-precision networks to quantize the networks. However, it always faces the serious challenges of accuracy degradation in practice. We first give a theoretical analysis that the diversity of synthetic samples is crucial for the data-free quantization, while in existing approaches, the synthetic data completely constrained by BN statistics experimentally exhibit severe homogenization at distribution and sample levels. This paper presents a generic Diverse Sample Generation (DSG) scheme for the generative data-free quantization, to mitigate detrimental homogenization. We first slack the statistics alignment for features in the BN layer to relax the distribution constraint. Then, we strengthen the loss impact of the specific BN layers for different samples and inhibit the correlation among samples in the generation process, to diversify samples from the statistical and spatial perspectives, respectively. Comprehensive experiments show that for large-scale image classification tasks, our DSG can consistently quantization performance on different neural architectures, especially under ultra-low bit-width. And data diversification caused by our DSG brings a general gain to various quantization-aware training and post-training quantization approaches, demonstrating its generality and effectiveness.
Haotong Qin, Yifu Ding 0001, Xiangguo Zhang, Jiakai Wang, Xianglong Liu 0001, Jiwen Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 A comprehensive evaluation framework for deep model robustness
abstract
Deep neural networks (DNNs) have achieved remarkable performance across a wide range of applications, while they are vulnerable to adversarial examples , which motivates the evaluation and benchmark of model robustness. However, current evaluations usually use simple metrics to study the performance of defenses, which are far from understanding the limitation and weaknesses of these defense methods. Thus, most proposed defenses are quickly shown to be attacked successfully, which results in the “arm race” phenomenon between attack and defense. To mitigate this problem, we establish a model robustness evaluation framework containing 23 comprehensive and rigorous metrics, which consider two key perspectives of adversarial learning (i.e., data and model). Through neuron coverage and data imperceptibility , we use data-oriented metrics to measure the integrity of test examples; by delving into model structure and behavior, we exploit model-oriented metrics to further evaluate robustness in the adversarial setting . To fully demonstrate the effectiveness of our framework, we conduct large-scale experiments on multiple datasets including CIFAR-10, SVHN, and ImageNet using different models and defenses with our open-source platform. Overall, our paper provides a comprehensive evaluation framework, where researchers could conduct comprehensive and fast evaluations using the open-source toolkit, and the analytical results could inspire deeper understanding and further improvement to the model robustness.
Jun Guo 0009, Jiakai Wang, Yuqing Ma, Xinghai Gao, Aishan Liu, Xianglong Liu 0001, Wenjun Wu 0001
Pattern Recognit.3
2022 Harnessing Perceptual Adversarial Patches for Crowd Counting
abstract
Crowd counting, which has been widely adopted for estimating the number of people in safety-critical scenes, is shown to be vulnerable to adversarial examples in the physical world (e.g., adversarial patches). Though harmful, adversarial examples are also valuable for evaluating and better understanding model robustness. However, existing adversarial example generation methods for crowd counting lack strong transferability among different black-box models, which limits their practicability for real-world systems. Motivated by the fact that attacking transferability is positively correlated to the model-invariant characteristics, this paper proposes the Perceptual Adversarial Patch (PAP) generation framework to tailor the adversarial perturbations for crowd counting scenes using the model-shared perceptual features. Specifically, we handcraft an adaptive crowd density weighting approach to capture the invariant scale perception features across various models and utilize the density guided attention to capture the model-shared position perception. Both of them are demonstrated to improve the attacking transferability of our adversarial patches. Extensive experiments show that our PAP could achieve state-of-the-art attacking performance in both the digital and physical world, and outperform previous proposals by large margins (at most +685.7 MAE and +699.5 MSE). Besides, we empirically demonstrate that adversarial training with our PAP can benefit the performance of vanilla models in alleviating several practical challenges in crowd counting scenarios, including generalization across datasets (up to -376.0 MAE and -354.9 MSE) and robustness towards complex backgrounds (up to -10.3 MAE and -16.4 MSE).
Shunchang Liu, Jiakai Wang, Aishan Liu, Yingwei Li 0002, Yijie Gao, Xianglong Liu 0001, Dacheng Tao
CCS2
2022 Defensive Patches for Robust Recognition in the Physical World
abstract
To operate in real-world high-stakes environments, deep learning systems have to endure noises that have been con-tinuously thwarting their robustness. Data-end defense, which improves robustness by operations on input data in-stead of modifying models, has attracted intensive attention due to its feasibility in practice. However, previous data-end defenses show low generalization against diverse noises and weak transferability across multiple models. Motivated by the fact that robust recognition depends on both local and global features, we propose a defensive patch generation framework to address these problems by helping mod-els better exploit these features. For the generalization against diverse noises, we inject class-specific identifiable patterns into a confined local patch prior, so that defensive patches could preserve more recognizable features towards specific classes, leading models for better recognition under noises. For the transferability across multiple models, we guide the defensive patches to capture more global fea-ture correlations within a class, so that they could activate model-shared global perceptions and transfer better among models. Our defensive patches show great potentials to im-prove application robustness in practice by simply sticking them around target objects. Extensive experiments show that we outperform others by large margins (improve 20+ % accuracy for both adversarial and corruption robustness on average in the digital and physical world).11Our codes are available at https://github.com/nlsde-safety-team/DefensivePatch.
Jiakai Wang, Zixin Yin, Aishan Liu, Renshuai Tao, Haotong Qin, Xianglong Liu 0001, Dacheng Tao
CVPR1
2022 Generating Transferable Adversarial Examples against Vision Transformers
abstract
Vision transformers (ViTs) are prevailing among several visual recognition tasks, therefore drawing intensive interest in generating adversarial examples against them. Different from CNNs, ViTs enjoy unique architectures, e.g., self-attention and image-embedding, which are commonly-shared features among various types of transformer-based models. However, existing adversarial methods suffer from weak transferable attacking ability due to the overlook of these architectural features. To address the problem, we propose an Architecture-oriented Transferable Attacking (ATA) framework to generate transferable adversarial examples by activating the uncertain attention and perturbing the sensitive embedding.Specifically, we first locate the patch-wise attentional regions that mostly affect model perception, therefore intensively activating the uncertainty of the attention mechanism and confusing the model decisions in turn.Furthermore, we search the pixel-wise attacking positions that are more likely to derange the embedded tokens using sensitive embedding perturbation, which could serve as a strong transferable attacking pattern.By jointly confusing the unique yet widely-used architectural features among transformer-based models, we can activate strong attacking transferability among diverse ViTs. Extensive experiments on large-scale dataset ImageNet using various popular transformers demonstrate that our ATA outperforms other baselines by large margins (at least +15% Attack Success Rate). Our code is available at https://github.com/nlsde-safety-team/ATA
Jiakai Wang, Zixin Yin, Ruihao Gong, Aishan Liu, Xianglong Liu 0001
ACM Multimedia2
2022 Universal Adversarial Patch Attack for Automatic Checkout Using Perceptual and Attentional Bias
abstract
Adversarial examples are inputs with imperceptible perturbations that easily mislead deep neural networks (DNNs). Recently, adversarial patch, with noise confined to a small and localized patch, has emerged for its easy feasibility in real-world scenarios. However, existing strategies failed to generate adversarial patches with strong generalization ability due to the ignorance of the inherent biases of models. In other words, the adversarial patches are always input-specific and fail to attack images from all classes or different models, especially unseen classes and black-box models. To address the problem, this paper proposes a bias-based framework to generate universal adversarial patches with strong generalization ability, which exploits the perceptual bias and attentional bias to improve the attacking ability. Regarding the perceptual bias, since DNNs are strongly biased towards textures, we exploit the hard examples which convey strong model uncertainties and extract a textural patch prior from them by adopting the style similarities. The patch prior is closer to decision boundaries and would promote attacks across classes. As for the attentional bias, motivated by the fact that different models share similar attention patterns towards the same image, we exploit this bias by confusing the model-shared similar attention patterns. Thus, the generated adversarial patches can obtain stronger transferability among different models. Taking Automatic Check-out (ACO) as the typical scenario, extensive experiments including white-box/black-box settings in both digital-world (RPC, the largest ACO related dataset) and physical-world scenario (Taobao and JD, the world's largest online shopping platforms) are conducted. Experimental results demonstrate that our proposed framework outperforms state-of-the-art adversarial patch attack methods.
Jiakai Wang, Aishan Liu, Xiao Bai 0001, Xianglong Liu 0001
IEEE Trans. Image Process.1
2021 Dual Attention Suppression Attack: Generate Adversarial Camouflage in Physical World
abstract
Deep learning models are vulnerable to adversarial examples. As a more threatening type for practical deep learning systems, physical adversarial examples have received extensive research attention in recent years. However, without exploiting the intrinsic characteristics such as model-agnostic and human-specific patterns, existing works generate weak adversarial perturbations in the physical world, which fall short of attacking across different models and show visually suspicious appearance. Motivated by the viewpoint that attention reflects the intrinsic characteristics of the recognition process, this paper proposes the Dual Attention Suppression (DAS) attack to generate visually-natural physical adversarial camouflages with strong transferability by suppressing both model and human attention. As for attacking, we generate transferable adversarial camouflages by distracting the model-shared similar attention patterns from the target to non-target regions. Meanwhile, based on the fact that human visual attention always focuses on salient items (e.g., suspicious distortions), we evade the human-specific bottom-up attention to generate visually-natural camouflages which are correlated to the scenario context. We conduct extensive experiments in both the digital and physical world for classification and detection tasks on up-to-date models (e.g., Yolo-V5) and demonstrate that our method outperforms state-of-the-art methods.1
Jiakai Wang, Aishan Liu, Zixin Yin, Shunchang Liu, Shiyu Tang, Xianglong Liu 0001
CVPR1
2021 Towards Real-world X-ray Security Inspection: A High-Quality Benchmark And Lateral Inhibition Module For Prohibited Items Detection
abstract
Prohibited items detection in X-ray images often plays an important role in protecting public safety, which often deals with color-monotonous and luster-insufficient objects, resulting in unsatisfactory performance. Till now, there have been rare studies touching this topic due to the lack of specialized high-quality datasets. In this work, we first present a High-quality X-ray (HiXray) security inspection image dataset, which contains 102,928 common prohibited items of 8 categories. It is the largest dataset of high quality for prohibited items detection, gathered from the real-world airport security inspection and annotated by professional security inspectors. Besides, for accurate prohibited item detection, we further propose the Lateral Inhibition Module (LIM) inspired by the fact that humans recognize these items by ignoring irrelevant information and focusing on identifiable characteristics, especially when objects are overlapped with each other. Specifically, LIM, the elaborately designed flexible additional module, suppresses the noisy information flowing maximumly by the Bidirectional Propagation (BP) module and activates the most identifiable charismatic, boundary, from four directions by Boundary Activation (BA) module. We evaluate our method extensively on HiXray and OPIXray and the results demonstrate that it outperforms SOTA detection methods.1
Renshuai Tao, Yanlu Wei, Xiangjian Jiang, Hainan Li, Haotong Qin, Jiakai Wang, Yuqing Ma, Libo Zhang 0001, Xianglong Liu 0001
ICCV6
2021 Adversarial Examples in Physical World
abstract
Although deep neural networks (DNNs) have already made fairly high achievements and a very wide range of impact, their vulnerability attracts lots of interest of researchers towards related studies about artificial intelligence (AI) safety and robustness this year. A series of works reveals that the current DNNs are always misled by elaborately designed adversarial examples. And unfortunately, this peculiarity also affects real-world AI applications and places them at potential risk. we are more interested in physical attacks due to their implementability in the real world. The study of physical attacks can effectively promote the application of AI techniques, which is of great significance to the security development of AI.
Jiakai Wang
IJCAI1
2021 Sequential alignment attention model for scene text recognition
Yan Wu 0013, Renshuai Tao, Jiakai Wang, Haotong Qin, Aishan Liu, Xianglong Liu 0001
J. Vis. Commun. Image Represent.4
2020 Bias-Based Universal Adversarial Patch Attack for Automatic Check-Out
Aishan Liu, Jiakai Wang, Xianglong Liu 0001, Bowen Cao, Chongzhi Zhang, Hang Yu 0016
ECCV (13)2
2020 A Parallel Implementation of Hypothesis-Oriented Multiple Hypothesis Tracking
abstract
Hypothesis-oriented Multiple Hypothesis Tracking (HOMHT) recursively generates hypotheses on the origins of measurements and manages them, therefore it is computationally intensive. To speed up HOMHT for tracking hundreds of targets in real time, we propose a parallel implementation of this algorithm which distributes hypotheses into independent worker threads residing in multiple CPU cores. The implementation in this paper is based on object-oriented programming: each hypothesis object manages its target data all by itself and the generation and pruning of a hypothesis is achieved by its copy constructor and destructor functions. We evaluate this method by tracking 150 targets through 3 heterogeneous sensors in real-time with 32-best hypotheses running in 1, 2, 4, 8, 16 and 32 worker threads respectively. The results validate the method's scalability in which measurement fusion latency is approximately inversely proportional to worker thread count. We also make a pressure test by tracking 500 targets through 3 sensors, and HOMHT is able to run concurrently in real-time with 32 worker threads.
Lin Wu 0006, Fei Wang 0014, Yongjun Xu 0001, Jiakai Wang
FUSION5