VLDB 2026 Research / reviewers in the wild / expert
Xiao Yang 0028
dblp:57/3385-28
· DBLP profile ↗
37ranked-venue papers
8as first author
33since 2021 · last 2025
0000-0001-9502-9962ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 7 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STAIR: Improving Safety Alignment with Introspective ReasoningabstractEnsuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and susceptibility to jailbreak attacks, primarily due to their reliance on direct refusals for malicious queries. In this paper, we propose STAIR, a novel framework that integrates SafeTy Alignment with Itrospective Reasoning. We enable LLMs to identify safety risks through step-by-step analysis by self-improving chain-of-thought (CoT) reasoning with safety awareness. STAIR first equips the model with a structured reasoning capability and then advances safety alignment via iterative preference optimization on step-level reasoning data generated using our newly proposed Safety-Informed Monte Carlo Tree Search (SI-MCTS). Specifically, we design a theoretically grounded reward for outcome evaluation to seek balance between helpfulness and safety. We further train a process reward model on this data to guide test-time searches for improved responses. Extensive experiments show that STAIR effectively mitigates harmful outputs while better preserving helpfulness, compared to instinctive alignment strategies. With test-time scaling, STAIR achieves a safety performance comparable to Claude-3.5 against popular jailbreak attacks. We have open-sourced our code, datasets and models at https://github.com/thu-ml/STAIR. Yichi Zhang 0012, Zeyu Xia 0003, Zhengwei Fang, Xiao Yang 0028, Ranjie Duan, Yinpeng Dong, Jun Zhu 0001 |
ICML | 6 |
| 2025 | From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language ModelsabstractHallucination in large vision-language models (LVLMs) is a significant challenge, i.e., generating objects that are not present in the visual input, which significantly compromises the reliability of models. Recent studies often attribute hallucinations to a lack of visual understanding, yet ignore a more fundamental issue: the model's inability to effectively extract or decouple visual features. In this paper, we revisit the hallucinations in LVLMs from an architectural perspective, investigating whether the primary cause lies in the visual encoder (feature extraction) or the modal alignment module (feature decoupling). Motivated by our preliminary findings, we propose a parameter-efficient fine-tuning strategy, PATCH, to mitigate hallucinations in LVLMs. This plug-and-play method can be integrated into various LVLMs, leveraging adaptive virtual tokens to extract object features from bounding boxes, thereby addressing hallucinations stemming from inadequate feature decoupling. PATCH achieves state-of-the-art performance across multiple multi-modal hallucination datasets and demonstrates significant improvements in general capabilities. We hope this work provides deeper insights into the underlying causes of hallucinations in LVLMs, fostering further advancements and innovation in this field. The code will be available at https://github.com/YuyingShang/PATCH. Yuying Shang, Xinyi Zeng, Yutao Zhu 0001, Xiao Yang 0028, Zhengwei Fang |
ACM Multimedia | 4 |
| 2025 | A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking
Chang Liu 0077, Yinpeng Dong, Wenzhao Xiang 0001, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001, Yuefeng Chen, Yuan He 0011, Hui Xue 0001, Shibao Zheng |
Int. J. Comput. Vis. | 4 |
| 2025 | Face3DAdv: Exploiting Robust Adversarial 3D Patches on Physical Face Recognition
Xiao Yang 0028, Longlong Xu, Tianyu Pang, Yinpeng Dong, Yikai Wang 0001, Hang Su 0006, Jun Zhu 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D EnvironmentsabstractAdversarial attacks in 3D environments have emerged as a critical threat to the reliability of visual perception systems, particularly in safety-sensitive applications such as identity verification and autonomous driving. These attacks employ adversarial patches and 3D objects to manipulate deep neural network (DNN) predictions by exploiting vulnerabilities within complex scenes. Existing defense mechanisms, such as adversarial training and purification, primarily employ passive strategies to enhance robustness. However, these approaches often rely on pre-defined assumptions about adversarial tactics, limiting their adaptability in dynamic 3D settings. To address these challenges, we introduce Reinforced Embodied Active Defense (Rein-EAD), a proactive defense framework that leverages adaptive exploration and interaction with the environment to improve perception robustness in 3D adversarial contexts. By implementing a multi-step objective that balances immediate prediction accuracy with predictive entropy minimization, Rein-EAD optimizes defense strategies over a multi-step horizon. Additionally, Rein-EAD involves an uncertainty-oriented reward-shaping mechanism that facilitates efficient policy updates, thereby reducing computational overhead and supporting real-world applicability without the need for differentiable environments. Comprehensive experiments validate the effectiveness of Rein-EAD, demonstrating a substantial reduction in attack success rates while preserving standard accuracy across diverse tasks. Notably, Rein-EAD exhibits robust generalization to unseen and adaptive attacks, making it suitable for real-world complex tasks, including 3D object classification, face recognition and autonomous driving. By integrating proactive policy learning with embodied scene interaction, Rein-EAD establishes a scalable and adaptable approach for securing DNN-based perception systems in dynamic and adversarial 3D environments. Xiao Yang 0028, Lingxuan Wu, Lizhong Wang, Chengyang Ying, Hang Su 0006, Jun Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | CamoEnv: Transferable and environment-consistent adversarial camouflage in autonomous driving
Xiao Yang 0028, Hang Su 0006, Shibao Zheng |
Pattern Recognit. Lett. | 2 |
| 2025 | ANF: Crafting Transferable Adversarial Point Clouds via Adversarial Noise FactorizationabstractTransfer-based adversarial attacks involve generating adversarial point clouds in surrogate models and transferring them to other models to assess 3D model robustness. However, current methods rely too much on surrogate model parameters, limiting transferability. In this work, we use Shapley value to identify positive and negative features, guiding optimization of adversarial noise in feature space. To effectively mislead the 3D classifier, we factorize the adversarial noise into positive and negative noise, with the former keeping the features of the adversarial point cloud close to the negative features, and the latter and the adversarial noise moving it away from the positive features. Finally, a novel adversarial point cloud attack method with Adversarial Noise Factorization is proposed, which is abbreviated asANF. ANF simultaneously optimizes the adversarial noise and its positive and negative noise in the feature space, only relying on partial network parameters, which significantly reduces the reliance on the surrogate model and improves the transferability of the adversarial point cloud. Experiments on well-recognized benchmark datasets show that the transferability of adversarial point clouds generated by ANF could be improved by more than 26.7$\%$on average over state-of-the-art transfer-based adversarial attack methods. Hai Chen, Shu Zhao 0005, Xiao Yang 0028, Huanqian Yan, Yuan He 0011, Hui Xue 0001, Fulan Qian, Hang Su 0006 |
IEEE Trans. Big Data | 3 |
| 2025 | ImAdv: Transferable Implicit Adversarial Attack for 3D Object Detectors in Autonomous Drivingabstract3D adversarial attacks have garnered significant attention in the realm of autonomous driving security due to their high feasibility and multi-view effectiveness. However, existing 3D attacks have limited transferability, primarily due to their overfitting to surrogate models. To address this limitation, we introduce a novel 3D adversarial attack method based on implicit texture modeling, termed ImAdv, against 3D object detection models. Specifically, ImAdv utilizes a positional encoder and a MLP to map the 3D coordinates of an object's surface to the RGB color space, thereby reformulating the object's texture within an implicit framework. This method significantly reduces the parameter number for color modeling, thus mitigating overfitting and improving transferability. Furthermore, we propose two innovative techniques to enhance the transferability, Random Texture Reset (RandReset) and Texture Model Averaging. RandReset randomly restores portions of the adversarial texture, increasing the training set diversity and mitigating overfitting. Texture Model Averaging employs self-ensembling of multiple texture checkpoints during the training phase to reduce overfitting in the final texture model. Comprehensive experiments demonstrate the superiority of our methods, which outperform previous methods by 17.18% in average black-box attack success rate. Additionally, our method shows strong transferability and practicality in zero-shot cross-task attacks and physical attacks. Xiao Yang 0028, Hang Su 0006, Shu Zhao 0005, Shibao Zheng |
IEEE Trans. Big Data | 2 |
| 2024 | Towards Transferable Targeted 3D Adversarial Attack in the Physical WorldabstractCompared with transferable untargeted attacks, transferable targeted adversarial attacks could specify the mis-classification categories of adversarial samples, posing a greater threat to security-critical tasks. In the meanwhile, 3D adversarial samples, due to their potential of multi-view robustness, can more comprehensively identify weak-nesses in existing deep learning systems, possessing great application value. However, the field of transferable targeted 3D adversarial attacks remains vacant. The goal of this work is to develop a more effective technique that could generate transferable targeted 3D adversarial examples, filling the gap in this field. To achieve this goal, we design a novel framework named TT3D that could rapidly reconstruct from few multi-view images into Transferable Targeted 3D textured meshes. While existing mesh-based texture optimization methods compute gradients in the high-dimensional mesh space and easily fall into local optima, leading to unsatisfactory transferability and distinct distortions, TT3D innovatively performs dual optimization towards both feature grid and Multi-layer Perceptron (MLP) parameters in the grid-based NeRF space, which significantly enhances black-box transferability while enjoying naturalness. Experimental results show that TT3D not only exhibits superior cross-model transferability but also maintains considerable adaptability across different renders and vision tasks. More importantly, we produce 3D adversarial examples with 3D printing techniques in the real world and verify their robust performance under various scenarios. Yinpeng Dong, Shouwei Ruan, Xiao Yang 0028, Hang Su 0006, Xingxing Wei 0001 |
CVPR | 4 |
| 2024 | Rethinking Model Ensemble in Transfer-based Adversarial AttacksabstractIt is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks without any knowledge of the victim model. An effective strategy to improve the transferability is attacking an ensemble of models. However, previous works simply average the outputs of different models, lacking an in-depth analysis on how and why model ensemble methods can strongly improve the transferability. In this paper, we rethink the ensemble in adversarial attacks and define the common weakness of model ensemble with two properties: 1) the flatness of loss landscape; and 2) the closeness to the local optimum of each model. We empirically and theoretically show that both properties are strongly correlated with the transferability and propose a Common Weakness Attack (CWA) to generate more transferable adversarial examples by promoting these two properties. Experimental results on both image classification and object detection tasks validate the effectiveness of our approach to improving the adversarial transferability, especially when attacking adversarially trained models. We also successfully apply our method to attack a black-box large vision-language model -- Google's Bard, showing the practical effectiveness. Code is available at \url{https://github.com/huanranchen/AdversarialAttacks}. Huanran Chen, Yichi Zhang 0012, Yinpeng Dong, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001 |
ICLR | 4 |
| 2024 | Embodied Active Defense: Leveraging Recurrent Feedback to Counter Adversarial PatchesabstractThe vulnerability of deep neural networks to adversarial patches has motivated numerous defense strategies for boosting model robustness. However, the prevailing defenses depend on single observation or pre-established adversary information to counter adversarial patches, often failing to be confronted with unseen or adaptive adversarial attacks and easily exhibiting unsatisfying performance in dynamic 3D environments. Inspired by active human perception and recurrent feedback mechanisms, we develop Embodied Active Defense (EAD), a proactive defensive strategy that actively contextualizes environmental information to address misaligned adversarial patches in 3D real-world settings. To achieve this, EAD develops two central recurrent sub-modules, i.e., a perception module and a policy module, to implement two critical functions of active vision. These models recurrently process a series of beliefs and observations, facilitating progressive refinement of their comprehension of the target object and enabling the development of strategic actions to counter adversarial patches in 3D environments. To optimize learning efficiency, we incorporate a differentiable approximation of environmental dynamics and deploy patches that are agnostic to the adversary’s strategies. Extensive experiments demonstrate that EAD substantially enhances robustness against a variety of patches within just a few steps through its action policy in safety-critical tasks (e.g., face recognition and object detection), without compromising standard accuracy. Furthermore, due to the attack-agnostic characteristic, EAD facilitates excellent generalization to unseen attacks, diminishing the averaged attack success rate by 95% across a range of unseen adversarial attacks. Lingxuan Wu, Xiao Yang 0028, Yinpeng Dong, Liuwei Xie, Hang Su 0006, Jun Zhu 0001 |
ICLR | 2 |
| 2024 | Robust Classification via a Single Diffusion ModelabstractDiffusion models have been applied to improve adversarial robustness of image classifiers by purifying the adversarial noises or generating realistic data for adversarial training. However, diffusion-based purification can be evaded by stronger adaptive attacks while adversarial training does not perform well under unseen threats, exhibiting inevitable limitations of these methods. To better harness the expressive power of diffusion models, this paper proposes Robust Diffusion Classifier (RDC), a generative classifier that is constructed from a pre-trained diffusion model to be adversarially robust. RDC first maximizes the data likelihood of a given input and then predicts the class probabilities of the optimized input using the conditional likelihood estimated by the diffusion model through Bayes’ theorem. To further reduce the computational cost, we propose a new diffusion backbone called multi-head diffusion and develop efficient sampling strategies. As RDC does not require training on particular adversarial attacks, we demonstrate that it is more generalizable to defend against multiple unseen threats. In particular, RDC achieves $75.67%$ robust accuracy against various $\ell_\infty$ norm-bounded adaptive attacks with $\epsilon_\infty=8/255$ on CIFAR-10, surpassing the previous state-of-the-art adversarial training models by $+4.77%$. The results highlight the potential of generative classifiers by employing pre-trained diffusion models for adversarial robustness compared with the commonly studied discriminative classifiers. Huanran Chen, Yinpeng Dong, Xiao Yang 0028, Chengqi Duan, Hang Su 0006, Jun Zhu 0001 |
ICML | 4 |
| 2024 | Efficient Black-box Adversarial Attacks via Bayesian Optimization Guided by a Function PriorabstractThis paper studies the challenging black-box adversarial attack that aims to generate adversarial examples against a black-box model by only using output feedback of the model to input queries. Some previous methods improve the query efficiency by incorporating the gradient of a surrogate white-box model into query-based attacks due to the adversarial transferability. However, the localized gradient is not informative enough, making these methods still query-intensive. In this paper, we propose a Prior-guided Bayesian Optimization (P-BO) algorithm that leverages the surrogate model as a global function prior in black-box adversarial attacks. As the surrogate model contains rich prior information of the black-box one, P-BO models the attack objective with a Gaussian process whose mean function is initialized as the surrogate model's loss. Our theoretical analysis on the regret bound indicates that the performance of P-BO may be affected by a bad prior. Therefore, we further propose an adaptive integration strategy to automatically adjust a coefficient on the function prior by minimizing the regret bound. Extensive experiments on image classifiers and large vision-language models demonstrate the superiority of the proposed algorithm in reducing queries and improving attack success rates compared with the state-of-the-art black-box attacks. Code is available at https://github.com/yibo-miao/PBO-Attack. Shuyu Cheng, Yibo Miao, Yinpeng Dong, Xiao Yang 0028, Xiao-Shan Gao, Jun Zhu 0001 |
ICML | 4 |
| 2024 | Diffusion Models are Certifiably Robust ClassifiersabstractGenerative learning, recognized for its effective modeling of data distributions, offers inherent advantages in handling out-of-distribution instances, especially for enhancing robustness to adversarial attacks. Among these, diffusion classifiers, utilizing powerful diffusion models, have demonstrated superior empirical robustness. However, a comprehensive theoretical understanding of their robustness is still lacking, raising concerns about their vulnerability to stronger future attacks. In this study, we prove that diffusion classifiers possess $O(1)$ Lipschitzness, and establish their certified robustness, demonstrating their inherent resilience. To achieve non-constant Lipschitzness, thereby obtaining much tighter certified robustness, we generalize diffusion classifiers to classify Gaussian-corrupted data. This involves deriving the evidence lower bounds (ELBOs) for these distributions, approximating the likelihood using the ELBO, and calculating classification probabilities via Bayes' theorem. Experimental results show the superior certified robustness of these Noised Diffusion Classifiers (NDCs). Notably, we achieve over 80\% and 70\% certified robustness on CIFAR-10 under adversarial perturbations with \(\ell_2\) norms less than 0.25 and 0.5, respectively, using a single off-the-shelf diffusion model without any additional data. Huanran Chen, Yinpeng Dong, Shitong Shao, Zhongkai Hao, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 5 |
| 2024 | Improving Robustness of 3D Point Cloud Recognition from a Fourier PerspectiveabstractAlthough 3D point cloud recognition has achieved substantial progress on standard benchmarks, the typical models are vulnerable to point cloud corruptions, leading to security threats in real-world applications. To improve the corruption robustness, various data augmentation methods have been studied, but they are mainly limited to the spatial domain. As the point cloud has low information density and significant spatial redundancy, it is challenging to analyze the effects of corruptions. In this paper, we focus on the frequency domain to observe the underlying structure of point clouds and their corruptions. Through graph Fourier transform (GFT), we observe a correlation between the corruption robustness of point cloud recognition models and their sensitivity to different frequency bands, which is measured by the GFT spectrum of the model’s Jacobian matrix. To reduce the sensitivity and improve the corruption robustness, we propose Frequency Adversarial Training (FAT) that adopts frequency-domain adversarial examples as data augmentation to train robust point cloud recognition models against corruptions. Theoretically, we provide a guarantee of FAT on its out-of-distribution generalization performance. Empirically, we conduct extensive experiments with various network architectures to validate the effectiveness of FAT, which achieves the new state-of-the-art results. Yibo Miao, Yinpeng Dong, Jinlai Zhang, Lijia Yu, Xiao Yang 0028, Xiao-Shan Gao |
NeurIPS | 5 |
| 2024 | MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language ModelsabstractDespite the superior capabilities of Multimodal Large Language Models (MLLMs) across diverse tasks, they still face significant trustworthiness challenges. Yet, current literature on the assessment of trustworthy MLLMs remains limited, lacking a holistic evaluation to offer thorough insights into future improvements. In this work, we establish MultiTrust, the first comprehensive and unified benchmark on the trustworthiness of MLLMs across five primary aspects: truthfulness, safety, robustness, fairness, and privacy. Our benchmark employs a rigorous evaluation strategy that addresses both multimodal risks and cross-modal impacts, encompassing 32 diverse tasks with self-curated datasets. Extensive experiments with 21 modern MLLMs reveal some previously unexplored trustworthiness issues and risks, highlighting the complexities introduced by the multimodality and underscoring the necessity for advanced methodologies to enhance their reliability. For instance, typical proprietary models still struggle with the perception of visually confusing images and are vulnerable to multimodal jailbreaking and adversarial attacks; MLLMs are more inclined to disclose privacy in text and reveal ideological and cultural biases even when paired with irrelevant images in inference, indicating that the multimodality amplifies the internal risks from base LLMs. Additionally, we release a scalable toolbox for standardized trustworthiness research, aiming to facilitate future advancements in this important field. Code and resources are publicly available at: https://multi-trust.github.io/. Yichi Zhang 0012, Yitong Sun 0002, Chang Liu 0077, Zhengwei Fang, Huanran Chen, Xiao Yang 0028, Xingxing Wei 0001, Hang Su 0006, Yinpeng Dong, Jun Zhu 0001 |
NeurIPS | 9 |
| 2024 | Efficient Adversarial Attack Strategy Against 3D Object Detection in Autonomous Driving SystemsabstractThe reliability and robustness of 3D object detection play an instrumental role in the practical deployment of autonomous driving systems. Despite previous research indicating that adversarial examples can negatively affect 3D object detection models, leading to misinterpretations of the environment, these models still maintain the capability to detect the majority of objects within adversarially manipulated point clouds. To further probe into the adversarial robustness of these models, we propose an effective adversarial attack method named IoU-S attack in this paper. We meticulously formulate the adversarial loss to adversely affect the decision-making behavior (such as localization, etc.) of 3D object detection, thereby compromising its ability to accurately interpret the environment. Owing to the significant relevance of this adversarial loss to 3D object detection tasks, we have integrated the IoU-S attack into three attack paradigms: point cloud perturbation, detachment, and attachment. Comprehensive experiments on the widely accepted nuScenes dataset illustrate that the IoU-S attack outperforms existing attack methods in both white-box and black-box scenarios (https://github.com/haichen-ber/IoU-S-Attack). It reinforces its potential to serve as a valuable method in understanding and enhancing the robustness of 3D object detection models against adversarial attacks. Hai Chen, Huanqian Yan, Xiao Yang 0028, Hang Su 0006, Shu Zhao 0005, Fulan Qian |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Benchmarking Robustness of 3D Object Detection to Common Corruptions in Autonomous Drivingabstract3D object detection is an important task in autonomous driving to perceive the surroundings. Despite the excellent performance, the existing 3D detectors lack the robustness to real-world corruptions caused by adverse weathers, sensor noises, etc., provoking concerns about the safety and reliability of autonomous driving systems. To comprehensively and rigorously benchmark the corruption robustness of 3D detectors, in this paper we design 27 types of common corruptions for both LiDAR and camera inputs considering realworld driving scenarios. By synthesizing these corruptions on public datasets, we establish three corruption robustness benchmarks-KITTI-C, nuScenes-C, and Waymo-C. Then, we conduct large-scale experiments on 24 diverse 3D object detection models to evaluate their corruption robustness. Based on the evaluation results, we draw several important findings, including: 1) motion-level corruptions are the most threatening ones that lead to significant performance drop of all models; 2) LiDAR-camerafusion models demonstrate better robustness; 3) camera-only models are extremely vulnerable to image corruptions, showing the indispensability of LiDAR point clouds. We release the benchmarks and codes at https://github.com/thu-ml/3D_Corruptions_AD to be helpful for future studies. Yinpeng Dong, Caixin Kang, Jinlai Zhang, Yikai Wang 0001, Xiao Yang 0028, Hang Su 0006, Xingxing Wei 0001, Jun Zhu 0001 |
CVPR | 6 |
| 2023 | Towards Effective Adversarial Textured 3D Meshes on Physical Face RecognitionabstractFace recognition is a prevailing authentication solution in numerous biometric applications. Physical adversarial attacks, as an important surrogate, can identify the weak-nesses of face recognition systems and evaluate their ro-bustness before deployed. However, most existing physical attacks are either detectable readily or ineffective against commercial recognition systems. The goal of this work is to develop a more reliable technique that can carry out an end-to-end evaluation of adversarial robustness for commercial systems. It requires that this technique can simultaneously deceive black-box recognition models and evade defensive mechanisms. To fulfill this, we design adversarial textured 3D meshes (AT3D) with an elaborate topology on a human face, which can be 3D-printed and pasted on the attacker's face to evade the defenses. However, the mesh-based op-timization regime calculates gradients in high-dimensional mesh space, and can be trapped into local optima with un-satisfactory transferability. To deviate from the mesh-based space, we propose to perturb the low-dimensional coefficient space based on 3D Morphable Model, which signifi-cantly improves black-box transferability meanwhile enjoying faster search efficiency and better visual quality. Exten-sive experiments in digital and physical scenarios show that our method effectively explores the security vulnerabilities of multiple popular commercial services, including three recognition A PIs, four anti-spoofing A PIs, two prevailing mobile phones and two automated access control systems. Xiao Yang 0028, Chang Liu 0077, Longlong Xu, Yikai Wang 0001, Yinpeng Dong, Ning Chen 0002, Hang Su 0006, Jun Zhu 0001 |
CVPR | 1 |
| 2023 | Root Pose Decomposition Towards Generic Non-rigid 3D Reconstruction with Monocular VideosabstractThis work focuses on the 3D reconstruction of non-rigid objects based on monocular RGB video sequences. Concretely, we aim at building high-fidelity models for generic object categories and casually captured scenes. To this end, we do not assume known root poses of objects, and do not utilize category-specific templates or dense pose priors. The key idea of our method, Root Pose Decomposition (RPD), is to maintain a per-frame root pose transformation, meanwhile building a dense field with local transformations to rectify the root pose. The optimization of local transformations is performed by point registration to the canonical space. We also adapt RPD to multi-object scenarios with object occlusions and individual differences. As a result, RPD allows non-rigid 3D reconstruction for complicated scenarios containing objects with large deformations, complex motion patterns, occlusions, and scale diversities of different individuals. Such a pipeline potentially scales to diverse sets of objects in the wild. We experimentally show that RPD surpasses state-of-the-art methods on the challenging DAVIS, OVIS, and AMA datasets. We provide video results in https://rpd-share.github.io. Yikai Wang 0001, Yinpeng Dong, Fuchun Sun 0001, Xiao Yang 0028 |
ICCV | 4 |
| 2023 | On Evaluating Adversarial Robustness of Large Vision-Language ModelsabstractLarge vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models such as ChatGPT. Nonetheless, multimodal generation exacerbates safety concerns, since adversaries may successfully evade the entire system by subtly manipulating the most vulnerable modality (e.g., vision). To this end, we propose evaluating the robustness of open-source large VLMs in the most realistic and high-risk setting, where adversaries have only black-box system access and seek to deceive the model into returning the targeted responses. In particular, we first craft targeted adversarial examples against pretrained models such as CLIP and BLIP, and then transfer these adversarial examples to other VLMs such as MiniGPT-4, LLaVA, UniDiffuser, BLIP-2, and Img2Prompt. In addition, we observe that black-box queries on these VLMs can further improve the effectiveness of targeted evasion, resulting in a surprisingly high success rate for generating targeted responses. Our findings provide a quantitative understanding regarding the adversarial vulnerability of large VLMs and call for a more thorough examination of their potential security flaws before deployment in practice. Our project page: https://yunqing-me.github.io/AttackVLM/. Yunqing Zhao, Tianyu Pang, Xiao Yang 0028, Chongxuan Li, Ngai-Man Cheung |
NeurIPS | 4 |
| 2023 | AdvFAS: A robust face anti-spoofing framework against adversarial examplesabstractEnsuring the reliability of face recognition systems against presentation attacks necessitates the deployment of face anti-spoofing techniques. Despite considerable advancements in this domain, the ability of even the most state-of-the-art methods to defend against adversarial examples remains elusive. While several adversarial defense strategies have been proposed, they typically suffer from constrained practicability due to inevitable trade-offs between universality, effectiveness, and efficiency. To overcome these challenges, we thoroughly delve into the coupled relationship between adversarial detection and face anti-spoofing. Based on this, we propose a robust face anti-spoofing framework, namely AdvFAS, that leverages two coupled scores to accurately distinguish between correctly detected and wrongly detected face images. Extensive experiments demonstrate the effectiveness of our framework in a variety of settings, including different attacks, datasets, and backbones, meanwhile enjoying high accuracy on clean examples. Moreover, we successfully apply the proposed method to detect real-world adversarial examples. Xiao Yang 0028, Mingzhi Ma, Bihui Chen, Jianteng Peng, Yandong Guo, Zhao-Xia Yin, Hang Su 0006 |
Comput. Vis. Image Underst. | 2 |
| 2022 | Boosting Transferability of Targeted Adversarial Examples via Hierarchical Generative Networks
Xiao Yang 0028, Yinpeng Dong, Tianyu Pang, Hang Su 0006, Jun Zhu 0001 |
ECCV (4) | 1 |
| 2022 | Exploring Memorization in Adversarial Training
Yinpeng Dong, Xiao Yang 0028, Tianyu Pang, Zhijie Deng, Hang Su 0006, Jun Zhu 0001 |
ICLR | 3 |
| 2022 | DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Shilong Liu 0004, Feng Li 0040, Hao Zhang 0097, Xiao Yang 0028, Xianbiao Qi, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICLR | 4 |
| 2022 | Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionabstractThe trade-off between robustness and accuracy has been widely studied in the adversarial literature. Although still controversial, the prevailing view is that this trade-off is inherent, either empirically or theoretically. Thus, we dig for the origin of this trade-off in adversarial training and find that it may stem from the improperly defined robust error, which imposes an inductive bias of local invariance — an overcorrection towards smoothness. Given this, we advocate employing local equivariance to describe the ideal behavior of a robust model, leading to a self-consistent robust error named SCORE. By definition, SCORE facilitates the reconciliation between robustness and accuracy, while still handling the worst-case uncertainty via robust optimization. By simply substituting KL divergence with variants of distance metrics, SCORE can be efficiently minimized. Empirically, our models achieve top-rank performance on RobustBench under AutoAttack. Besides, SCORE provides instructive insights for explaining the overfitting phenomenon and semantic input gradients observed on robust models. Tianyu Pang, Xiao Yang 0028, Jun Zhu 0001, Shuicheng Yan |
ICML | 3 |
| 2022 | Towards generalizable detection of face forgery via self-guided model-agnostic learning
Xiao Yang 0028, Shilong Liu 0004, Yinpeng Dong, Hang Su 0006, Lei Zhang 0001, Jun Zhu 0001 |
Pattern Recognit. Lett. | 1 |
| 2021 | LiBRe: A Practical Bayesian Approach to Adversarial DetectionabstractDespite their appealing flexibility, deep neural networks (DNNs) are vulnerable against adversarial examples. Various adversarial defense strategies have been proposed to resolve this problem, but they typically demonstrate restricted practicability owing to unsurmountable compromise on universality, effectiveness, or efficiency. In this work, we propose a more practical approach, Lightweight Bayesian Refinement (LiBRe), in the spirit of leveraging Bayesian neural networks (BNNs) for adversarial detection. Empowered by the task and attack agnostic modeling under Bayes principle, LiBRe can endow a variety of pre-trained task-dependent DNNs with the ability of defending heterogeneous adversarial attacks at a low cost. We develop and integrate advanced learning techniques to make LiBRe appropriate for adversarial detection. Concretely, we build the few-layer deep ensemble variational and adopt the pre-training & fine-tuning workflow to boost the effectiveness and efficiency of LiBRe. We further provide a novel insight to realise adversarial detection-oriented uncertainty quantification without inefficiently crafting adversarial examples during training. Extensive empirical studies covering a wide range of scenarios verify the practicability of LiBRe. We also conduct thorough ablation studies to evidence the superiority of our modeling and learning strategies.1 Zhijie Deng, Xiao Yang 0028, Shizhen Xu, Hang Su 0006, Jun Zhu 0001 |
CVPR | 2 |
| 2021 | Unsupervised Part Segmentation Through Disentangling Appearance and ShapeabstractWe study the problem, of unsupervised discovery and segmentation of object parts, which, as an intermediate local representation, are capable of finding intrinsic object structure and providing more explainable recognition results. Recent unsupervised methods have greatly relaxed the dependency on annotated data which are costly to obtain, but still rely on additional information such as object segmentation mask or saliency map. To remove such a dependency and further improve the part segmentation performance, we develop a novel approach by disentangling the appearance and shape representations of object parts followed with reconstruction losses without using additional object mask information. To avoid degenerated solutions, a bottleneck block is designed to squeeze and expand the appearance representation, leading to a more effective disentanglement between geometry and appearance. Combined with a self-supervised part classification loss and an improved geometry concentration constraint, we can segment more consistent parts with semantic meanings. Comprehensive experiments on a wide variety of objects such as face, bird, and PASCAL VOC objects demonstrate the effectiveness of the proposed method. Shilong Liu 0004, Lei Zhang 0001, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001 |
CVPR | 3 |
| 2021 | Black-box Detection of Backdoor Attacks with Limited Information and DataabstractAlthough deep neural networks (DNNs) have made rapid progress in recent years, they are vulnerable in adversarial environments. A malicious backdoor could be embedded in a model by poisoning the training dataset, whose intention is to make the infected model give wrong predictions during inference when the specific trigger appears. To mitigate the potential threats of backdoor attacks, various backdoor detection and defense methods have been proposed. However, the existing techniques usually require the poisoned training data or access to the white-box model, which is commonly unavailable in practice. In this paper, we propose a black-box backdoor detection (B3D) method to identify backdoor attacks with only query access to the model. We introduce a gradient-free optimization algorithm to reverse-engineer the potential trigger for each class, which helps to reveal the existence of backdoor attacks. In addition to backdoor detection, we also propose a simple strategy for reliable predictions using the identified backdoored models. Extensive experiments on hundreds of DNN models trained on several datasets corroborate the effectiveness of our method under the black-box setting against various backdoor attacks. Yinpeng Dong, Xiao Yang 0028, Zhijie Deng, Tianyu Pang, Zihao Xiao 0002, Hang Su 0006, Jun Zhu 0001 |
ICCV | 2 |
| 2021 | Towards Face Encryption by Generating Adversarial Identity MasksabstractAs billions of personal data being shared through social media and network, the data privacy and security have drawn an increasing attention. Several attempts have been made to alleviate the leakage of identity information from face photos, with the aid of, e.g., image obfuscation techniques. However, most of the present results are either perceptually unsatisfactory or ineffective against face recognition systems. Our goal in this paper is to develop a technique that can encrypt the personal photos such that they can protect users from unauthorized face recognition systems but remain visually identical to the original version for human beings. To achieve this, we propose a targeted identity-protection iterative method (TIP-IM) to generate adversarial identity masks which can be overlaid on facial images, such that the original identities can be concealed without sacrificing the visual quality. Extensive experiments demonstrate that TIP-IM provides 95%+ protection success rate against various state-of-the-art face recognition models under practical test scenarios. Besides, we also show the practical and effective applicability of our method on a commercial API service. Xiao Yang 0028, Yinpeng Dong, Tianyu Pang, Hang Su 0006, Jun Zhu 0001, Yuefeng Chen, Hui Xue 0001 |
ICCV | 1 |
| 2021 | Bag of Tricks for Adversarial Training
Tianyu Pang, Xiao Yang 0028, Yinpeng Dong, Hang Su 0006, Jun Zhu 0001 |
ICLR | 2 |
| 2021 | Accumulative Poisoning Attacks on Real-time DataabstractCollecting training data from untrusted sources exposes machine learning services to poisoning adversaries, who maliciously manipulate training data to degrade the model accuracy. When trained on offline datasets, poisoning adversaries have to inject the poisoned data in advance before training, and the order of feeding these poisoned batches into the model is stochastic. In contrast, practical systems are more usually trained/fine-tuned on sequentially captured real-time data, in which case poisoning adversaries could dynamically poison each data batch according to the current model state. In this paper, we focus on the real-time settings and propose a new attacking strategy, which affiliates an accumulative phase with poisoning attacks to secretly (i.e., without affecting accuracy) magnify the destructive effect of a (poisoned) trigger batch. By mimicking online learning and federated learning on MNIST and CIFAR-10, we show that model accuracy significantly drops by a single update step on the trigger batch after the accumulative phase. Our work validates that a well-designed but straightforward attacking strategy can dramatically amplify the poisoning effects, with no need to explore complex techniques. Tianyu Pang, Xiao Yang 0028, Yinpeng Dong, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 2 |
| 2020 | Benchmarking Adversarial Robustness on Image ClassificationabstractDeep neural networks are vulnerable to adversarial examples, which becomes one of the most important research problems in the development of deep learning. While a lot of efforts have been made in recent years, it is of great significance to perform correct and complete evaluations of the adversarial attack and defense algorithms. In this paper, we establish a comprehensive, rigorous, and coherent benchmark to evaluate adversarial robustness on image classification tasks. After briefly reviewing plenty of representative attack and defense methods, we perform large-scale experiments with two robustness curves as the fair-minded evaluation criteria to fully understand the performance of these methods. Based on the evaluation results, we draw several important findings that can provide insights for future research, including: 1) The relative robustness between models can change across different attack configurations, thus it is encouraged to adopt the robustness curves to evaluate adversarial robustness; 2) As one of the most effective defense techniques, adversarial training can generalize across different threat models; 3) Randomization-based defenses are more robust to query-based black-box attacks. Yinpeng Dong, Qi-An Fu, Xiao Yang 0028, Tianyu Pang, Hang Su 0006, Zihao Xiao 0002, Jun Zhu 0001 |
CVPR | 3 |
| 2020 | Design and Interpretation of Universal Adversarial Patches in Face Detection
Xiao Yang 0028, Fangyun Wei, Hongyang Zhang 0001, Jun Zhu 0001 |
ECCV (17) | 1 |
| 2020 | Boosting Adversarial Training with Hypersphere EmbeddingabstractAdversarial training (AT) is one of the most effective defenses against adversarial attacks for deep learning models. In this work, we advocate incorporating the hypersphere embedding (HE) mechanism into the AT procedure by regularizing the features onto compact manifolds, which constitutes a lightweight yet effective module to blend in the strength of representation learning. Our extensive analyses reveal that AT and HE are well coupled to benefit the robustness of the adversarially trained models from several aspects. We validate the effectiveness and adaptability of HE by embedding it into the popular AT frameworks including PGD-AT, ALP, and TRADES, as well as the FreeAT and FastAT strategies. In the experiments, we evaluate our methods under a wide range of adversarial attacks on the CIFAR-10 and ImageNet datasets, which verifies that integrating HE can consistently enhance the model robustness for each AT framework with little extra computation. Tianyu Pang, Xiao Yang 0028, Yinpeng Dong, Taufik Xu, Jun Zhu 0001, Hang Su 0006 |
NeurIPS | 2 |
| 2018 | Recognizing Minimal Facial Sketch by Generating Photorealistic Faces With the Guidance of Descriptive AttributesabstractCross-modal sketch-photo recognition is of vital importance in law enforcement and public security. Most existing methods are dedicated to bridging the gap between the low-level visual features of sketches and photo images, which is limited due to intrinsic differences in pixel values. In this paper, based on the intuition that sketches and photo images are highly correlated in the semantic domain, we propose to jointly utilize the low-level visual features and high-level facial attributes to enhance the representation ability of sketches. More specifically, a Multi-Modal Conditional GAN (MMC-GAN) is proposed to generate face images for further face recognition based on the generated images. During training, an identity-preserving constraint is further introduced to improve the discriminative ability of the synthetic images. Extensive experiments demonstrate that the effectiveness of attribute-aided face synthesis and recognition. Xiao Yang 0028, Hang Su 0006, Qin Zhou 0002, Xinzhe Li 0002, Shibao Zheng |
ICASSP | 1 |