Hanwei Zhang 0001

dblp:144/8890-1 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-9690-6952ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability
abstract
Explainable AI (XAI) is crucial for building transparent and trustworthy machine learning systems, especially in high-stakes domains. Concept Bottleneck Models (CBMs) have emerged as a promising ante-hoc approach that provides interpretable, concept-level explanations by explicitly modeling human-understandable concepts. However, existing CBMs often suffer from poor locality faithfulness, failing to spatially align concepts with meaningful image regions, which limits their interpretability and reliability. In this work, we propose SL-CBM (CBM with Semantic Locality), a novel extension that enforces locality faithfulness by generating spatially coherent saliency maps at both concept and class levels. SL-CBM integrates a 1 × 1 convolutional layer with a cross-attention mechanism to enhance alignment between concepts, image regions, and final predictions. Unlike prior methods, SL-CBM produces faithful saliency maps inherently tied to the model’s internal reasoning, facilitating more effective debugging and intervention. Extensive experiments on image datasets demonstrate that SL-CBM substantially improves locality faithfulness, explanation quality, and intervention efficacy while maintaining competitive classification accuracy. Our ablation studies highlight the importance of contrastive and entropy-based regularization for balancing accuracy, sparsity, and faithfulness. Overall, SL-CBM bridges the gap between concept-based reasoning and spatial explainability, setting a new standard for interpretable and trustworthy concept-based models.
Hanwei Zhang 0001, Luo Cheng, Rui Wen 0002, Yang Zhang 0016, Lijun Zhang 0001, Holger Hermanns
AAAI1
2026 Revisiting Transferable Adversarial Images: Systemization, Evaluation, and New Insights
abstract
Transferable adversarial images raise critical security concerns for computer vision systems in real-world, black-box attack scenarios. Although many transfer attacks have been proposed, existing research lacks a systematic and comprehensive evaluation. In this paper, we systemize transfer attacks into five categories around the general machine learning pipeline and provide the first comprehensive evaluation, with 23 representative attacks against 11 representative defenses, including the recent, transfer-oriented defense and the real-world Google Cloud Vision. In particular, we identify two main problems of existing evaluations: (1) for attack transferability, lack of intra-category analyses with fair hyperparameter settings, and (2) for attack stealthiness, lack of diverse measures. Our evaluation results validate that these problems have indeed caused misleading conclusions and missing points, and addressing them leads to new, consensus-challenging insights, such as (1) an early attack, DI, even outperforms all similar follow-up ones, (2) the state-of-the-art (white-box) defense, DiffPure, is even vulnerable to (black-box) transfer attacks, and (3) even under the same $L_{p}$Lp constraint, different attacks yield dramatically different stealthiness results regarding diverse imperceptibility metrics, finer-grained measures, and a user study. We hope that our analyses will serve as guidance on properly evaluating transferable adversarial images and advance the design of attacks and defenses.
Zhengyu Zhao 0001, Hanwei Zhang 0001, Renjue Li, Ronan Sicre, Laurent Amsaleg, Michael Backes 0001, Qi Li 0002, Qian Wang 0002, Chao Shen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 MPAM-3DGS: Multi-Parametric Adversarial Manipulation for 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) is gaining popularity in fields such as robotics, autonomous driving, and virtual reality, due to its effectiveness and efficiency. Given that some tasks involve high risks, it is crucial to investigate the adversarial robustness of 3DGS and its downstream tasks—a topic that remains largely unexplored. In this study, we introduce a framework, Multi-Parametric Adversarial Manipulation for 3D Gaussian Splatting (MPAM-3DGS), that allows to attack 3DGS and its downstream tasks, such as object detection and classification, by perturbing a specified subset of parameters. Leveraging this framework, we examine the adversarial sensitivity of each 3DGS parameter and propose two strategies to attack multiple parameters based on our observations. To our knowledge, this is the first study to explore the adversarial robustness of 3DGS. Our experimental results demonstrate the effectiveness of our attacks on downstream tasks and the invisibility of perturbations in 3DGS. The code can be found at https://github.com/jiang-wenxiang/MPAM-3DGS.
Wenxiang Jiang 0002, Hanwei Zhang 0001, Zhongwen Guo, Tianao Zhang, Hao Wang 0003
ICASSP2
2025 HFE-RWKV: High-Frequency Enhanced RWKV Model for Efficient Left Ventricle Segmentation in Pediatric Echocardiograms
abstract
Automated ventricular function analysis can improve healthcare in resource-scarce areas, but current segmentation methods struggle with accurately delineating the irregular shape of the left ventricle due to a lack of emphasis on exploring the high-frequency target boundary features, and computational inefficiency is another concern. To address the two challenges, we turn to a novel and efficient basic structure, RWKV, and propose High-Frequency Enhanced RWKV (HFE-RWKV) for accurate and efficient left ventricle segmentation. Specifically, we propose the HFE-RWKV block as the encoder’s core to augment the high-frequency component, which is also the boundary area of the left ventricles in pediatric echocardiograms. In this way, the target boundaries can be explored more adequately during feature extraction. We propose space-frequency consistency loss to refine the shape of predicted masks further. Specifically, our new loss function incorporates spatial and frequency domain loss components to jointly refine predicted mask shapes in cases where current spatial-domain segmentation losses cannot be optimized further. Experiments on two public datasets prove our HFE-RWKV’s superiority in accuracy and efficiency. Specifically, our HFE-RWKV outperforms U-Mamba [12] by 2% in Dice Similarity Coefficient (DSC) while using only 67% of the parameters and 26% of the computational complexity. The code is available at https://github.com/yezizi1022/HFE-RWKV.
Hanwei Zhang 0001, Lijun Zhang 0001
ICASSP4
2025 Eidos revisited: Expanding Efficient, imperceptible adversarial attacks on 3D point clouds
Luo Cheng, Hanwei Zhang 0001, Qisong He, Wei Huang 0035, Renjue Li, Xiaowei Huang 0001, Holger Hermanns, Lijun Zhang 0001
J. Syst. Archit.2
2025 A Novel Robustness-Enhancing Adversarial Defense Approach to AI-Powered Sea State Estimation for Autonomous Marine Vessels
abstract
Sea state information is significant for the guide of maritime activities of autonomous vessels. The sea state estimation (SSE) model, powered by artificial intelligence (AI), has shown great effectiveness but is susceptible to malicious data attacks. These attacks can lead to significant declines in the system’s performance and result in incorrect predictions about the sea state. This study introduces SecureSSE, a strategy for protecting SSE models in autonomous marine vessels from adversarial attacks. This approach incorporates three main components: 1) the multiscale feature extraction learning (MFEL) module; 2) the feature convolution aggregation learning (FCAL) module; and 3) the perturbation examples training (PET) module. The PET module is specifically crafted to create perturbation examples that are in line with unaltered data, leveraging the capabilities of both the MFEL and FCAL modules to efficiently extract and integrate detailed features from ship motion data. Our proposed SecureSSE approach is shown to significantly improve the resilience of deep learning models against potential attacks. Through experimental testing, we have validated the effectiveness of this method in enhancing SSE. Additional ablation studies highlight the critical role of each module within the SecureSSE framework. To our knowledge, this is the first study to address adversarial attacks in this context and to propose a comprehensive defense mechanism for SSE systems in autonomous marine vessels.
Xu Cheng 0003, Fan Shi 0001, Hanwei Zhang 0001, Hongning Dai, Houxiang Zhang, Shengyong Chen
IEEE Trans. Syst. Man Cybern. Syst.4
2024 NeRFail: Neural Radiance Fields-Based Multiview Adversarial Attack
abstract
Adversarial attacks, i.e., generating adversarial perturbations with a small magnitude to deceive deep neural networks, are important for investigating and improving model trustworthiness. Traditionally, the topic was scoped within 2D images without considering 3D multiview information. Benefiting from Neural Radiance Fields (NeRF), one can easily reconstruct a 3D scene with a Multi-Layer Perceptron (MLP) from given 2D views and synthesize photo-realistic renderings of novel vantages. This opens up a door to discussing the possibility of undertaking to attack multiview NeRF network with downstream tasks from different rendering angles, which we denote Neural Radiance Fiels-based multiview adversarial Attack (NeRFail). The goal is, given one scene and a subset of views, to deceive the recognition results of agnostic view angles as well as given views. To do so, we propose a transformation mapping from pixels to 3D points such that our attack generates multiview adversarial perturbations by attacking a subset of images with different views, intending to prevent the downstream classifier from correctly predicting images rendered by NeRF from other views. Experiments show that our multiview adversarial perturbations successfully obfuscate the downstream classifier at both known and unknown views. Notably, when retraining another NeRF on the perturbed training data, we show that the perturbation can be inherited and reproduced. The code can be found at https://github.com/jiang-wenxiang/NeRFail.
Wenxiang Jiang 0002, Hanwei Zhang 0001, Xi Wang 0002, Zhongwen Guo, Hao Wang 0003
AAAI2
2024 Saliency Maps Give a False Sense of Explanability to Image Classifiers: An Empirical Evaluation across Methods and Metrics
Hanwei Zhang 0001, Felipe Torres Figueroa, Holger Hermanns
ACML1
2024 IPA-NeRF: Illusory Poisoning Attack Against Neural Radiance Fields
abstract
Neural Radiance Field (NeRF) represents a significant advancement in computer vision, offering implicit neural network-based scene representation and novel view synthesis capabilities. Its applications span diverse fields including robotics, urban mapping, autonomous navigation, virtual reality/augmented reality, etc., some of which are considered high-risk AI applications. However, despite its widespread adoption, the robustness and security of NeRF remain largely unexplored. In this study, we contribute to this area by introducing the Illusory Poisoning Attack against Neural Radiance Fields (IPA-NeRF). This attack involves embedding a hidden backdoor view into NeRF, allowing it to produce predetermined outputs, i.e. illusory, when presented with the specified backdoor view while maintaining normal performance with standard inputs. Our attack is specifically designed to deceive users or downstream models at a particular position while ensuring that any abnormalities in NeRF remain undetectable from other viewpoints. Experimental results demonstrate the effectiveness of our Illusory Poisoning Attack, successfully presenting the desired illusory on the specified viewpoint without impacting other views. Notably, we achieve this attack by introducing small perturbations solely to the training set. The code can be found at https://github.com/jiang-wenxiang/IPA-NeRF.
Wenxiang Jiang 0002, Hanwei Zhang 0001, Shuo Zhao 0001, Zhongwen Guo, Hao Wang 0003
ECAI2
2024 Traceability and Accountability by Construction
Julius Wenzel, Maximilian A. Köhl, Sarah Sterz, Hanwei Zhang 0001, Andreas Schmidt 0003, Christof Fetzer, Holger Hermanns
ISoLA (4)4
2024 Eidos: Efficient, Imperceptible Adversarial 3D Point Clouds
Hanwei Zhang 0001, Luo Cheng, Qisong He, Wei Huang 0035, Renjue Li, Ronan Sicre, Xiaowei Huang 0001, Holger Hermanns, Lijun Zhang 0001
SETTA1
2024 Opti-CAM: Optimizing saliency maps for interpretability
abstract
Methods based on class activation maps (CAM) provide a simple mechanism to interpret predictions of convolutional neural networks by using linear combinations of feature maps as saliency maps. By contrast, masking-based methods optimize a saliency map directly in the image space or learn it by training another network on additional data. In this work we introduce Opti-CAM, combining ideas from CAM-based and masking-based approaches. Our saliency map is a linear combination of feature maps, where weights are optimized per image such that the logit of the masked image for a given class is maximized. We also fix a fundamental flaw in two of the most common evaluation metrics of attribution methods. On several datasets, Opti-CAM largely outperforms other CAM-based approaches according to the most relevant classification metrics. We provide empirical evidence supporting that localization and classifier interpretability are not necessarily aligned.
Hanwei Zhang 0001, Felipe Torres, Ronan Sicre, Yannis Avrithis, Stéphane Ayache
Comput. Vis. Image Underst.1
2023 DP-Net: Learning Discriminative Parts for Image Recognition
abstract
This paper1presents Discriminative Part Network (DP-Net), a deep architecture with strong interpretation capabilities, which exploits a pretrained Convolutional Neural Network (CNN) combined with a part-based recognition module. This system learns and detects parts in the images that are discriminative among categories, without the need for fine-tuning the CNN, making it more scalable than other part-based models. While part-based approaches naturally offer interpretable representations, we propose explanations at image and category levels and introduce specific constraints on the part learning process to make them more discrimative.
Ronan Sicre, Hanwei Zhang 0001, Julien Dejasmin, Chiheb Daaloul, Stéphane Ayache, Thierry Artières
ICIP2
2021 Walking on the Edge: Fast, Low-Distortion Adversarial Examples
abstract
Adversarial examples of deep neural networks are receiving ever increasing attention because they help in understanding and reducing the sensitivity to their input. This is natural given the increasing applications of deep neural networks in our everyday lives. When white-box attacks are almost always successful, it is typically only the distortion of the perturbations that matters in their evaluation. In this work, we argue that speed is important as well, especially when considering that fast attacks are required by adversarial training. Given more time, iterative methods can always find better solutions. We investigate this speed-distortion trade-off in some depth and introduce a new attack called boundary projection (BP) that improves upon existing methods by a large margin. Our key idea is that the classification boundary is a manifold in the image space: we therefore quickly reach the boundary and then optimize distortion on this manifold.
Hanwei Zhang 0001, Yannis Avrithis, Teddy Furon, Laurent Amsaleg
IEEE Trans. Inf. Forensics Secur.1
2020 Smooth adversarial examples
abstract
Abstract This paper investigates the visual quality of the adversarial examples. Recent papers propose to smooth the perturbations to get rid of high frequency artifacts. In this work, smoothing has a different meaning as it perceptually shapes the perturbation according to the visual content of the image to be attacked. The perturbation becomes locally smooth on the flat areas of the input image, but it may be noisy on its textured areas and sharp across its edges.This operation relies on Laplacian smoothing, well-known in graph signal processing, which we integrate in the attack pipeline. We benchmark several attacks with and without smoothing under a white box scenario and evaluate their transferability. Despite the additional constraint of smoothness, our attack has the same probability of success at lower distortion.
Hanwei Zhang 0001, Yannis Avrithis, Teddy Furon, Laurent Amsaleg
EURASIP J. Inf. Secur.1