Zhenbo Shi

dblp:296/4458 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-4230-1552ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs
abstract
Yiming Huang, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng, Cuiyun Gao, Peiyi Han, Chuanyi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yiming Huang 0001, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng, Cuiyun Gao 0001, Peiyi Han, Chuanyi Liu
ACL (1)2
2026 A survey on anomaly segmentation in urban scene understanding with image data
Yuxuan Zhang 0007, Shuchang Wang, Zhenbo Shi, Wei Yang 0011
Knowl. Based Syst.3
2025 Stop Diverse OOD Attacks: Knowledge Ensemble for Reliable Defense
abstract
Enhancing defense through model ensemble is an emerging trend, where the challenge lies in how to use ensemble knowledge to counter Out-of-Distribution (OOD) attacks. In this paper, we propose the Reliable Defense Ensemble (REE) to address this issue. REE optimizes the ensemble knowledge of models through aggregation and enhances multidimensional robust performance through collaboration. It employs the Dynamic Synergy Amplification for weight allocation and strategy adjustment. Furthermore, we design a new Kernel Anomaly Smoothing Detection Module, which detects anomalous attacks using a smoothing feature function based on Gaussian kernel mean embedding and a multi-layer feedback structure. Particularly, we build a framework that uses reinforcement learning to iteratively fine-tune the parameters of inter-model communication and consensus. Extensive experimental results show that REE outperforms current state-of-the-art methods by a large margin in defending against OOD attacks.
Zhenbo Shi, Yuxuan Zhang 0007, Shuchang Wang, Zhidong Yu, Wei Yang 0011, Liusheng Huang
AAAI1
2025 AAKR: Adversarial Attack-based Knowledge Retention for Continual Semantic Segmentation
abstract
In the context of Continual Semantic Segmentation (CSS), replay-based methods tend to achieve better performance than knowledge distillation-based ones, as the former utilizes additional data to transfer old knowledge. However, this advantage is at the cost of necessitating additional space for storing the generative model and extra time for continual training. To address this predicament, we propose a novel CSS framework, namely Adversarial Attack-based Knowledge Retention (AAKR). The AKKR framework generates specific adversarial samples by adding images, and uses them to retain old knowledge. Specifically, we leverage adversarial attacks to generate adversarial images for incremental samples. By imposing additional constraints within these attacks, we enhance the transfer of old knowledge, thereby reinforcing the understanding of previously learned information. Furthermore, we design an attack probability module that adjusts adversarial attack directions based on training feedback. This module effectively encourages the new model to learn old knowledge from poorly protected classes, significantly improving knowledge transfer effectiveness. Our comprehensive experiments demonstrate the efficacy of AAKR, and showcase that AAKR surpasses state-of-the-art competitors on benchmark datasets.
Zhidong Yu, Jiajun Hu, Zhenbo Shi, Wei Yang 0011
AAAI4
2025 RP-PGD: Boosting Segmentation Robustness with a Region-and-Prototype Based Adversarial Attack
abstract
Adversarial attack and defense have been extensively explored in classification tasks, but their study in semantic segmentation remains limited. Moreover, current attacks fail to act as strong underlying attacks for adversarial training (AT), making it difficult to achieve segmentation robustness against strong attacks. In this paper, we present RP-PGD, a novel Region-and-Prototype based Projected Gradient Descent attack tailored to fool segmentation models. In particular, we propose a region-based attack, which leverages a spatial-temporal way to separate the pixels into three disjoint regions, and highlights the attack on the crucial True Region and Boundary Region. Moreover, we introduce a prototype-based attack to disrupt the feature space, further enhancing the attack capability. To boost the robustness of segmentation models, we inject adversaries generated by RP-PGD into the clean data and perform AT. Extensive experiments on multiple datasets showcase that RP-PGD generates adversaries with faster convergence and stronger attack effectiveness, surpassing state-of-the-art attacks by a large margin. Consequently, RP-PGD serves as a strong underlying attack for segmentation models to perform AT, assisting them in defending against a variety of strong attacks without incurring additional computational costs during inference.
Yuxuan Zhang 0007, Zhenbo Shi, Shuchang Wang, Wei Yang 0011, Shaowei Wang 0003, Yinxing Xue
AAAI2
2025 Tip the Scales: Achieving Balance in Adversarial Examples Across Modalities
abstract
In the field of multimodal learning, controlling the training of unimodal encoders from different perspectives is a primary approach to addressing Training Imbalance. However, the inherent capacity limitations of the modality affect the model’s capability. Therefore, generating adversarial examples that can achieve balanced transferability remains a challenging and perplexing problem. In this paper, we propose the InterModality Balanced Attack (MOBA) to address this problem. MOBA leverages Aggregated Modality Perturbation (AMP), which exploits the unbalanced effects of text and image perturbations to maximize the impact on the victim model. AMP capitalizes on the intrinsic feature connections between modalities during the optimization process, adjusting perturbations through Cross-Modality Discrepancy Loss to enhance attack success rates. Additionally, we devise the Transferability-Enhanced Evolution (TEE) to overcome the issue of diminished attack transferability due to model capacity limitations. TEE employs Transfer-Driven Optimization Loss to alleviate overfitting in single models, thereby enhancing the generalization ability.
Zhenbo Shi, Zhidong Yu, Yuxuan Zhang 0007, Shuchang Wang, Wei Yang 0011, Liusheng Huang
ICASSP1
2025 Leaving No OOD Instance Behind: Instance-Level OOD Fine-Tuning for Anomaly Segmentation
abstract
Out-of-distribution (OOD) fine-tuning has emerged as a promising approach for anomaly segmentation. Current OOD fine-tuning strategies typically employ global-level objectives, aiming to guide segmentation models to accurately predict a large number of anomaly pixels. However, these strategies often perform poorly on small anomalies. To address this issue, we propose an instance-level OOD fine-tuning framework, dubbed LNOIB (Leaving No OOD Instance Behind). We start by theoretically analyzing why global-level objectives fail to segment small anomalies. Building on this analysis, we introduce a simple yet effective instance-level objective. Moreover, we propose a feature separation objective to explicitly constrain the representations of anomalies, which are prone to be smoothed by their in-distribution (ID) surroundings. LNOIB integrates these objectives to enhance the segmentation of small anomalies and serves as a paradigm adaptable to existing OOD fine-tuning strategies, without introducing additional inference cost. Experimental results show that integrating LNOIB into various OOD fine-tuning strategies yields significant improvements, particularly in component-level results, highlighting its strength in comprehensive anomaly segmentation.
Yuxuan Zhang 0007, Zhenbo Shi, Shuchang Wang, Zhidong Yu, Shaowei Wang 0003, Wei Yang 0011
NeurIPS2
2025 On filling the intra-class and inter-class gaps for few-shot segmentation
Yuxuan Zhang 0007, Shuchang Wang, Zhenbo Shi, Wei Yang 0011
Expert Syst. Appl.3
2025 A Unified Perspective From Diffuse Deviation to Target Hijacking
abstract
In the developing field of visual object tracking, the robustness and resilience of detection modules against adversarial perturbations is critical. Traditional attacks have shown limitations in maintaining long-term deception, which is mainly reflected in that they are often only effective for a short period of time, or only have an impact on a specific single frame of images, rather than continuously and effectively mislead objects in continuous video sequences. Therefore, the tracker tends to quickly recover to the correct tracking state in the face of continuously changing adversarial attacks, showing robustness against occasional detection anomalies. In order to solve these problems, we propose Multi-Strategy Adversarial Attack (MSAA). MSAA imposes specific constraints on the decision-making ability of the model, regulating the priority of modification to candidate bounding boxes, including offset and size. In addition, the strategy to construct a predefined hijacking trajectory includes a direction-aware perturbation and a center matching scheme to hijack the feature-aware module to a predefined target. To the best of our knowledge, this is the first time that a unified perspective is adopted to address the problem from diffuse deviation to target hijacking. Our method not only enhances the persistence and concealment of attacks, but also achieves more precise control in multi-target scenarios, which has not been fully addressed in traditional adversarial attack methods. Experiments show that MSAA greatly outperforms state-of-the-art attack methods on multiple public datasets.
Zhenbo Shi, Zhidong Yu, Yuxuan Zhang 0007, Wei Yang 0011, Liusheng Huang
IEEE Trans. Dependable Secur. Comput.1
2024 Attacks on Continual Semantic Segmentation by Perturbing Incremental Samples
abstract
As an essential computer vision task, Continual Semantic Segmentation (CSS) has received a lot of attention. However, security issues regarding this task have not been fully studied. To bridge this gap, we study the problem of attacks in CSS in this paper. We first propose a new task, namely, attacks on incremental samples in CSS, and reveal that the attacks on incremental samples corrupt the performance of CSS in both old and new classes. Moreover, we present an adversarial sample generation method based on class shift, namely Class Shift Attack (CS-Attack), which is an offline and easy-to-implement approach for CSS. CS-Attack is able to significantly degrade the performance of models on both old and new classes without knowledge of the incremental learning approach, which undermines the original purpose of the incremental learning, i.e., learning new classes while retaining old knowledge. Experiments show that on the popular datasets Pascal VOC, ADE20k, and Cityscapes, our approach easily degrades the performance of currently popular CSS methods, which reveals the importance of security in CSS.
Zhidong Yu, Wei Yang 0011, Xike Xie, Zhenbo Shi
AAAI4
2024 TIKP: Text-to-Image Knowledge Preservation for Continual Semantic Segmentation
abstract
Continual Semantic Segmentation (CSS) is an emerging trend, where catastrophic forgetting has been a perplexing problem. In this paper, we propose a Text-to-Image Knowledge Preservation (TIKP) framework to address this issue. TIKP applies Text-to-Image techniques to CSS by automatically generating prompts and content adaptation. It extracts associations between the labels of seen data and constructs text-level prompts based on these associations, which are preserved and maintained at each incremental step. During training, these prompts generate correlated images to mitigate the catastrophic forgetting. Particularly, as the generated images may have different distributions from the original data, TIKP transfers the knowledge by a content adaption loss, which determines the role played by the generated images in incremental training based on the similarity. In addition, for the classifier, we use the previous model from a different perspective: misclassifying new classes into old objects instead of the background. We propose a knowledge distillation loss based on wrong labels, enabling us to attribute varying weights to individual objects during the distillation process. Extensive experiments conducted in the same setting show that TIKP outperforms state-of-the-art methods by a large margin on benchmark datasets.
Zhidong Yu, Wei Yang 0011, Xike Xie, Zhenbo Shi
AAAI4
2024 GenSeg: On Generating Unified Adversary for Segmentation
Yuxuan Zhang 0007, Zhenbo Shi, Wei Yang 0011, Shuchang Wang, Shaowei Wang 0003, Yinxing Xue
IJCAI2
2024 PFFAA: Prototype-based Feature and Frequency Alteration Attack for Semantic Segmentation
abstract
Recent research has confirmed the possibility of adversarial attacks on deep models. However, these methods typically assume that the surrogate model has access to the target domain, which is difficult to achieve in practical scenarios. To address this limitation, this paper introduces a novel cross-domain attack method tailored for semantic segmentation, named Prototype-based Feature and Frequency Alteration Attack (PFFAA). This approach empowers a surrogate model to efficiently deceive the black-box victim model without requiring access to the target data. Specifically, through limited queries on the victim model, bidirectional relationships are established between the target classes of the victim model and the source classes of the surrogate model, enabling the extraction of prototypes for these classes. During the attack process, the features of each source class are perturbed to move these features away from their respective prototypes. Moreover, we propose substituting frequency information from images used to train the surrogate model into the frequency domain of the test images to modify texture and structure, thus further enhancing the attack efficacy. Experimental results across multiple datasets and victim models validate that PFFAA achieves state-of-the-art performances.
Zhidong Yu, Zhenbo Shi, Wei Yang 0011
ACM Multimedia2
2023 Reinforcement Learning-based Adversarial Attacks on Object Detectors using Reward Shaping
abstract
In the field of object detector attacks, previous methods primarily rely on fixed gradient optimization or patch-based cover techniques, often leading to suboptimal attack performance and excessive distortions. To address these limitations, we propose a novel attack method, Interactive Reinforcement-based Sparse Attack (IRSA), which employs Reinforcement Learning (RL) to discover the vulnerabilities of object detectors and systematically generate erroneous results. Specifically, we formulate the process of seeking optimal margins for adversarial examples as a Markov Decision Process (MDP). We tackle the RL convergence difficulty through innovative reward functions and a composite optimization method for effective and efficient policy training. Moreover, the perturbations generated by IRSA are more subtle and difficult to detect while requiring less computational effort. Our method also demonstrates strong generalization capabilities against various object detectors. In summary, IRSA is a refined, efficient, and scalable interactive, iterative, end-to-end algorithm.
Zhenbo Shi, Wei Yang 0011, Zhenbo Xu, Zhidong Yu, Liusheng Huang
ACM Multimedia1
2022 Shape Prior Guided Attack: Sparser Perturbations on 3D Point Clouds
abstract
Deep neural networks are extremely vulnerable to malicious input data. As 3D data is increasingly used in vision tasks such as robots, autonomous driving and drones, the internal robustness of the classification models for 3D point cloud has received widespread attention. In this paper, we propose a novel method named SPGA (Shape Prior Guided Attack) to generate adversarial point cloud examples. We use shape prior information to make perturbations sparser and thus achieve imperceptible attacks. In particular, we propose a Spatially Logical Block (SLB) to apply adversarial points through sliding in the oriented bounding box. Moreover, we design an algorithm called FOFA for this type of task, which further refines the adversarial attack in the process of breaking down complicated problems into sub-problems. Compared with the methods of global perturbation, our attack method consumes significantly fewer computations, making it more efficient. Most importantly of all, SPGA can generate examples with a higher attack success rate (even in a defensive situation), less perturbation budget and stronger transferability.
Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Zhidong Yu, Liusheng Huang
AAAI1
2022 AtHom: Two Divergent Attentions Stimulated By Homomorphic Training in Text-to-Image Synthesis
abstract
Image generation from text is a challenging and ill-posed task. Images generated from previous methods usually have low semantic consistency with texts and the achieved resolution is limited. To generate semantically consistent high-resolution images, we propose a novel method named AtHom, in which two attention modules are developed to extract the relationships from both independent modality and unified modality. The first is a novel Independent Modality Attention Module (IAM), which is presented to find out semantically important areas in generated images and to extract the informative context in texts. The second is a new module named Unified Semantic Space Attention Module (UAM), which is utilized to find out the relationships between extracted text context and essential areas in generated images. In particular, to bring the semantic features of texts and images closer in a unified semantic space, AtHom incorporates a homomorphic training mode by exploiting an extra discriminator to distinguish between two different modalities. Extensive experiments show that our AtHom surpasses previous methods by large margins.
Zhenbo Shi, Zhi Chen 0026, Zhenbo Xu, Wei Yang 0011, Liusheng Huang
ACM Multimedia1
2021 VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant Keypoints
abstract
In this paper, we propose VK-Net, a neural network that learns to discover a set of category-specific keypoints from a single point cloud in an unsupervised manner. VK-Net is able to generate semantically consistent and rotation invariant keypoints across objects of the same category and different views. Particularly, we find that utilizing learned keypoints for the task of point cloud registration outperforms other traditional and learning-based approaches. Given the paired source and target point clouds, we can construct keypoint correspondences from learned keypoints using VK-Net. These keypoint correspondences are then employed to calculate a good pose initialization, after which an ICP is utilized to refine the registration. Extensive experiments on the ShapeNet dataset demonstrate that our model outperforms the state-of-the-art methods by a large margin.
Zhi Chen 0026, Wei Yang 0011, Zhenbo Xu, Zhenbo Shi, Liusheng Huang
ICASSP4
2021 Mask4D: 4D Convolution Network for Light Field Occlusion Removal
abstract
Current light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D, which is a 4D convolution-based encoder-decoder network. We propose to keep the spatial layout of SAIs and construct all SAIs as a 5D input tensor to fully exploit the spatial connection information between SAIs. In particular, except for center view reconstruction, we jointly predict the occlusion mask to disentangle the occlusion mask from the occluded content. Extensive evaluations demonstrate that our Mask4D surpasses the state-of-the-art approaches across different datasets. Moreover, visualizations show that Mask4D predicts the occlusion mask precisely and the reconstructed center view looks more realistic than other approaches. Our code will be publicly available.
Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Zhenbo Shi, Liusheng Huang
ICASSP5
2021 Adversarial Attacks on Object Detectors with Limited Perturbations
abstract
Deep convolutional neural networks are widely witnessed vulnerable to adversarial attacks. Recently, great progress has been achieved in attacking object detectors. However, current attacks neglect the practical utility and rely on global perturbations on the target image with a large number of patches or pixels. In this paper, we present a novel attack framework named DTTACK to fool both one-stage and two-stage object detectors with limited perturbations. A novel divergent patch shape consisting of four intersecting lines is proposed to effectively affect deep convolutional feature extraction with limited pixels. In particular, we introduce an instance-aware heat map as a self-attention module to help DTTACK focus on salient object areas, which further improves the attacking performance. Extensive experiments on PASCAL-VOC, MS-COCO, as well as an online detection system demonstrate that DTTACK surpasses the state-of-the-art methods by large margins.
Zhenbo Shi, Wei Yang 0011, Zhenbo Xu, Zhi Chen 0026, Liusheng Huang
ICASSP1
2021 Continuous Copy-Paste for One-stage Multi-object Tracking and Segmentation
abstract
Current one-step multi-object tracking and segmentation (MOTS) methods lag behind recent two-step methods. By separating the instance segmentation stage from the tracking stage, two-step methods can exploit non-video datasets as extra data for training instance segmentation. Moreover, instances belonging to different IDs on different frames, rather than limited numbers of instances in raw consecutive frames, can be gathered to allow more effective hard example mining in the training of trackers. In this paper, we bridge this gap by presenting a novel data augmentation strategy named continuous copy-paste (CCP). Our intuition behind CCP is to fully exploit the pixel-wise annotations provided by MOTS to actively increase the number of instances as well as unique instance IDs in training. Without any modifications to frameworks, current MOTS methods achieve significant performance gains when trained with CCP. Based on CCP, we propose the first effective one-stage online MOTS method named CCPNet, which generates instance masks as well as the tracking results in one shot. Our CCPNet surpasses all state-of-the-art methods by large margins (3.8% higher sMOTSA and 4.1% higher MOTSA for pedestrians on the KITTI MOTS Validation) and ranks 1st on the KITTI MOTS leaderboard. Evaluations across three datasets also demonstrate the effectiveness of both CCP and CCPNet. Our codes are publicly available at: https://github.com/detectRecog/CCP.
Zhenbo Xu, Ajin Meng, Zhenbo Shi, Wei Yang 0011, Zhi Chen 0026, Liusheng Huang
ICCV3
2021 Private FLI: Anti-Gradient Leakage Recovery Data Privacy Architecture
abstract
While machine learning brings convenience, it also faces the issue of data privacy. For privacy issues, most researches focus on implementing homomorphic encryption or differential privacy to protect data, while ignoring the potential threats caused by the leakage of model parameters. However, a malicious attacker can still recover sensitive data information through model parameters. On the one hand, traditional methods cannot take both high accuracy and low computation time into account. On the other hand, they cannot resist the reconstruction attack from the model's parameter. In order to address this problem, this paper designs a privacy protection framework named FLI, which is inspired by public key infrastructure. In FLI, all participants and the server are trained and aggregated under one framework based on federated learning, which includes key exchange and shares with the idea of homomorphic encryption. Under the algorithm we design, the malicious adversary cannot recover effective information after obtaining the transformed parameters, while the server can still perform effective parameter aggregation. To evaluate the performance of FLI, we conduct extensive experiments. The experimental results show that the computation time is within an acceptable range while ensuring high accuracy.
Huichao Wang, Wei Yang 0011, Bangzhou Xin, Yangyang Geng, Zhenbo Shi, Liusheng Huang
IJCNN5
2021 AggNet for Self-supervised Monocular Depth Estimation: Go An Aggressive Step Furthe
abstract
Without appealing to exhaustive labeled data, self-supervised monocular depth estimation (MDE) plays a fundamental role in computer vision. Previous methods usually adopt a one-stage MDE network, which is insufficient to achieve high performance. In this paper, we dig deep into this task to propose an aggressive framework termed AggNet. The framework is based on a training-only progressive two-stage module to perform pseudo counter-surveillance as well as a simple yet effective dual-warp loss function between image pairs. In particular, we first propose a residual module, which follows the MDE network to learn a refined depth. The residual module takes both the initial depth generated from MDE and the initial color image as input to generate refined depth with residual depth learning. Then, the refined depth is leveraged to supervise the initial depth simultaneously during the training period. For inference, only the MDE network is retained to regress depth from a single image, which gains better performance without introducing extra computation. In addition to self-distillation loss, a simple yet effective dual-warp consistency loss is introduced to encourage the MDE network to keep depth consistency between stereo image pairs. Extensive experiments show that our AggNet achieves state-of-the-art performance on the KITTI and Make3D datasets.
Zhi Chen 0026, Xiaoqing Ye, Liang Du 0004, Wei Yang 0011, Liusheng Huang, Xiao Tan 0001, Zhenbo Shi, Fumin Shen, Errui Ding
ACM Multimedia7