Yuxuan Zhang 0007

dblp:126/5240-7 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0002-1447-7120ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A survey on anomaly segmentation in urban scene understanding with image data
Yuxuan Zhang 0007, Shuchang Wang, Zhenbo Shi, Wei Yang 0011
Knowl. Based Syst.1
2025 Stop Diverse OOD Attacks: Knowledge Ensemble for Reliable Defense
abstract
Enhancing defense through model ensemble is an emerging trend, where the challenge lies in how to use ensemble knowledge to counter Out-of-Distribution (OOD) attacks. In this paper, we propose the Reliable Defense Ensemble (REE) to address this issue. REE optimizes the ensemble knowledge of models through aggregation and enhances multidimensional robust performance through collaboration. It employs the Dynamic Synergy Amplification for weight allocation and strategy adjustment. Furthermore, we design a new Kernel Anomaly Smoothing Detection Module, which detects anomalous attacks using a smoothing feature function based on Gaussian kernel mean embedding and a multi-layer feedback structure. Particularly, we build a framework that uses reinforcement learning to iteratively fine-tune the parameters of inter-model communication and consensus. Extensive experimental results show that REE outperforms current state-of-the-art methods by a large margin in defending against OOD attacks.
Zhenbo Shi, Yuxuan Zhang 0007, Shuchang Wang, Zhidong Yu, Wei Yang 0011, Liusheng Huang
AAAI3
2025 RP-PGD: Boosting Segmentation Robustness with a Region-and-Prototype Based Adversarial Attack
abstract
Adversarial attack and defense have been extensively explored in classification tasks, but their study in semantic segmentation remains limited. Moreover, current attacks fail to act as strong underlying attacks for adversarial training (AT), making it difficult to achieve segmentation robustness against strong attacks. In this paper, we present RP-PGD, a novel Region-and-Prototype based Projected Gradient Descent attack tailored to fool segmentation models. In particular, we propose a region-based attack, which leverages a spatial-temporal way to separate the pixels into three disjoint regions, and highlights the attack on the crucial True Region and Boundary Region. Moreover, we introduce a prototype-based attack to disrupt the feature space, further enhancing the attack capability. To boost the robustness of segmentation models, we inject adversaries generated by RP-PGD into the clean data and perform AT. Extensive experiments on multiple datasets showcase that RP-PGD generates adversaries with faster convergence and stronger attack effectiveness, surpassing state-of-the-art attacks by a large margin. Consequently, RP-PGD serves as a strong underlying attack for segmentation models to perform AT, assisting them in defending against a variety of strong attacks without incurring additional computational costs during inference.
Yuxuan Zhang 0007, Zhenbo Shi, Shuchang Wang, Wei Yang 0011, Shaowei Wang 0003, Yinxing Xue
AAAI1
2025 Tip the Scales: Achieving Balance in Adversarial Examples Across Modalities
abstract
In the field of multimodal learning, controlling the training of unimodal encoders from different perspectives is a primary approach to addressing Training Imbalance. However, the inherent capacity limitations of the modality affect the model’s capability. Therefore, generating adversarial examples that can achieve balanced transferability remains a challenging and perplexing problem. In this paper, we propose the InterModality Balanced Attack (MOBA) to address this problem. MOBA leverages Aggregated Modality Perturbation (AMP), which exploits the unbalanced effects of text and image perturbations to maximize the impact on the victim model. AMP capitalizes on the intrinsic feature connections between modalities during the optimization process, adjusting perturbations through Cross-Modality Discrepancy Loss to enhance attack success rates. Additionally, we devise the Transferability-Enhanced Evolution (TEE) to overcome the issue of diminished attack transferability due to model capacity limitations. TEE employs Transfer-Driven Optimization Loss to alleviate overfitting in single models, thereby enhancing the generalization ability.
Zhenbo Shi, Zhidong Yu, Yuxuan Zhang 0007, Shuchang Wang, Wei Yang 0011, Liusheng Huang
ICASSP3
2025 Leaving No OOD Instance Behind: Instance-Level OOD Fine-Tuning for Anomaly Segmentation
abstract
Out-of-distribution (OOD) fine-tuning has emerged as a promising approach for anomaly segmentation. Current OOD fine-tuning strategies typically employ global-level objectives, aiming to guide segmentation models to accurately predict a large number of anomaly pixels. However, these strategies often perform poorly on small anomalies. To address this issue, we propose an instance-level OOD fine-tuning framework, dubbed LNOIB (Leaving No OOD Instance Behind). We start by theoretically analyzing why global-level objectives fail to segment small anomalies. Building on this analysis, we introduce a simple yet effective instance-level objective. Moreover, we propose a feature separation objective to explicitly constrain the representations of anomalies, which are prone to be smoothed by their in-distribution (ID) surroundings. LNOIB integrates these objectives to enhance the segmentation of small anomalies and serves as a paradigm adaptable to existing OOD fine-tuning strategies, without introducing additional inference cost. Experimental results show that integrating LNOIB into various OOD fine-tuning strategies yields significant improvements, particularly in component-level results, highlighting its strength in comprehensive anomaly segmentation.
Yuxuan Zhang 0007, Zhenbo Shi, Shuchang Wang, Zhidong Yu, Shaowei Wang 0003, Wei Yang 0011
NeurIPS1
2025 On filling the intra-class and inter-class gaps for few-shot segmentation
Yuxuan Zhang 0007, Shuchang Wang, Zhenbo Shi, Wei Yang 0011
Expert Syst. Appl.1
2025 A Unified Perspective From Diffuse Deviation to Target Hijacking
abstract
In the developing field of visual object tracking, the robustness and resilience of detection modules against adversarial perturbations is critical. Traditional attacks have shown limitations in maintaining long-term deception, which is mainly reflected in that they are often only effective for a short period of time, or only have an impact on a specific single frame of images, rather than continuously and effectively mislead objects in continuous video sequences. Therefore, the tracker tends to quickly recover to the correct tracking state in the face of continuously changing adversarial attacks, showing robustness against occasional detection anomalies. In order to solve these problems, we propose Multi-Strategy Adversarial Attack (MSAA). MSAA imposes specific constraints on the decision-making ability of the model, regulating the priority of modification to candidate bounding boxes, including offset and size. In addition, the strategy to construct a predefined hijacking trajectory includes a direction-aware perturbation and a center matching scheme to hijack the feature-aware module to a predefined target. To the best of our knowledge, this is the first time that a unified perspective is adopted to address the problem from diffuse deviation to target hijacking. Our method not only enhances the persistence and concealment of attacks, but also achieves more precise control in multi-target scenarios, which has not been fully addressed in traditional adversarial attack methods. Experiments show that MSAA greatly outperforms state-of-the-art attack methods on multiple public datasets.
Zhenbo Shi, Zhidong Yu, Yuxuan Zhang 0007, Wei Yang 0011, Liusheng Huang
IEEE Trans. Dependable Secur. Comput.3
2024 GenSeg: On Generating Unified Adversary for Segmentation
Yuxuan Zhang 0007, Zhenbo Shi, Wei Yang 0011, Shuchang Wang, Shaowei Wang 0003, Yinxing Xue
IJCAI1
2023 BAProto: Boundary-Aware Prototype for High-quality Instance Segmentation
abstract
To date, the boundary quality remains unsatisfactory in instance segmentation, resulting in increasing attention to the mask refinement mechanism. Such a mechanism is supposed to be accurate, efficient and generic to the existing models. Yet, few methods met the three factors simultaneously. In this paper, to address this issue, we propose a Boundary-Aware Prototype (BAProto) for boundary refinement, which conducts pixel-wise prediction through the similarity of boundary representation and the specific prototype. Such a prototype is obtained by a memory unit for comprehensive learning. To our best knowledge, BAProto is the first approach that satisfies the above three factors at the same time. In particular, we elaborately design a three-phase segmentation loss, focusing on the learning of different regions to extract discriminating boundary representations for prototype establishment. Extensive experimental results show that BAProto is precise, efficient and model-agnostic on COCO and Cityscapes datasets.
Yuxuan Zhang 0007, Wei Yang 0011
ICME1
2023 FGNet: Towards Filling the Intra-class and Inter-class Gaps for Few-shot Segmentation
abstract
Current few-shot segmentation (FSS) approaches have made tremendous achievements based on prototypical learning techniques. However, due to the scarcity of the support data provided, FSS methods still suffer from the intra-class and inter-class gaps. In this paper, we propose a uniform network to fill both the gaps, termed FGNet. It consists of the novel design of a Self-Adaptive Module (SAM) to emphasize the query feature to generate an enhanced prototype for self-alignment. Such a prototype caters to each query sample itself since it contains the underlying intra-instance information, which gets around the intra-class appearance gap. Moreover, we design an Inter-class Feature Separation Module (IFSM) to separate the feature space of the target class from other classes, which contributes to bridging the inter-class gap. In addition, we present several new losses and a method termed B-SLIC, which help to further enhance the separation performance of FGNet. Experimental results show that FGNet reduces both the gaps for FSS by SAM and IFSM respectively, and achieves state-of-the-art performances on both PASCAL-5i and COCO-20i datasets compared with previous top-performing approaches.
Yuxuan Zhang 0007, Wei Yang 0011, Shaowei Wang 0003
IJCAI1
2022 BSOLO: Boundary-Aware One-Stage Instance Segmentation SOLO
abstract
Current one-stage instance segmentation methods ignore the boundary information of masks, resulting in coarse masks that are far from the ground truth. In this paper, we propose a boundary-aware method to refine boundary information, called BSOLO. The core idea of BSOLO is to design a Hungarian-Algorithm-based boundary loss to calculate matching costs between boundaries. This loss effectively measures the difference between boundaries and suits for boundary regression, contributing to generating refined instance masks with high-quality boundaries. Besides, we propose a Feature Fusion Network (FFN) to capture long-range dependency. Through constructing the relationship between pixels, such a module is beneficial for predicting masks with large or uncontinuous region. Furthermore, we introduce a Prototype Attention Module (PAM) for mask assembling through channel attention, which enhances informative features and spotlights important prototypes. To evaluate the performance of BSOLO, we conduct extensive experiments. Experimental results show that BSOLO achieves 39.3 AP on MS COCO test-dev2017, outperforming SOLO and other methods by a large margin. We hope that BSOLO broadens the perspective for designing more valid boundary constraints.
Yuxuan Zhang 0007, Wei Yang 0011
ICASSP1