Tao Sun 0009

dblp:74/3590-9 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-6926-0543ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Backdooring Vision-Language Models with Out-Of-Distribution Data
abstract
The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to generate detailed text descriptions from visual inputs. Despite their growing importance, the security of VLMs, particularly against backdoor attacks, is under explored. Moreover, prior works often assume attackers have access to the original training data, which is often unrealistic. In this paper, we address a more practical and challenging scenario where attackers must rely solely on Out-Of-Distribution (OOD) data. We introduce VLOOD (Backdoor Vision-Language Models using Out-of-Distribution Data), a novel approach with two key contributions: (1) demonstrating backdoor attacks on VLMs in complex image-to-text tasks while minimizing degradation of the original semantics under poisoned inputs, and (2) proposing innovative techniques for backdoor injection without requiring any access to the original training data. Our evaluation on image captioning and visual question answering (VQA) tasks confirms the effectiveness of VLOOD, revealing a critical security vulnerability in VLMs and laying the foundation for future research on securing multimodal models against sophisticated threats.
Weimin Lyu, Jiachen Yao, Saumya Gupta, Lu Pang 0006, Tao Sun 0009, Lingjie Yi, Lijie Hu, Haibin Ling, Chao Chen 0012
ICLR5
2025 PivotAlign: Improve Semi-Supervised Learning by Learning Intra-Class Heterogeneity and Aligning with Pivots
abstract
Self-supervised learning plays an important role in current state-of-the-art semi-supervised learning (SSL) methods. These methods learn inter-class heterogeneity among data and generate pseudo-labels based on class level representations. However, they often neglect intra-class heterogeneity, resulting in the under-exploitation of finer-grained semantic relationships within classes. To address this limitation, we introduce PivotAlign, a novel SSL approach that aims to 1) learn hierarchical representations to detect both interclass and intra-class semantic relationships, and 2) refine pseudo-labels based on learned representations with a class-debiasing strategy. Specifically, we first learn a set of pivots as sub-prototypes of classes. We then train representations so that features align with the assigned pivot and are hierarchically grouped based on both inter-class and intra-class heterogeneity. This allows us to capture both inter-class and intra-class semantic relationships among data and leverage them to better assign and refine pseudo-labels. Additionally, since SSL methods are prone to bias toward classes that are easier to learn, we further re-balance class predictions to alleviate this class bias. We demonstrate the effectiveness of PivotAlign on various SSL benchmarks, where PivotAlign achieves state-of-the-art performances. The source code will be released upon publication of the work.
Lingjie Yi, Tao Sun 0009, Yikai Zhang 0003, Songzhu Zheng, Weimin Lyu, Haibin Ling, Chao Chen 0012
WACV2
2023 Domain Adaptation with Adversarial Training on Penultimate Activations
abstract
Enhancing model prediction confidence on target data is an important objective in Unsupervised Domain Adaptation (UDA). In this paper, we explore adversarial training on penultimate activations, i.e., input features of the final linear classification layer. We show that this strategy is more efficient and better correlated with the objective of boosting prediction confidence than adversarial training on input images or intermediate features, as used in previous works. Furthermore, with activation normalization commonly used in domain adaptation to reduce domain gap, we derive two variants and systematically analyze the effects of normalization on our adversarial training. This is illustrated both in theory and through empirical analysis on real adaptation tasks. Extensive experiments are conducted on popular UDA benchmarks under both standard setting and source-data free setting. The results validate that our method achieves the best scores against previous arts. Code is available at https://github.com/tsun/APA.
Tao Sun 0009, Cheng Lu 0006, Haibin Ling
AAAI1
2023 Backdoor Cleansing with Unlabeled Data
abstract
Due to the increasing computational demand of Deep Neural Networks (DNNs), companies and organizations have begun to outsource the training process. However, the externally trained DNNs can potentially be backdoor attacked. It is crucial to defend against such attacks, i.e., to postprocess a suspicious model so that its backdoor behavior is mitigated while its normal prediction power on clean inputs remain uncompromised. To remove the abnormal backdoor behavior, existing methods mostly rely on additional labeled clean samples. However, such requirement may be unrealistic as the training data are often unavailable to end users. In this paper, we investigate the possibility of circumventing such barrier. We propose a novel defense method that does not require training labels. Through a carefully designed layer-wise weight reinitialization and knowledge distillation, our method can effectively cleanse backdoor behaviors of a suspicious network with negligible compromise in its normal behavior. In experiments, we show that our method, trained without labels, is on-par with state-of-the-art defense methods trained using labels. We also observe promising defense results even on out-of-distribution data. This makes our method very practical. Code is available at: https://github.com/luluppang/BCU.
Lu Pang 0006, Tao Sun 0009, Haibin Ling, Chao Chen 0012
CVPR2
2023 Local Context-Aware Active Domain Adaptation
abstract
Active Domain Adaptation (ADA) queries the labels of a small number of selected target samples to help adapting a model from a source domain to a target domain. The local context of queried data is important, especially when the domain gap is large. However, this has not been fully explored by existing ADA works. In this paper, we propose a Local context-aware ADA framework, named LADA, to address this issue. To select informative target samples, we devise a novel criterion based on the local inconsistency of model predictions. Since the labeling budget is usually small, fine-tuning model on only queried data can be inefficient. We progressively augment labeled target data with the confident neighbors in a class-balanced manner. Experiments validate that the proposed criterion chooses more informative target samples than existing active selection strategies. Furthermore, our full method clearly surpasses recent ADA arts on various benchmarks. Code is available at https://github.com/tsun/LADA.
Tao Sun 0009, Cheng Lu 0006, Haibin Ling
ICCV1
2022 Safe Self-Refinement for Transformer-based Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to leverage a label-rich source domain to solve tasks on a related unlabeled target domain. It is a challenging problem especially when a large domain gap lies between the source and target domains. In this paper we propose a novel solution named SSRT (Safe Self-Refinement for Transformer-based domain adaptation), which brings improvement from two aspects. First, encouraged by the success of vision transformers in various vision tasks, we arm SSRT with a transformer backbone. We find that the combination of vision transformer with simple adversarial adaptation surpasses best reported Convolutional Neural Network (CNN)-based results on the challenging DomainNet benchmark, showing its strong transferable feature representation. Second, to reduce the risk of model collapse and improve the effectiveness of knowledge transfer between domains with large gaps, we propose a Safe Self-Refinement strategy. Specifically, SSRT utilizes predictions of perturbed target domain data to refine the model. Since the model capacity of vision transformer is large and predictions in such challenging tasks can be noisy, a safe training mechanism is designed to adaptively adjust learning configuration. Extensive evaluations are conducted on several widely tested UDA benchmarks and SSRT achieves consistently the best performances, including 85.43% on Office-Home, 88.76% on VisDA-2017 and 45.2% on DomainNet.
Tao Sun 0009, Cheng Lu 0006, Tianshuo Zhang, Haibin Ling
CVPR1
2022 Prior Knowledge Guided Unsupervised Domain Adaptation
Tao Sun 0009, Cheng Lu 0006, Haibin Ling
ECCV (33)1
2022 End-to-End Full Projector Compensation
abstract
Full projector compensation aims to modify a projector input image to compensate for both geometric and photometric disturbance of the projection surface. Traditional methods usually solve the two parts separately and may suffer from suboptimal solutions. In this paper, we propose the first end-to-end differentiable solution, named CompenNeSt++, to solve the two problems jointly. First, we propose a novel geometric correction subnet, named WarpingNet, which is designed with a cascaded coarse-to-fine structure to learn the sampling grid directly from sampling images. Second, we propose a novel photometric compensation subnet, named CompenNeSt, which is designed with a siamese architecture to capture the photometric interactions between the projection surface and the projected images, and to use such information to compensate the geometrically corrected images. By concatenating WarpingNet with CompenNeSt, CompenNeSt++ accomplishes full projector compensation and is end-to-end trainable. Third, to improve practicability, we propose a novel synthetic data-based pre-training strategy to significantly reduce the number of training images and training time. Moreover, we construct the first setup-independent full compensation benchmark to facilitate future studies. In thorough experiments, our method shows clear advantages over prior art with promising compensation quality and meanwhile being practically convenient.
Bingyao Huang, Tao Sun 0009, Haibin Ling
IEEE Trans. Pattern Anal. Mach. Intell.2