EDBT 2026 Demo / reviewers in the wild / expert
Zhaoyu Chen 0001
dblp:119/8788-1
· DBLP profile ↗
48ranked-venue papers
7as first author
48since 2021 · last 2026
0000-0002-7112-2596ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 36 since 2021Artificial intelligence and machine learning · 29 · 6 first-author · 29 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual GroundingabstractMultimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the principle of “Seeing is Believing”, we introduce VBackChecker, a novel reference-free hallucination detection framework that verifies the consistency of MLLM-generated responses with visual inputs, by leveraging a pixel-level Grounding LLM equipped with reasoning and referring segmentation capabilities. This referencefree framework not only effectively handles rich-context scenarios, but also offers interpretability. To facilitate this, an innovative pipeline is accordingly designed for generating instruction-tuning data (R-Instruct), featuring richcontext descriptions, grounding masks, and hard negative samples. We further establish R 2 -HalBench, a new hallucination benchmark for MLLMs, which, unlike previous benchmarks, encompasses real-world, rich-context descriptions from 18 MLLMs with high-quality annotations, spanning diverse object-, attribute-, and relationship-level details. VBackChecker outperforms prior complex frameworks and achieves state-of-the-art performance on R^2 -HalBench, even rivaling GPT-4o’s capabilities in hallucination detection. It also surpasses prior methods in the pixel-level grounding task, achieving over a 10% improvement. Pinxue Guo, Chongruo Wu, Xinyu Zhou 0006, Lingyi Hong, Zhaoyu Chen 0001, Kaixun Jiang, Sen-Ching S. Cheung, Wei Zhang 0016 |
AAAI | 5 |
| 2026 | Delving into the adversarial robustness of semantic segmentation with decision-based black-box attacks
Zhaoyu Chen 0001, Zhengyang Shan, Jingwen Chang, Kaixun Jiang, Dingkang Yang, Yiting Cheng 0001 |
Neural Networks | 1 |
| 2026 | LVOS: A Benchmark for Large-Scale Long-Term Video Object SegmentationabstractVideo object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shelf VOS models, part of the existing VOS benchmarks mainly focuses on short-term videos, where objects remain visible most of the time. However, these benchmarks may not fully capture challenges encountered in practical applications, and the absence of long-term datasets restricts further investigation of VOS in realistic scenarios. Thus, we propose a novel benchmark named LVOS, comprising 720 videos with 296,401 frames and 407,945 high-quality annotations. Videos in LVOS last 1.14 minutes on average. Each video includes various attributes, especially challenges encountered in the wild, such as long-term reappearing and cross-temporal similar objects. Compared to previous benchmarks, our LVOS better reflects VOS models' performance in real scenarios. Based on LVOS, we evaluate 15 existing VOS models under 3 different settings and conduct a comprehensive analysis. On LVOS, these models suffer a large performance drop, highlighting the challenge of achieving precise tracking and segmentation in real-world scenarios. Attribute-based analysis indicates that one of the significant factors contributing to accuracy decline is the increased video length, interacting with complex challenges such as long-term reappearance, cross-temporal confusion, and occlusion, which emphasize LVOS's crucial role. We hope our LVOS can advance development of VOS in real scenes. Lingyi Hong, Zhongying Liu, Chenzhi Tan, Yuang Feng, Xinyu Zhou 0006, Pinxue Guo, Zhaoyu Chen 0001, Shuyong Gao, Wei Zhang 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | ClickVOS: Click Video Object SegmentationabstractVideo Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS. Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Debiased Multimodal Understanding for Human Language SequencesabstractHuman multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module. Zhi Xu 0010, Dingkang Yang, Mingcheng Li, Zhaoyu Chen 0001, Jiawei Chen 0012, Jinjie Wei, Lihua Zhang 0002 |
AAAI | 5 |
| 2025 | Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful ComparatorsabstractDespite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the internal factual knowledge of target LLMs. In addition, hallucinations typically exhibit multifaceted patterns in downstream tasks, limiting the model's holistic performance across tasks. In this paper, we propose a Comparator-driven Decoding-Time (CDT) framework to alleviate the response hallucination. Firstly, we construct hallucinatory and truthful comparators with multi-task fine-tuning samples. In this case, we present an instruction prototype-guided mixture of experts strategy to enhance the ability of the corresponding comparators to capture different hallucination or truthfulness patterns in distinct task instructions. CDT constrains next-token predictions to factuality-robust distributions by contrasting the logit differences between the target LLMs and these comparators. Systematic experiments on multiple downstream tasks show that our framework can significantly improve the model performance and response factuality. Dingkang Yang, Dongling Xiao, Jinjie Wei, Mingcheng Li, Zhaoyu Chen 0001, Ke Li 0015, Lihua Zhang 0002 |
AAAI | 5 |
| 2025 | Pruning for Sparse Diffusion Models Based on Gradient FlowabstractDiffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-shot structure pruning to derive lightweight DMs from pre-trained ones, but this approach often leads to a significant drop in generation quality and may result in the removal of crucial weights. Thus we propose a iterative pruning method based on gradient flow, including the gradient flow pruning process and the gradient flow pruning criterion. We employ a progressive soft pruning strategy to maintain the continuity of the mask matrix and guide it along the gradient flow of the energy function based on the pruning criterion in sparse space, thereby avoiding the sudden information loss typically caused by one-shot pruning. Gradient-flow based criterion prune parameters whose removal increases the gradient norm of loss function and can enable fast convergence for a pruned model in iterative pruning stage. Our extensive experiments on widely used datasets demonstrate that our method achieves superior performance in efficiency and consistency with pre-trained models. Ben Wan, Tianyi Zheng 0001, Zhaoyu Chen 0001, Yuxiao Wang 0004, Jia Wang 0004 |
ICASSP | 3 |
| 2025 | General Compression Framework for Efficient Transformer Object TrackingabstractPrevious works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker. Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001 |
ICCV | 7 |
| 2025 | Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
Kaixun Jiang, Zhaoyu Chen 0001, Bo Li 0115, Weifeng Ge |
ICCV | 3 |
| 2025 | Boosting Adversarial Transferability with Spatial Adversarial AlignmentabstractDeep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods still show limited transferability, partiovovocularly in cross-architecture scenarios, such as from CNN to ViT. To achieve high transferability, we propose a technique termed Spatial Adversarial Alignment (SAA), which employs an alignment loss and leverages a witness model to fine-tune the surrogate model. Specifically, SAA consists of two key parts: spatial-aware alignment and adversarial-aware alignment. First, we minimize the divergences of features between the two models in both global and local regions, facilitating spatial alignment. Second, we introduce a self-adversarial strategy that leverages adversarial examples to impose further constraints, aligning features from an adversarial perspective. Through this alignment, the surrogate model is trained to concentrate on the common features extracted by the witness model. This facilitates adversarial attacks on these shared features, thereby yielding perturbations that exhibit enhanced transferability. Extensive experiments on various architectures on ImageNet show that aligned surrogate models based on SAA can provide higher transferable adversarial examples, especially in cross-architecture attacks. Zhaoyu Chen 0001, Haijing Guo, Kaixun Jiang, Jiyuan Fu, Xinyu Zhou 0006, Dingkang Yang, Hao Tang 0005, Bo Li 0115 |
NeurIPS | 1 |
| 2025 | Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences AlignmentabstractPreference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial alignment involves two inherently conflicting preferences: visual consistency and attack effectiveness, which often lead to unstable optimization and reward hacking (e.g., reducing visual quality to improve attack success). To address this, we propose APA (Adversary Preferences Alignment), a two-stage framework that decouples conflicting preferences and optimizes each with differentiable rewards. In the first stage, APA fine-tunes LoRA to improve visual consistency using rule-based similarity reward. In the second stage, APA updates either the image latent or prompt embedding based on feedback from a substitute classifier, guided by trajectory-level and step-wise rewards. To enhance black-box transferability, we further incorporate a diffusion augmentation strategy. Experiments demonstrate that APA achieves significantly better attack transferability while maintaining high visual consistency, inspiring further research to approach adversarial attacks from an alignment perspective. Kaixun Jiang, Zhaoyu Chen 0001, Haijing Guo, Jiyuan Fu, Pinxue Guo, Hao Tang 0005, Bo Li 0115 |
NeurIPS | 2 |
| 2025 | Dynamic Semantic-Aware Correlation Modeling for UAV TrackingabstractUAV tracking can be widely applied in scenarios such as disaster rescue, environmental monitoring, and logistics transportation. However, existing UAV tracking methods predominantly emphasize speed and lack exploration in semantic awareness, which hinders the search region from extracting accurate localization information from the template. The limitation results in suboptimal performance under typical UAV tracking challenges such as camera motion, fast motion, and low resolution, etc. To address this issue, we propose a dynamic semantic aware correlation modeling tracking framework. The core of our framework is a Dynamic Semantic Relevance Generator, which, in combination with the correlation map from the Transformer, explore semantic relevance. The approach enhances the search region's ability to extract important information from the template, improving accuracy and robustness under the aforementioned challenges.
Additionally, to enhance the tracking speed, we design a pruning method for the proposed framework. Therefore, we present multiple model variants that achieve trade-offs between speed and accuracy, enabling flexible deployment according to the available computational resources. Experimental results validate the effectiveness of our method, achieving competitive performance on multiple UAV tracking datasets. Xinyu Zhou 0006, Tongxin Pan, Lingyi Hong, Pinxue Guo, Haijing Guo, Zhaoyu Chen 0001, Kaixun Jiang |
NeurIPS | 6 |
| 2025 | Enhancing the accuracy of Generative Adversarial Networks with Fokker-Planck Equations
Ben Wan, Tianyi Zheng 0001, Zhaoyu Chen 0001, Jia Wang 0004 |
Neurocomputing | 3 |
| 2025 | Exploring the adversarial robustness of face forgery detection with decision-based black-box attacks
Zhaoyu Chen 0001, Bo Li 0115, Kaixun Jiang, Shuang Wu 0001, Shouhong Ding |
Knowl. Based Syst. | 1 |
| 2025 | VideoPure: Diffusion-Based Adversarial Purification for Video RecognitionabstractRecent work indicates that video recognition models are vulnerable to adversarial examples, posing a serious security risk to downstream applications. However, current research has primarily focused on adversarial attacks, with limited work exploring defense mechanisms. Furthermore, due to the spatial-temporal complexity of videos, existing video defense methods face issues of high cost, overfitting, and limited defense performance. Recently, diffusion-based adversarial purification methods have achieved robust defense performance in the image domain. However, due to the additional temporal dimension in videos, directly applying these diffusion-based adversarial purification methods to the video domain suffers performance and efficiency degradation. To achieve an efficient and effective video adversarial defense method, we propose the first diffusion-based video purification framework to improve video recognition models’ adversarial robustness: VideoPure. Given an adversarial example, we first employ temporal DDIM inversion to transform the input distribution into a temporally consistent and trajectory-defined distribution, covering adversarial noise while preserving more video structure. Then, during DDIM denoising, we leverage intermediate results at each denoising step and conduct guided spatial-temporal optimization, removing adversarial noise while maintaining temporal consistency. Finally, we input the list of optimized intermediate results into the video recognition model for multi-step voting to obtain the predicted class. We investigate the defense performance of our method against state-of-the-art black-box, gray-box, and adaptive attacks on benchmark datasets and models. Compared with other adversarial purification methods, our method overall demonstrates better defense performance against different attacks. Moreover, our method can be applied as a flexible defense plugin for video recognition models. Our code is available at https://github.com/deep-kaixun/VideoPure. Kaixun Jiang, Zhaoyu Chen 0001, Jiyuan Fu, Lingyi Hong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Out of Thin Air: Exploring Data-Free Adversarial Robustness DistillationabstractAdversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impractical to deploy in natural high-security scenes due to these methods rely entirely on original or publicly available data with a similar distribution. In fact, these data are almost always private, specific, and distinctive for scenes that require high robustness. To tackle these issues, we propose a challenging but significant task called Data-Free Adversarial Robustness Distillation (DFARD), which aims to train small, easily deployable, robust models without relying on data. We demonstrate that the challenge lies in the lower upper bound of knowledge transfer information, making it crucial to mining and transferring knowledge more efficiently. Inspired by human education, we design a plug-and-play Interactive Temperature Adjustment (ITA) strategy to improve the efficiency of knowledge transfer and propose an Adaptive Generator Balance (AGB) module to retain more data information. Our method uses adaptive hyperparameters to avoid a large number of parameter tuning, which significantly outperforms the combination of existing techniques. Meanwhile, our method achieves stable and reliable performance on multiple benchmarks. Zhaoyu Chen 0001, Dingkang Yang, Pinxue Guo, Kaixun Jiang, Lizhe Qi |
AAAI | 2 |
| 2024 | Large Vision-Language Models as Emotion Recognizers in Context Awareness
Yuxuan Lei, Dingkang Yang, Zhaoyu Chen 0001, Jiawei Chen 0012, Peng Zhai, Lihua Zhang 0002 |
ACML | 3 |
| 2024 | OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient TuningabstractVisual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility, tracking tasks can be divided into RGB tracking and RGB+X (e.g. RGB+N, and RGB+D) tracking. Despite the different input modalities, the core aspect of tracking is the temporal matching. Based on this common ground, we present a general framework to unify various tracking tasks, termed as One Tracker. One- Tracker first performs a large-scale pre-training on a RGB tracker called Foundation Tracker. This pretraining phase equips the Foundation Tracker with a stable ability to estimate the location of the target object. Then we regard other modality information as prompt and build Prompt Tracker upon Foundation Tracker. Through freezing the Foundation Tracker and only adjusting some additional trainable parameters, Prompt Tracker inhibits the strong localization ability from Foundation Tracker and achieves parameter- efficient finetuning on downstream RGB+X tracking tasks. To evaluate the effectiveness of our general framework OneTracker, which is consisted of Foundation Tracker and Prompt Tracker, we conduct extensive experiments on 6 popular tracking tasks across 11 benchmarks and our One- Tracker outperforms other models and achieves state-of-the-art performance. Lingyi Hong, Shilin Yan, Renrui Zhang, Wanyun Li, Xinyu Zhou 0006, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001 |
CVPR | 10 |
| 2024 | De-Confounded Data-Free Knowledge Distillation for Handling Distribution ShiftsabstractData-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset. Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi |
CVPR | 3 |
| 2024 | Self-cooperation Knowledge Distillation for Novel Class Discovery
Zhaoyu Chen 0001, Dingkang Yang, Yunquan Sun, Lizhe Qi |
ECCV (70) | 2 |
| 2024 | Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002 |
ECCV (58) | 6 |
| 2024 | FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-SpoofingabstractWhile existing face anti-spoofing (FAS) methods have achieved high performance on in-domain datasets, good generalization is crucial for their real-world application. Previous domain generalizable FAS methods have attempted to identify common features of live samples from different domains in the spatial domain, but finding such features is challenging. To address this issue, we propose a solution to the FAS problem in the frequency domain. Specifically, we propose a novel Frequency-domain Augmentation Masked Image Model Framework (FAMIM) for domain generalizable face anti-spoofing. Our approach involves mixing different style information into the input face image in the frequency domain and then reconstructing the original face image. This reconstruction task is designed to make our encoder insensitive to the domain style information of the input face image. Additionally, we design a label fusion module in FAMIM to ensure that the encoder extracts liveness or spoof features in the masked image. Our extensive experiments on four widely used domain generalizable FAS datasets demonstrate that our method achieves state-of-the-art performance. Tianyi Zheng 0001, Qinji Yu, Zhaoyu Chen 0001, Jia Wang 0004 |
ICASSP | 3 |
| 2024 | X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
Pinxue Guo, Wanyun Li, Lingyi Hong, Xinyu Zhou 0006, Zhaoyu Chen 0001, Kaixun Jiang, Wei Zhang 0016 |
ACM Multimedia | 6 |
| 2024 | TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
Xinyu Zhou 0006, Kaixun Jiang, Lingyi Hong, Pinxue Guo, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 6 |
| 2024 | Sampling to Distill: Knowledge Transfer from Open-World DataabstractData-Free Knowledge Distillation (DFKD) is a novel task that aims to train high-performance student models using only the pre-trained teacher network without original training data. Most of the existing DFKD methods rely heavily on additional generation modules to synthesize the substitution data resulting in high computational costs and ignoring the massive amounts of easily accessible, low-cost, unlabeled open-world data. Meanwhile, existing methods ignore the domain shift issue between the substitution data and the original data, resulting in knowledge from teachers not always trustworthy and structured knowledge from data becoming a crucial supplement. To tackle the issue, we propose a novel Open-world Data Sampling Distillation (ODSD) method for the DFKD task without the redundant generation process. First, we try to sample open-world data close to the original data's distribution by an adaptive sampling module and introduce a low-noise representation to alleviate the domain shift issue. Then, we build structured relationships of multiple data examples to exploit data knowledge through the student model itself and the teacher's structured representation. Extensive experiments on CIFAR-10, CIFAR-100, NYUv2, and ImageNet show that our ODSD method achieves state-of-the-art performance with lower FLOPs and parameters. Especially, we improve 1.50%-9.59% accuracy on the ImageNet dataset and avoid training the separate generator for each class. Zhaoyu Chen 0001, Jie Zhang 0107, Dingkang Yang, Zuhao Ge, Yang Liu 0246, Siao Liu, Yunquan Sun, Lizhe Qi |
ACM Multimedia | 2 |
| 2024 | Boosting the transferability of adversarial attacks with global momentum initialization
Zhaoyu Chen 0001, Kaixun Jiang, Dingkang Yang, Lingyi Hong, Pinxue Guo, Haijing Guo |
Expert Syst. Appl. | 2 |
| 2024 | Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded TrainingabstractUnderstanding emotions from diverse contexts has received widespread attention in computer vision communities. The core philosophy of Context-Aware Emotion Recognition (CAER) is to provide valuable semantic cues for recognizing the emotions of target persons by leveraging rich contextual information. Current approaches invariably focus on designing sophisticated structures to extract perceptually critical representations from contexts. Nevertheless, a long-neglected dilemma is that a severe context bias in existing datasets results in an unbalanced distribution of emotional states among different contexts, causing biased visual representation learning. From a causal demystification perspective, the harmful bias is identified as a confounder that misleads existing models to learn spurious correlations based on likelihood estimation, limiting the models' performance. To address the issue, we embrace causal inference to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a customized causal graph. Subsequently, we present a Contextual Causal Intervention Module (CCIM) to de-confound the confounder, which is built upon backdoor adjustment theory to facilitate seeking approximate causal effects during model training. As a plug-and-play component, CCIM can easily integrate with existing approaches and bring significant improvements. Systematic experiments on three datasets demonstrate the effectiveness of our CCIM. Dingkang Yang, Kun Yang 0010, Haopeng Kuang, Zhaoyu Chen 0001, Lihua Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Context De-Confounded Emotion RecognitionabstractContext-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight. Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002 |
CVPR | 2 |
| 2023 | Adversarial Contrastive Distillation with Adaptive DenoisingabstractAdversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not always make correct predictions, interfering with the student’s robust performance. Besides, in the previous ARD methods, the robustness comes entirely from one-to-one imitation, ignoring the relationship between examples. To this end, we propose a novel structured ARD method called Contrastive Relationship DeNoise Distillation (CRDND). We design an adaptive compensation module to model the instability of the teacher. Moreover, we utilize the contrastive relationship to explore implicit robustness knowledge among multiple examples. Experimental results on multiple attack benchmarks show CRDND can transfer robust knowledge efficiently and achieves state-of-the-art performance. Zhaoyu Chen 0001, Dingkang Yang, Yang Liu 0246, Siao Liu, Lizhe Qi |
ICASSP | 2 |
| 2023 | Explicit and Implicit Knowledge Distillation via Unlabeled DataabstractData-free knowledge distillation is a challenging model lightweight task for scenarios in which the original dataset is not available. Previous methods require a lot of extra computational costs to update one or more generators and their naive imitate-learning lead to lower distillation efficiency. Based on these observations, we first propose an efficient unlabeled sample selection method to replace high computational generators and focus on improving the training efficiency of the selected samples. Then, a class-dropping mechanism is designed to suppress the label noise caused by the data domain shifts. Finally, we propose a distillation method that incorporates explicit features and implicit structured relations to improve the effect of distillation. Experimental results show that our method can quickly converge and obtain higher accuracy than other state-of-the-art methods. Zuhao Ge, Zhaoyu Chen 0001, Chuangjia Ma, Yunquan Sun, Lizhe Qi |
ICASSP | 3 |
| 2023 | LVOS: A Benchmark for Long-term Video Object SegmentationabstractExisting video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of VOS on the application in realistic scenarios. So, in this paper, we present a new benchmark dataset named LVOS, which consists of 220 videos with a total duration of 421 minutes. To the best of our knowledge, LVOS is the first densely annotated long-term VOS dataset. The videos in our LVOS last 1.59 minutes on average, which is 20 times longer than videos in existing VOS datasets. Each video includes various attributes, especially challenges deriving from the wild, such as long-term reappearing and cross-temporal similar objeccts. Based on LVOS, we assess existing video object segmentation algorithms and propose a Diverse Dynamic Memory network (DDMemory) that consists of three complementary memory banks to exploit temporal information adequately. The experimental results demonstrate the strength and weaknesses of prior methods, pointing promising directions for further study. Data and code are available at https://lingyihongfd.github.io/lvos.github.io/. Lingyi Hong, Zhongying Liu, Wei Zhang 0016, Pinxue Guo, Zhaoyu Chen 0001 |
ICCV | 6 |
| 2023 | Efficient Decision-based Black-box Patch Attacks on Video RecognitionabstractAlthough Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have not been well investigated. Further, decision-based attacks, where attackers only access the predicted hard labels by querying threat models, have not been well explored on video models either, even if they are practical in real-world video recognition scenes. The absence of such studies leads to a huge gap in the robustness assessment for video models. To bridge this gap, this work first explores decision-based patch attacks on video models. We analyze that the huge parameter space brought by videos and the minimal information returned by decision-based models both greatly increase the attack difficulty and query burden. To achieve a query-efficient attack, we propose a spatial-temporal differential evolution (STDE) framework. First, STDE introduces target videos as patch textures and only adds patches on keyframes that are adaptively selected by temporal difference. Second, STDE takes minimizing the patch area as the optimization objective and adopts spatial-temporal mutation and crossover to search for the global optimum without falling into the local optimum. Experiments show STDE has demonstrated state-of-the-art performance in terms of threat, efficiency and imperceptibility. Hence, STDE has the potential to be a powerful tool for evaluating the robustness of video recognition models. Kaixun Jiang, Zhaoyu Chen 0001, Dingkang Yang, Bo Li 0115, Yan Wang 0068 |
ICCV | 2 |
| 2023 | Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement AugmentationabstractLearning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficiency, suffering from serve performance degradation. In this paper, we first conduct qualitative analysis and illuminate the main causes: (i) high-variance gradient magnitudes and (ii) gradient conflicts existed in various augmentation methods. To alleviate these issues, we propose a general policy gradient optimization framework, named Conflict-aware Gradient Agreement Augmentation (CG2A), and better integrate augmentation combination into visual RL algorithms to address the generalization bias. In particular, CG2A develops a Gradient Agreement Solver to adaptively balance the varying gradient magnitudes, and introduces a Soft Gradient Surgery strategy to alleviate the gradient conflicts. Extensive experiments demonstrate that CG2A significantly improves the generalization performance and sample efficiency of visual RL algorithms. Siao Liu, Zhaoyu Chen 0001, Yang Liu 0246, Dingkang Yang, Zhile Zhao, Ziqing Zhou, Xie Yi, Wei Li 0055, Zhongxue Gan 0001 |
ICCV | 2 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 10 |
| 2023 | GSFormer: Geometric-Spatial Transformer on Point Cloud CompletionabstractPoint cloud completion aims to complete the objects’ shape from incomplete 3D objects. Most works based on encoder-decoder lose the local geometric details of global features when encoding partial points. Besides, the decoder lacks the exploration of the correlation between global and local features. To solve these problems, 1) we propose a Geometric Transformer to learn the global and local geometric details of incomplete point clouds by learning their shape prior geometric information in the encoder, which is beneficial to generate geometric keypoints. The generated geometric keypoints contain the global structure information and local geometric details of the complete point cloud. 2) We propose a Spatial Transformer in the decoder, which can adaptively select neighborhood features to learn the long-distance geometric relationship between upsampling points. Experimental results show that our method achieves better performance on PCN and Shapenet-55/34 datasets. Yijun Long, Zhaoyu Chen 0001, Hong Lu 0001 |
ICME | 2 |
| 2023 | Towards Decision-based Sparse Attacks on Video RecognitionabstractRecent studies indicate that sparse attacks threaten the security of deep learning models, which modify only a small set of pixels in the input based on the l0 norm constraint. While existing research has primarily focused on sparse attacks against image models, there is a notable gap in evaluating the robustness of video recognition models. To bridge this gap, we are the first to study sparse video attacks and propose an attack framework named V-DSA in the most challenging decision-based setting, in which threat models only return the predicted hard label. Specifically, V-DSA comprises two modules: a Cross-Modal Generator (CMG) for query-free transfer attacks on each frame and an Optical flow Grouping Evolution algorithm (OGE) for query-efficient spatial-temporal attacks. CMG passes each frame to generate the transfer video as the starting point of the attack based on the feature similarity between image classification and video recognition models. OGE first initializes populations based on transfer video and then leverages optical flow to establish the temporal connection of the perturbed pixels in each frame, which can reduce the parameter space and break the temporal relationship between frames specifically. Finally, OGE complements the above optical flow modeling by grouping evolution which can realize the coarse-to-fine attack to avoid falling into the local optimum. In addition, OGE makes the perturbation with temporal coherence while balancing the number of perturbed pixels per frame, further increasing the imperceptibility of the attack. Extensive experiments demonstrate that V-DSA achieves state-of-the-art performance in terms of both threat effectiveness and imperceptibility. We hope V-DSA can provide valuable insights into the security of video recognition systems. Kaixun Jiang, Zhaoyu Chen 0001, Xinyu Zhou 0006, Lingyi Hong, Bo Li 0115, Yan Wang 0068 |
ACM Multimedia | 2 |
| 2023 | Exploring the Adversarial Robustness of Video Object Segmentation via One-shot Adversarial AttacksabstractVideo object segmentation (VOS) is a fundamental task for computer vision and multimedia. Despite significant progress of VOS models in recent works, there has been little research on the VOS models' adversarial robustness, posing serious security risks in the VOS models' practical applications (e.g., autonomous driving and video surveillance). Adversarial robustness refers to the ability of the model to resist malicious attacks on adversarial examples. To address this gap, we propose a one-shot adversarial robustness evaluation framework (i.e., the adversary only perturbs the first frame) for VOS models, including white-box and black-box attacks. For white-box attacks, we introduce Objective Attention (OA) and Boundary Attention (BA) mechanisms to enhance the attention of attack on objects from both pixel and object levels while mitigating issues such as multi-objects attack imbalance, attack bias towards the background, and boundary reservation. For black-box attacks, we propose the Video Diverse Input (VDI) module, which utilizes data augmentation to simulate historical information, improving our method's black-box transferability. We conduct extensive experiments to evaluate the adversarial robustness of VOS models with different structures. Our experimental results reveal that existing VOS models are more vulnerable to our attacks (both white-box and black-box) compared to other state-of-the-art attacks. We further analyze the influence of different designs (e.g., memory and matching mechanisms) on adversarial robustness. Finally, we provide insights for designing more secure VOS models in the future. Kaixun Jiang, Lingyi Hong, Zhaoyu Chen 0001, Pinxue Guo, Zeng Tao, Yan Wang 0068 |
ACM Multimedia | 3 |
| 2023 | Freq-HD: An Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for in-the-Wild Facial Expression Recognition in VideosabstractThe in-the-wild dynamic facial expression recognition (DFER) has been challenging due to several high-dynamics factors such as limited dynamic expression-related frames and variable non-expression noise in facial expression sequences. To provide more expression-related clips for DFER models, we propose a novel and interpretable frequency-based method (Freq-HD) for high-dynamics affective clip selection. It can select clips containing pure expression changes from sequences and aid different DFER network structures in recognizing in-the-wild dynamic facial expressions more accurately and efficiently. We first design a novel spatial-temporal frequency analysis (STFA) module to compute the dynamics values of each clip by using sliding windows and spatial-temporal frequency analysis. Moreover, we propose a multi-band complementary selection (MBC) module to amend the inappropriate reaction of the dynamics values of different spatial frequency bands in STFA when expression-irrelevant noise occurs. Specifically, the MBC uses an ingenious mapping method to generate the inhibitory factors to complement and separate the dynamics of expressions and non-expressions in different frequency bands. The Freq-HD can select the most expression-correlated clips and the consisting frames, which could be incorporated into any existing DFER models. We extensively evaluate the Freq-HD on two in-the-wild datasets and four DFER baselines, showing that our method significantly improves the subsequent network performance while using fewer input frames and reducing computation cost. More ablation studies and visualization analysis provide further empirical evidence of the effectiveness of our method. Zeng Tao, Yan Wang 0068, Zhaoyu Chen 0001, Boyang Wang 0003, Shaoqi Yan, Kaixun Jiang, Shuyong Gao |
ACM Multimedia | 3 |
| 2023 | A Capture to Registration Framework for Realistic Image Super-Resolution in the Industry EnvironmentabstractThe acquisition and processing of visual data in industrial environments are of paramount importance. High-resolution (HR) images offer superior clarity and richer textural detail compared to low-resolution (LR) images. On the one hand, owing to the incorporation of richer information, HR images demonstrate substantially enhanced performance compared to LR images in downstream applications, such as anomaly detection. On the other hand, they provide valuable insights to designers and quality inspectors who require a detailed understanding of the images. Currently, the majority of research on super-resolution focuses on natural scenes such as cities and fields, however, the development of datasets for industrial scenes is still in its infancy. To address the image distortion in building realistic LR-HR image pairs in the industry environment, we design a capture to registration framework. It consists of the standard imaging system, physical calibration of the imaging system, as well as the rigid to elastic registration of the LR-HR image pairs. Thus, we build the first realistic industrial sence super-resolution dataset (IndSR), comprises of 50 sets of calibrated images with three scale factors and five typical defects. To benchmark IndSR, we employ quantitative, qualitative, and task-oriented studies to evaluate the representative super-resolution and anomaly detection methods. Besides, we systematically investigate and discuss the performances and results of the existing SISR methods to advance research in the field of super-resolution in industry environment. The IndSR dataset can be available from https://byw4ng.github.io/IndSR/. Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Junxiong Lin, Zeng Tao, Pinxue Guo, Zhaoyu Chen 0001, Kaixun Jiang, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 7 |
| 2023 | Content-based Unrestricted Adversarial AttackabstractUnrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. However, current works usually sacrifice unrestricted degrees and subjectively select some image content to guarantee the photorealism of unrestricted adversarial examples, which limits its attack performance. To ensure the photorealism of adversarial examples and boost attack performance, we propose a novel unrestricted attack framework called Content-based Unrestricted Adversarial Attack. By leveraging a low-dimensional manifold that represents natural images, we map the images onto the manifold and optimize them along its adversarial direction. Therefore, within this framework, we implement Adversarial Content Attack (ACA) based on Stable Diffusion and can generate high transferable unrestricted adversarial examples with various adversarial contents. Extensive experimentation and visualization demonstrate the efficacy of ACA, particularly in surpassing state-of-the-art attacks by an average of 13.3-50.4\% and 16.8-48.0\% in normally trained models and defense methods, respectively. Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Kaixun Jiang, Shouhong Ding |
NeurIPS | 1 |
| 2023 | Memory Network With Pixel-Level Spatio-Temporal Learning for Visual Object TrackingabstractMaking full use of temporal and spatial information is critical to cope with the appearance changes of objects in visual object tracking. However, existing methods in the tracking field, which employ a memory network at frame level to learn this information, bring redundancy and cannot build long-term relationships among historical frames due to the limited memory size. In this paper, we propose a novel memory network, Pixel-level Spatio-Temporal Memory (PSTM), which organizes object features in an efficient way to leverage temporal and spatial context information. Specifically, PSTM is constructed and updated by a memory writer, which includes a pixel-level updating strategy to maintain the temporal consistency and dynamically memorize the noteworthy variations. Furthermore, in order to exploit relationships between the object and search region and precisely estimate the state of the object, we propose a memory reader, Pixel-wise Matching and Refinement module (PMR), and model spatial context without a complex manual-designed mechanism. Comprehensive experiments and comparisons on challenging large-scale benchmarks, including GOT-10k, TrackingNet, LaSOT, OTB2015, VOT2020, and NfS, have demonstrated the effectiveness of our proposed method, which performs favorably against state-of-the-art trackers. Zechu Zhou, Xinyu Zhou 0006, Zhaoyu Chen 0001, Pinxue Guo, Qian-Yu Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Query-Efficient Decision-Based Black-Box Patch AttackabstractDeep neural networks (DNNs) have been showed to be highly vulnerable to imperceptible adversarial perturbations. As a complementary type of adversary, patch attacks that introduce perceptible perturbations to the images have attracted the interest of researchers. Existing patch attacks rely on the architecture of the model or the probabilities of predictions and perform poorly in the decision-based setting, which can still construct a perturbation with the minimal information exposed – the top-1 predicted label. In this work, we first explore the decision-based patch attack. To enhance the attack efficiency, we model the patches using paired key-points and use targeted images as the initialization of patches, and parameter optimizations are all performed on the integer domain. Then, we propose a differential evolutionary algorithm named DevoPatch for query-efficient decision-based patch attacks. Experiments demonstrate that DevoPatch outperforms the state-of-the-art black-box patch attacks in terms of patch area and attack success rate within a given query budget on image classification and face verification. Additionally, we conduct the vulnerability evaluation of ViT and MLP on image classification in the decision-based patch attack setting for the first time. Using DevoPatch, we can evaluate the robustness of models to black-box patch attacks. We believe this method could inspire the design and deployment of robust vision models based on various DNN architectures in the future. Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Shouhong Ding |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating DeepfakesabstractMalicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models, leading them to generate distorted outputs. Despite achieving impressive results, these adversarial watermarks have low image-level and model-level transferability, meaning that they can protect only one facial image from one specific deepfake model. To address these issues, we propose a novel solution that can generate a Cross-Model Universal Adversarial Watermark (CMUA-Watermark), protecting a large number of facial images from multiple deepfake models. Specifically, we begin by proposing a cross-model universal attack pipeline that attacks multiple deepfake models iteratively. Then, we design a two-level perturbation fusion strategy to alleviate the conflict between the adversarial watermarks generated by different facial images and models. Moreover, we address the key problem in cross-model optimization with a heuristic approach to automatically find the suitable attack step sizes for different models, further weakening the model-level conflict. Finally, we introduce a more reasonable and comprehensive evaluation method to fully test the proposed method and compare it with existing ones. Extensive experimental results demonstrate that the proposed CMUA-Watermark can effectively distort the fake facial images generated by multiple deepfake models while achieving a better performance than existing methods. Our code is available at https://github.com/VDIGPKU/CMUA-Watermark. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Jingdong Chen, Weisi Lin, Kai-Kuang Ma |
AAAI | 3 |
| 2022 | Towards Practical Certifiable Patch Defense with Vision TransformerabstractPatch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks. Existing certifiable patch defenses sacrifice the clean accuracy of classifiers and only obtain a low certified accuracy on toy datasets. Furthermore, the clean and certified accuracy of these methods is still significantly lower than the accuracy of normal classification networks, which limits their application in practice. To move towards a practical certifiable patch defense, we introduce Vision Transformer (ViT) into the framework of Derandomized Smoothing (DS). Specifically, we propose a progressive smoothed image modeling task to train Vision Transformer, which can capture the more discriminable local context of an image while preserving the global semantic information. For efficient inference and deployment in the real world, we innovatively reconstruct the global self-attention structure of the original ViT into isolated band unit self-attention. On ImageNet, under 2% area patch attacks our method achieves 41.70% certified accuracy, a nearly 1-fold increase over the previous best method (26.00%). Simultaneously, our method achieves 78.58% clean accuracy, which is quite close to the normal ResNet-101 accuracy. Extensive experiments show that our method obtains state-of-the-art clean and certified accuracy with inferring efficiently on CIFAR-10 and ImageNet. Zhaoyu Chen 0001, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding |
CVPR | 1 |
| 2022 | Shape Matters: Deformable Patch Attack
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Jianghe Xu, Shouhong Ding |
ECCV (4) | 1 |
| 2022 | Efficient Universal Shuffle Attack for Visual Object TrackingabstractRecently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing attacks are difficult to implement in reality due to the real-time of tracking and the re-initialization mechanism. To address these issues, we propose an offline universal adversarial attack called Efficient Universal Shuffle Attack. It takes only one perturbation to cause the tracker malfunction on all videos. To improve the computational efficiency and attack performance, we propose a greedy gradient strategy and a triple loss to efficiently capture and attack model-specific feature representations through the gradients. Experimental results show that EUSA can significantly reduce the performance of state-of-the-art trackers on OTB2015 and VOT2018. Siao Liu, Zhaoyu Chen 0001, Wei Li 0055, Jiwei Zhu, Zhongxue Gan 0001 |
ICASSP | 2 |
| 2022 | DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosabstractCurrent works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset. Yan Wang 0068, Yixuan Sun, Wei Song 0007, Shuyong Gao, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 6 |
| 2021 | Rpattack: Refined Patch Attack on General Object DetectorsabstractNowadays, general object detectors like YOLO and Faster R-CNN as well as their variants are widely exploited in many applications. Many works have revealed that these detectors are extremely vulnerable to adversarial patch attacks. The perturbed regions generated by previous patch-based attack works on object detectors are very large which are not necessary for attacking and perceptible for human eyes. To generate much less but more efficient perturbation, we propose a novel patch-based method for attacking general object detectors. Firstly, we propose a patch selection and refining scheme to find the pixels which have the greatest importance for attack and remove the inconsequential perturbations gradually. Then, for a stable ensemble attack, we balance the gradients of detectors to avoid over-optimizing one of them during the training phase. Our RPAttack can achieve an amazing missed detection rate of 100% for both Yolo v4 and Faster R-CNN while only modifies 0.32% pixels on VOC 2007 test set. Our code is available at https://github.com/VDIGPKU/RPAttack. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Kai-Kuang Ma |
ICME | 3 |