Shuai Jia

dblp:142/5236 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
16since 2021 · last 2026
0009-0000-4900-465XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Parameter-Free Fine-tuning via Redundancy Elimination for Vision Foundation Models
abstract
Vision foundation models (VFMs) have demonstrated remarkable capabilities in learning universal visual representations. However, adapting these models to downstream tasks conventionally requires parameter updates, with even parameter-efficient fine-tuning methods necessitating the modification of thousands to millions of weights. In this paper, we investigate the redundancies in the segment anything model (SAM) and then propose a novel parameter-free fine-tuning method. Unlike traditional fine-tuning methods that adjust parameters, our method emphasizes selecting, reusing, and enhancing pre-trained features, offering a new perspective on fine-tuning foundation models. Specifically, we introduce a channel selection algorithm based on the model's output difference to identify redundant and effective channels. By selectively replacing the redundant channels with more effective ones, we filter out less useful features and reuse more task-irrelevant features to downstream tasks, thereby enhancing the task-specific feature representation. Experiments on both out-of-domain and in-domain datasets demonstrate the efficiency and effectiveness of our method in different vision tasks (e.g., image segmentation, depth estimation and image classification). Notably, our approach can seamlessly integrate with existing fine-tuning strategies (e.g., LoRA, Adapter), further boosting the performance of already fine-tuned models. Moreover, since our channel selection involves only model inference, our method significantly reduces GPU memory overhead.
Jiahuan Long, Tingsong Jiang, Wen Yao 0001, Yizhe Xiong, Zhengqin Xu, Shuai Jia, Chao Ma 0004
AAAI6
2026 Patch-Discontinuity Mining for Generalized Deepfake Detection
abstract
The advancement of generative artificial intelligence has led to the creation of more diverse and realistic fake facial images. This poses serious threats to personal privacy and can contribute to the spread of misinformation. Existing deepfake detection methods usually utilize prior knowledge about forged clues to design complex modules, achieving excellent performance in the intra-domain settings. However, their performance usually suffers from a significant decline in unseen forgery patterns. It is thus desirable to develop a generalized deepfake detection method using a neat network structure. In this paper, we propose a simple yet efficient framework to transfer a powerful large-scale vision model like ViT to the downstream deepfake detection task, namely the generalized deepfake detection framework (GenDF). Concretely, we first propose a deepfake-specific representation learning (DSRL) scheme to learn different discontinuity patterns across patches inside a fake facial image and continuity between patches within a real counterpart in a low-dimensional space. To further alleviate the distribution mismatch between generic real images and human facial images consisting of both real and fake, we introduce a feature space redistribution (FSR) scheme to separately optimize the distributions of real and fake feature space, enabling the model to learn more distinctive representations. Furthermore, to enhance the generalization performance on unseen forgery patterns produced by constantly evolving facial manipulation techniques and diverse variations on real faces, we propose a classification-invariant feature augmentation (CIFAug) function without trainable parameters. CIFAug expands the scopes of real and fake feature space along directions orthogonal to the classification direction, enabling the model to learn more generalizable features while preserving discrimination. Extensive experiments demonstrate that our method achieves state-of-the-art generalization performance in cross-domain and cross-manipulation settings with only 0.28M trainable parameters.
Huanhuan Yuan, Yang Ping, Zhengqin Xu, Junyi Cao, Shuai Jia, Chao Ma 0004
IEEE Trans. Multim.5
2025 Robust SAM: On the Adversarial Robustness of Vision Foundation Models
abstract
The Segment Anything Model (SAM) is a widely used vision foundation model with diverse applications, including image segmentation, detection, and tracking. Given SAM's wide applications, understanding its robustness against adversarial attacks is crucial for real-world deployment. However, research on SAM's robustness is still in its early stages. Existing attacks often overlook the role of prompts in evaluating SAM's robustness, and there has been insufficient exploration of defense methods to balance the robustness and accuracy. To address these gaps, this paper proposes an adversarial robustness framework designed to evaluate and enhance the robustness of SAM. Specifically, we introduce a cross-prompt attack method to enhance the attack transferability across different prompt types. Besides attacking, we propose a few-parameter adaptation strategy to defend SAM against various adversarial attacks. To balance robustness and accuracy, we use the singular value decomposition (SVD) to constrain the space of trainable parameters, where only singular values are adaptable. Experiments demonstrate that our cross-prompt attack method outperforms previous approaches in terms of attack success rate on both SAM and SAM 2. By adapting only 512 parameters, we achieve at least a 15% improvement in mean intersection over union (mIoU) against various adversarial attacks. Compared to previous defense methods, our approach enhances the robustness of SAM while maximally maintaining its original performance.
Jiahuan Long, Zhengqin Xu, Tingsong Jiang, Wen Yao 0001, Shuai Jia, Chao Ma 0004, Xiaoqian Chen
AAAI5
2025 SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
abstract
Automating code review with Large Language Models (LLMs) shows immense promise, yet practical adoption is hampered by their lack of reliability, context-awareness, and control. To address this, we propose Specification-Grounded Code Review (SGCR), a framework that grounds LLMs in human-authored specifications to produce trustworthy and relevant feedback. SGCR features a novel dual-pathway architecture: an explicit path ensures deterministic compliance with predefined rules derived from these specifications, while an implicit path heuristically discovers and verifies issues beyond those rules. Deployed in a live industrial environment at HiThink Research, SGCR’s suggestions achieved a 42% developer adoption rate—a 90.9% relative improvement over a baseline LLM (22%). Our work demonstrates that specification-grounding is a powerful paradigm for bridging the gap between the generative power of LLMs and the rigorous reliability demands of software engineering.
Bingcheng Mao, Shuai Jia, Yujie Ding, Dongming Han, Bin Cao 0004
ASE3
2025 CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared Detectors
abstract
Adversarial patches are widely used to evaluate the robustness of object detection systems in real-world scenarios. These patches were initially designed to deceive single-modal detectors (e.g., visible or infrared) and have recently been extended to target visible-infrared dual-modal detectors. However, existing dual-modal adversarial patch attacks have limited attack effectiveness across diverse physical scenarios. To address this, we propose CDUPatch, a universal cross-modal patch attack against visible-infrared object detectors across scales, views, and scenarios. Specifically, we observe that color variations lead to different levels of thermal absorption, resulting in temperature differences in infrared imaging. Leveraging this property, we propose an RGB-to-infrared adapter that maps RGB patches to infrared patches, enabling unified optimization of cross-modal patches. By learning an optimal color distribution on the adversarial patch, we can manipulate its thermal response and generate an adversarial infrared texture. Additionally, we introduce a multi-scale clipping strategy and construct a new visible-infrared dataset, MSDrone, which contains aerial vehicle images in varying scales and perspectives. These data augmentation strategies enhance the robustness of our patch in real-world conditions. Experiments on four benchmark datasets (e.g., DroneVehicle, LLVIP, VisDrone, MSDrone) show that our method outperforms existing patch attacks in the digital domain. Extensive physical tests further confirm strong transferability across scales, views, and scenarios. Attack demos are provided in the supplementary materials.
Jiahuan Long, Wen Yao 0001, Tingsong Jiang, Shuai Jia, Junqi Wu 0002, Xiaohu Zheng, Chao Ma 0004
ACM Multimedia5
2025 Robust Deep Object Tracking against Adversarial Attacks
Shuai Jia, Chao Ma 0004, Yibing Song, Xiaokang Yang 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.1
2025 Gradient-based sparse voxel attacks on point cloud object detection
Junqi Wu 0002, Wen Yao 0001, Shuai Jia, Tingsong Jiang, Weien Zhou, Chao Ma 0004, Xiaoqian Chen
Pattern Recognit.3
2024 PapMOT: Exploring Adversarial Patch Attack Against Multiple Object Tracking
Jiahuan Long, Tingsong Jiang, Wen Yao 0001, Shuai Jia, Weien Zhou, Chao Ma 0004, Xiaoqian Chen
ECCV (51)4
2024 Prompt Learning with Quaternion Networks
abstract
Multimodal pre-trained models have shown impressive potential in enhancing performance on downstream tasks. However, existing fusion strategies for modalities primarily rely on explicit interaction structures that fail to capture the diverse aspects and patterns inherent in input data. This yields limited performance in zero-shot contexts, especially when fine-grained classifications and abstract interpretations are required. To address this, we propose an effective approach, namely Prompt Learning with Quaternion Networks (QNet), for semantic alignment across diverse modalities. QNet employs a quaternion hidden space where the mutually orthogonal imaginary axes capture rich intermodal semantic spatial correlations from various perspectives. Hierarchical features across multilayers are utilized to encode intricate interdependencies within various modalities with reduced parameters. Our experiments on 11 datasets demonstrate that QNet outperforms state-of-the-art prompt learning techniques in base-to-novel generalization, cross-dataset transfer, and domain transfer scenarios with fewer learnable parameters. The source code is available at https://github.com/VISION-SJTU/QNet.
Boya Shi, Zhengqin Xu, Shuai Jia, Chao Ma 0004
ICLR3
2022 Exploring Frequency Adversarial Attacks for Face Forgery Detection
abstract
Various facial manipulation techniques have drawn seri-ous public concerns in morality, security, and privacy. Al- though existing face forgery classifiers achieve promising performance on detecting fake images, these methods are vulnerable to adversarial examples with injected impercep- tible perturbations on the pixels. Meanwhile, many face forgery detectors always utilize the frequency diversity be-tween real and fake faces as a crucial clue. In this paper, in- stead of injecting adversarial perturbations into the spatial domain, we propose a frequency adversarial attack method against face forgery detectors. Concretely, we apply dis-crete cosine transform (DCT) on the input images and in-troduce a fusion module to capture the salient region of ad-versary in the frequency domain. Compared with existing adversarial attacks (e.g. FGSM, PGD) in the spatial do-main, our method is more imperceptible to human observers and does not degrade the visual quality of the original images. Moreover, inspired by the idea of meta-learning, we also propose a hybrid adversarial attack that performs at-tacks in both the spatial and frequency domains. Exten-sive experiments indicate that the proposed method fools not only the spatial-based detectors but also the state-of- the-art frequency-based detectors effectively. In addition, the proposed frequency attack enhances the transferability across face forgery detectors as black-box attacks.
Shuai Jia, Chao Ma 0004, Taiping Yao, Bangjie Yin, Shouhong Ding, Xiaokang Yang 0001
CVPR1
2022 Adv-Attribute: Inconspicuous and Transferable Adversarial Attack on Face Recognition
abstract
Deep learning models have shown their vulnerability when dealing with adversarial attacks. Existing attacks almost perform on low-level instances, such as pixels and super-pixels, and rarely exploit semantic clues. For face recognition attacks, existing methods typically generate the l_p-norm perturbations on pixels, however, resulting in low attack transferability and high vulnerability to denoising defense models. In this work, instead of performing perturbations on the low-level pixels, we propose to generate attacks through perturbing on the high-level semantics to improve attack transferability. Specifically, a unified flexible framework, Adversarial Attributes (Adv-Attribute), is designed to generate inconspicuous and transferable attacks on face recognition, which crafts the adversarial noise and adds it into different attributes based on the guidance of the difference in face recognition features from the target. Moreover, the importance-aware attribute selection and the multi-objective optimization strategy are introduced to further ensure the balance of stealthiness and attacking strength. Extensive experiments on the FFHQ and CelebA-HQ datasets show that the proposed Adv-Attribute method achieves the state-of-the-art attacking success rates while maintaining better visual effects against recent attack methods.
Shuai Jia, Bangjie Yin, Taiping Yao, Shouhong Ding, Chunhua Shen, Xiaokang Yang 0001, Chao Ma 0004
NeurIPS1
2022 Finding complete minimum driver node set with guaranteed control capacity
Shuai Jia, Yugeng Xi 0001, Dewei Li 0001, Haibin Shao
Neurocomputing1
2022 Brain-Inspired Experience Reinforcement Model for Bin Packing in Varying Environments
abstract
Bin-packing problem (BPP) is a typical combinatorial optimization problem whose decision-making process is NP-hard. This article examines BPPs in varying environments, where random number and shape of items are to be packed in different instances. The objective is to find a unified model to derive optimal decision process that maximizes the utilization of bins. To this end, by mimicking the experience-based reasoning process of humans, this article proposes a novel brain-inspired experience reinforcement model, which takes advantage of both biological and engineering systems. By learning experience from similar situations, the model is adaptive, such as the human brain for sophisticated scenarios and varying environments. The proposed model mimics the functional coordination among brain regions by knowledge representation and knowledge extraction modules. The former one corresponds to the part of information processing and experience storage. The latter one includes two parts that can train reasoning strategies and improve the decision performance. The proposed model is applied to instances of random number and shape of items of BPP. The obtained results outperform the state-of-the-art methods for BPPs in varying environments.
Dewei Li 0001, Shuai Jia, Haibin Shao
IEEE Trans. Neural Networks Learn. Syst.3
2021 IoU Attack: Towards Temporally Coherent Black-Box Adversarial Attack for Visual Object Tracking
abstract
Adversarial attack arises due to the vulnerability of deep neural networks to perceive input samples injected with imperceptible perturbations. Recently, adversarial attack has been applied to visual object tracking to evaluate the robustness of deep trackers. Assuming that the model structures of deep trackers are known, a variety of white-box attack approaches to visual tracking have demonstrated promising results. However, the model knowledge about deep trackers is usually unavailable in real applications. In this paper, we propose a decision-based black-box attack method for visual object tracking. In contrast to existing black-box adversarial attack methods that deal with static images for image classification, we propose IoU attack that sequentially generates perturbations based on the predicted IoU scores from both current and historical frames. By decreasing the IoU scores, the proposed attack method degrades the accuracy of temporal coherent bounding boxes (i.e., object motions) accordingly. In addition, we transfer the learned perturbations to the next few frames to initialize temporal motion attack. We validate the proposed IoU attack on state-of-the-art deep trackers (i.e., detection based, correlation filter based, and long-term trackers). Extensive experiments on the benchmark datasets indicate the effectiveness of the proposed IoU attack method. The source code is available at https://github.com/VISION-SJTU/IoUattack.
Shuai Jia, Yibing Song, Chao Ma 0004, Xiaokang Yang 0001
CVPR1
2021 Mix-hops Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has drawn considerable research interest since it can robustly accommodate dynamic circumstances and complex backgrounds. By modeling the human body skeletons as graph structure, graph convolution network (GCN) has achieved great success in this field. However, these methods based on GCN are difficult to capture global features and relations only through a single-layer network. To obtain the relations between distant skeleton joints of human body, it is necessary to stack multiple graph convolution layers. In addition, the topology of the graph needs to be set manually in graph convolution operation and it is shared in all layers and time dimensions. In this paper, a novel mix-hops graph convolutional network (MHGCN) is proposed to recognize human action from skeleton data. The proposed module can fuse local features with global features through a layer of graph convolutional network. Besides, the topological structure of graph in our model changes with the time dimension and it can be individually learned in an end-to-end way through the BP algorithm. The experiments on several benchmark datasets show remarkable performance for human action recognition, demonstrating the effectiveness of our method.
Dewei Li 0001, Shuai Jia
IJCNN3
2021 Contrastive Cycle Consistency Learning for Unsupervised Visual Tracking
Chao Ma 0004, Shuai Jia, Shugong Xu
PRCV (1)3
2020 Robust Tracking Against Adversarial Attacks
Shuai Jia, Chao Ma 0004, Yibing Song, Xiaokang Yang 0001
ECCV (19)1
2020 Reinforcement learning with actor-critic for knowledge graph reasoning
Dewei Li 0001, Yugeng Xi 0001, Shuai Jia
Sci. China Inf. Sci.4
2019 A Deep Learning Approach for Dog Face Verification and Recognition
Guillaume Mougeot, Dewei Li 0001, Shuai Jia
PRICAI (3)3