Bo Li 0115

dblp:50/3402-115 · DBLP profile ↗
← Back
45ranked-venue papers
2as first author
45since 2021 · last 2026
0000-0001-7817-0665ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 29 since 2021Artificial intelligence and machine learning · 28 · 1 first-author · 28 since 2021Security and privacy · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Disentangled self-supervised video camouflaged object detection and salient object detection
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li
Neural Networks3
2026 Bidirectional Beta-Tuned Diffusion Model
abstract
Diffusion models have gained significant attention in the field of generative modeling due to their capability to produce high-quality samples. However, recent studies show that applying a uniform treatment to all distributions during the training of diffusion models is sub-optimal. In this paper, we present a comprehensive theoretical analysis of the forward process in diffusion models. Our findings indicate that distribution variations are not uniform throughout the diffusion process, with the sharpest changes occurring during the initial stages. Moreover, we observe that the initial distribution converges to a Gaussian distribution at an exponential rate, indicating that different initial distributions rapidly become quite similar during the forward diffusion process. Consequently, employing a uniform timestep sampling strategy does not effectively capture these dynamics, potentially leading to sub-optimal training outcomes for diffusion models. To remedy this, we introduce the Bidirectional Beta-Tuned Diffusion Model (BB-TDM). The BB-TDM leverages the Beta distribution to design the timestep sampling distribution and enhance the separation between different initial distributions during the diffusion process. By selecting appropriate parameters, the BB-TDM ensures that the timestep sampling distribution is aligned with the properties of the forward diffusion process and moderates the convergence speed of different initial distributions. Extensive experiments across various benchmark datasets on different diffusion models confirm the efficacy of the proposed BB-TDM.
Tianyi Zheng 0001, Jiayang Zou, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 DepthMaster: Taming Diffusion Models for Monocular Depth Estimation
abstract
Monocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference speed. Recent methods adopt a single-step deterministic paradigm to improve inference efficiency while maintaining comparable performance. However, they overlook the gap between generative and discriminative features, leading to suboptimal results. In this work, we propose DepthMaster, a single-step diffusion model designed to adapt generative features for the discriminative depth estimation task. First, to mitigate overfitting to texture details introduced by generative features, we propose a Feature Alignment module, which incorporates high-quality semantic features to enhance the denoising network's representation capability. Second, to address the lack of fine-grained details in the single-step deterministic framework, we propose a Fourier Enhancement module to adaptively balance low-frequency structure and high-frequency details. We adopt a two-stage training strategy to fully leverage the potential of the two modules. In the first stage, we focus on learning the global scene structure with the Feature Alignment module, while in the second stage, we exploit the Fourier Enhancement module to improve the visual quality. Through these efforts, our model achieves state-of-the-art performance in terms of generalization and detail preservation, outperforming other diffusion-based methods across various datasets. Our project page can be found at https://indu1ge.github.io/DepthMaster_page.
Ziyang Song 0001, Zerong Wang, Bo Li 0115, Hao Zhang 0063, Ruijie Zhu 0002, Li Liu 0067, Peng-Tao Jiang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 SEMat: Semantic Enhanced Natural Image Interactive Matting
abstract
Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen SAM and only train a lightweight matting decoder by end-to-end matting losses, which do not fully exploit the potential of the pre-trained SAM. Thus, we propose SEMat which revamps the network architecture and training objectives. For network architecture, the proposed feature-aligned transformer learns to extract fine-grained edge and transparency features. The proposed matte-aligned decoder aims to segment matting-specific objects and convert coarse masks into high-precision mattes. For training objectives, the proposed regularization and trimap loss aim to retain the prior from the pre-trained model and push the matting logits extracted from the mask decoder to contain trimap-based semantic information. Extensive experiments across seven diverse datasets demonstrate the superior performance of our method, proving its efficacy in interactive natural image matting. Code is available at https://github.com/XiaRho/SEMat.
Ruihao Xia, Peng-Tao Jiang, Hao Zhang 0063, Qianru Sun, Yang Tang 0001, Bo Li 0115, Pan Zhou 0002
IEEE Trans. Circuits Syst. Video Technol.7
2025 Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
Kaixun Jiang, Zhaoyu Chen 0001, Bo Li 0115, Weifeng Ge
ICCV4
2025 High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
abstract
In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets comprising billions of image-text pairs, such as SD V2.1, have revolutionized text-to-image synthesis by delivering exceptional quality, fine detail resolution, and strong contextual awareness, making them an attractive solution for high-resolution image segmentation. To this end, we propose DiffDIS, a diffusion-driven segmentation model that taps into the potential of the pre-trained U-Net within diffusion models, specifically designed for high-resolution, fine-grained object segmentation. By leveraging the robust generalization capabilities and rich, versatile image representation prior of the SD models, coupled with a task-specific stable one-step denoising approach, we significantly reduce the inference time while preserving high-fidelity, detailed generation. Additionally, we introduce an auxiliary edge generation task to not only enhance the preservation of fine details of the object boundaries, but reconcile the probabilistic nature of diffusion with the deterministic demands of segmentation. With these refined strategies in place, DiffDIS serves as a rapid object mask generation model, specifically optimized for generating detailed binary maps at high resolutions, while demonstrating impressive accuracy and swift processing. Experiments on the DIS5K dataset demonstrate the superiority of DiffDIS, achieving state-of-the-art results through a streamlined inference process. The source code will be publicly available at \href{https://github.com/qianyu-dlut/DiffDIS}{DiffDIS}.
Qian Yu 0015, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115, Lihe Zhang, Huchuan Lu
ICLR5
2025 Boosting Adversarial Transferability with Spatial Adversarial Alignment
abstract
Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods still show limited transferability, partiovovocularly in cross-architecture scenarios, such as from CNN to ViT. To achieve high transferability, we propose a technique termed Spatial Adversarial Alignment (SAA), which employs an alignment loss and leverages a witness model to fine-tune the surrogate model. Specifically, SAA consists of two key parts: spatial-aware alignment and adversarial-aware alignment. First, we minimize the divergences of features between the two models in both global and local regions, facilitating spatial alignment. Second, we introduce a self-adversarial strategy that leverages adversarial examples to impose further constraints, aligning features from an adversarial perspective. Through this alignment, the surrogate model is trained to concentrate on the common features extracted by the witness model. This facilitates adversarial attacks on these shared features, thereby yielding perturbations that exhibit enhanced transferability. Extensive experiments on various architectures on ImageNet show that aligned surrogate models based on SAA can provide higher transferable adversarial examples, especially in cross-architecture attacks.
Zhaoyu Chen 0001, Haijing Guo, Kaixun Jiang, Jiyuan Fu, Xinyu Zhou 0006, Dingkang Yang, Hao Tang 0005, Bo Li 0115
NeurIPS8
2025 Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment
abstract
Preference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial alignment involves two inherently conflicting preferences: visual consistency and attack effectiveness, which often lead to unstable optimization and reward hacking (e.g., reducing visual quality to improve attack success). To address this, we propose APA (Adversary Preferences Alignment), a two-stage framework that decouples conflicting preferences and optimizes each with differentiable rewards. In the first stage, APA fine-tunes LoRA to improve visual consistency using rule-based similarity reward. In the second stage, APA updates either the image latent or prompt embedding based on feedback from a substitute classifier, guided by trajectory-level and step-wise rewards. To enhance black-box transferability, we further incorporate a diffusion augmentation strategy. Experiments demonstrate that APA achieves significantly better attack transferability while maintaining high visual consistency, inspiring further research to approach adversarial attacks from an alignment perspective.
Kaixun Jiang, Zhaoyu Chen 0001, Haijing Guo, Jiyuan Fu, Pinxue Guo, Hao Tang 0005, Bo Li 0115
NeurIPS8
2025 Towards Training-Free Open-World Segmentation via Image Prompt Foundation Models
Lv Tang, Peng-Tao Jiang, Haoke Xiao, Bo Li 0115
Int. J. Comput. Vis.4
2025 EnfoMax: Domain entropy and mutual information maximization for domain generalized face anti-spoofing
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
Neurocomputing2
2025 CSBNet: Leveraging Edge Intelligence for Multigranularity Low-Light Image Enhancement
abstract
Low-light (LOL) conditions constantly restrict the performance of Internet of Things (IoT) image sensors, thereby impacting image quality and the precision of visual data analysis. The emerging edge intelligence is crucial for LOL image enhancement in improving image quality and data support reliability for IoT systems, which in turn fosters the intelligence and automation progress of the IoT. The enhancement of LOL images necessitates the restoration of both contextual information and spatial details, maintaining the semantic content of the original image and the point-to-point correspondence between inputs and outputs. However, existing methods predominantly concentrate on one aspect, either contextual information or spatial details, making it difficult to simultaneously balance both. To overcome this challenge, we introduce a novel two-branch network, the context-space balance network (CSBNet), and tailored for LOL image enhancement. It comprises a contextual information recovery network (CIRNet), which adeptly extracts contextual information from multiscale LOL images, and a spatial information recovery network (SIRNet), which is designed to preserve spatial details at the original resolution. We also implement a context-space feature fusion (CSFF) module to seamlessly integrate contextual information with spatial details. Qualitative and quantitative experimental results demonstrate that our CSBNet can better handle various kinds of degradations in lowlight images compared with state-of-the-art solutions on the benchmark LOL dataset. The source code of CSBNet is available athttps://github.com/Loong161/CSBNet.
Yong Wang 0053, Lijun Jiang, Zilong Du, Bo Li 0115, Wenming Yang
IEEE Internet Things J.4
2025 Exploring the adversarial robustness of face forgery detection with decision-based black-box attacks
Zhaoyu Chen 0001, Bo Li 0115, Kaixun Jiang, Shuang Wu 0001, Shouhong Ding
Knowl. Based Syst.2
2024 ASAM: Boosting Segment Anything Model with Adversarial Tuning
abstract
In the evolving landscape of computer vision, foundation models have emerged as pivotal tools, exhibiting ex-ceptional adaptability to a myriad of tasks. Among these, the Segment Anything Model (SAM) by Meta AI has distin-guished itself in image segmentation. However, SAM, like its counterparts, encounters limitations in specific niche ap-plications, prompting a quest for enhancement strategies that do not compromise its inherent capabilities. This pa-per introduces ASAM, a novel methodology that amplifies SAM's performance through adversarial tuning. We har-ness the potential of natural adversarial examples, inspired by their successful implementation in natural language pro-cessing. By utilizing a stable diffusion model, we augment a subset (1%) of the SA-1B dataset, generating adversar-ial instances that are more representative of natural variations rather than conventional imperceptible perturbations. Our approach maintains the photorealism of adversarial ex-amples and ensures alignment with original mask annotations, thereby preserving the integrity of the segmentation task. The fine-tuned ASAM demonstrates significant im-provements across a diverse range of segmentation tasks without necessitating additional data or architectural mod-ifications. The results of our extensive evaluations confirm that ASAM establishes new benchmarks in segmentation tasks, thereby contributing to the advancement of foundational models in computer vision. Our project page is in https://asam2024.github.io/.
Bo Li 0115, Haoke Xiao, Lv Tang
CVPR1
2024 Re-Thinking Data Availability Attacks Against Deep Neural Networks
abstract
The unauthorized use of personal data for commercial purposes and the covert acquisition of private data for training machine learning models continue to raise concerns. To address these issues, researchers have proposed availability attacks that aim to render data unexploitable. However, many availability attack methods can be easily disrupted by adversarial training. Although some robust methods can resist adversarial training, their protective effects are limited. In this paper, we re-examine the existing availability attack methods and propose a novel two-stage min-max-min optimization paradigm to generate robust unlearnable noise. The inner min stage is utilized to generate unlearnable noise, while the outer min-max stage simulates the training process of the poisoned model. Additionally, we formulate the attack effects and use it to constrain the optimization objective. Comprehensive experiments have revealed that the noise generated by our method can lead to a decline in test accuracy for adversarially trained poisoned models by up to approximately 30%, in comparison to SOTA methods.11Code is available at EuterpeK/Rethinking-Data-Availability-Attacks
Bin Fang 0009, Bo Li 0115, Shuang Wu 0001, Shouhong Ding, Ran Yi 0002, Lizhuang Ma
CVPR2
2024 SAFNet: Selective Alignment Fusion Network for Efficient HDR Imaging
Lingtong Kong, Bo Li 0115, Yike Xiong, Hao Zhang 0063, Jinwei Chen 0003
ECCV (26)2
2024 Mono-ViFI: A Unified Learning Framework for Self-supervised Single and Multi-frame Monocular Depth Estimation
Lingtong Kong, Bo Li 0115, Zerong Wang, Jinwei Chen 0003
ECCV (45)3
2024 Beta-Tuned Timestep Diffusion Model
Tianyi Zheng 0001, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
ECCV (3)7
2024 Zero-Shot Co-Salient Object Detection Framework
abstract
Co-salient Object Detection (CoSOD) endeavors to replicate the human visual system’s capacity to recognize common and salient objects within a collection of images. Despite recent advancements in deep learning models, these models still rely on training with well-annotated CoSOD datasets. The exploration of training-free zero-shot CoSOD frameworks has been limited. In this paper, taking inspiration from the zero-shot transfer capabilities of foundational computer vision models, we introduce the first zero-shot CoSOD framework that harnesses these models without any training process. To achieve this, we introduce two novel components in our proposed framework: the group prompt generation (GPG) module and the co-saliency map generation (CMP) module. We evaluate the framework’s performance on widely-used datasets and observe impressive results. Our approach surpasses existing unsupervised methods and even outperforms fully supervised methods developed before 2020, while remaining competitive with some fully supervised methods developed before 2022.
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li
ICASSP3
2024 Improving Adversarial Energy-Based Model via Diffusion Process
abstract
Generative models have shown strong generation ability while efficient likelihood estimation is less explored. Energy-based models (EBMs) define a flexible energy function to parameterize unnormalized densities efficiently but are notorious for being difficult to train. Adversarial EBMs introduce a generator to form a minimax training game to avoid expensive MCMC sampling used in traditional EBMs, but a noticeable gap between adversarial EBMs and other strong generative models still exists. Inspired by diffusion-based models, we embedded EBMs into each denoising step to split a long-generated process into several smaller steps. Besides, we employ a symmetric Jeffrey divergence and introduce a variational posterior distribution for the generator's training to address the main challenges that exist in adversarial EBMs. Our experiments show significant improvement in generation compared to existing adversarial EBMs, while also providing a useful energy function for efficient density estimation.
Cong Geng, Tian Han 0001, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Søren Hauberg, Bo Li 0115
ICML7
2024 Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
abstract
In this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilities of Multimodal Large Language Models (MLLMs). Recognizing the inherent limitations of current COD methodologies, which predominantly rely on supervised learning models demanding extensive and accurately annotated datasets, resulting in weak generalization, our research proposes a zero-shot MMCPF that circumvents these challenges. Although MLLMs hold significant potential for broad applications, their effectiveness in COD is hindered and they would make misinterpretations of camouflaged objects. To address this challenge, we further propose a strategic enhancement called the Chain of Visual Perception (CoVP), which significantly improves the perceptual capabilities of MLLMs in camouflaged scenes by leveraging both linguistic and visual cues more effectively. We validate the effectiveness of MMCPF on five widely used COD datasets, containing CAMO, COD10K, NC4K, MoCA-Mask and OVCamo. Experiments show that MMCPF can outperform all existing state-of-the-art zero-shot COD methods, and achieve competitive performance compared to weakly-supervised and fully-supervised methods, which demonstrates the potential of MMCPF. The Github link of this paper is https://github.com/luckybird1994/MMCPF.
Lv Tang, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115
ACM Multimedia6
2024 Non-uniform Timestep Sampling: Towards Faster Diffusion Model Training
abstract
Diffusion models have garnered significant success in generative tasks, emerging as the predominant model in this domain. Despite their success, the substantial computational resources required for training diffusion models restrict their practical applications. In this paper, we resort to the optimal transport theory to accelerate the training of diffusion models, providing an in-depth analysis of the forward diffusion process. It shows that the upper bound on the Wasserstein distance of the distribution between any two timesteps in the diffusion process is an exponential decrease of the initial distance by a factor of times. This finding suggests that the state distribution of the diffusion model has a non-uniform rate of change at different points in time, thus highlighting the different importance of the diffusion timestep. To this end, we propose a novel non-uniform timestep sampling method based on the Bernoulli distribution, which favors more frequent sampling in significant timestep intervals. The key idea is to make the model focus on timesteps with larger differences, thus accelerating the training of the diffusion model. Experiments on benchmark datasets reveal that the proposed method significantly reduces the computational overhead while improving the quality of the generated images.
Tianyi Zheng 0001, Cong Geng, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115
ACM Multimedia8
2024 Crafter: Facial Feature Crafting against Inversion-based Identity Theft on Deep Models
Liyao Xiang, Hao Zhang 0063, Xinbing Wang, Chenghu Zhou, Bo Li 0115
NDSS7
2024 From Composited to Real-World: Transformer-Based Natural Image Matting
abstract
The task of image matting is an active research area in computer vision, and various trimap-free methods have been proposed to improve its performance. However, these methods do not consider the gap between composited and real-world images, resulting in limited generalization ability. To address this issue, we propose a domain alignment (DA) module that consists of local region-wise alignment (LRA) and global harmonious alignment (GHA). The LRA aligns the most diverse pixels in the transparent regions of the foreground between composited and real images. On the other hand, the GHA aligns the global image harmonization for both composited and real images, which helps the network choose the appropriate semantics for real harmonious images. Additionally, we design a transformer-based network with dynamic attention pruning (DAP) mechanism to accurately locate domain-sensitive regions, allowing the DA module to work more effectively. Furthermore, we introduce a new dataset, the Real-world Matting Dataset (RM-1k), to advance the real-world matting task. Our proposed method is evaluated on two composited benchmarks (Composite-1k and Distinctions-646) and two real-world datasets (AIM-500 and RM-1k), and the results show that our method achieves robust performance on both composited and real-world images.
Lv Tang, Yijie Zhong 0001, Bo Li 0115
IEEE Trans. Circuits Syst. Video Technol.4
2024 MFAE: Masked Frequency Autoencoders for Domain Generalization Face Anti-Spoofing
abstract
The generalizable face anti-spoofing (FAS) has attracted much attention recently. Even though many existing methods perform well under intra-domain settings, the model’s performance in the unseen domain is not satisfying. In this paper, we shift our attention to the frequency domain to seek a solution. Specifically, we examine the characteristics of different frequency band components of FAS images and observe that the model’s cross-domain performance is very sensitive to low-frequency features. To alleviate this sensitivity and improve the model’s performance in FAS cross-domain tasks, we propose a new approach called Masked Frequency Autoencoders (MFAE). MFAE randomly masks a portion of frequencies on the low-frequency spectrum of the image and then reconstructs the image from the resulting embedding. This innovative Masked Image Modeling (MIM) strategy can be used as a self-supervised task for pre-training vision transformers (ViTs), which can reduce the ViT encoder’s sensitivity to domain shifts. Additionally, we add an auxiliary content-regularization decoder in our MFAE to encourage the encoder to be insensitive to low-frequency features. The results show that the model insensitive to low-frequency features performs well on extensive public datasets and outperforms other state-of-the-art methods in cross-domain FAS tasks.
Tianyi Zheng 0001, Bo Li 0115, Shuang Wu 0001, Ben Wan, Guodong Mu, Shice Liu, Shouhong Ding, Jia Wang 0004
IEEE Trans. Inf. Forensics Secur.2
2023 Delving into the Adversarial Robustness of Federated Learning
abstract
In Federated Learning (FL), models are as fragile as centrally trained models against adversarial examples. However, the adversarial robustness of federated learning remains largely unexplored. This paper casts light on the challenge of adversarial robustness of federated learning. To facilitate a better understanding of the adversarial vulnerability of the existing FL methods, we conduct comprehensive robustness evaluations on various attacks and adversarial training methods. Moreover, we reveal the negative impacts induced by directly adopting adversarial training in FL, which seriously hurts the test accuracy, especially in non-IID settings. In this work, we propose a novel algorithm called Decision Boundary based Federated Adversarial Training (DBFAT), which consists of two components (local re-weighting and global regularization) to improve both accuracy and robustness of FL systems. Extensive experiments on multiple datasets demonstrate that DBFAT consistently outperforms other baselines under both IID and non-IID settings.
Jie Zhang 0081, Bo Li 0115, Chen Chen 0043, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001
AAAI2
2023 Attack Can Benefit: An Adversarial Approach to Recognizing Facial Expressions under Noisy Annotations
abstract
The real-world Facial Expression Recognition (FER) datasets usually exhibit complex scenarios with coupled noise annotations and imbalanced classes distribution, which undoubtedly impede the development of FER methods. To address the aforementioned issues, in this paper, we propose a novel and flexible method to spot noisy labels by leveraging adversarial attack, termed as Geometry Aware Adversarial Vulnerability Estimation (GAAVE). Different from existing state-of-the-art methods of noisy label learning (NLL), our method has no reliance on additional information and is thus easy to generalize to the large-scale real-world FER datasets. Besides, the combination of Dataset Splitting module and Subset Refactoring module mitigates the impact of class imbalance, and the Self-Annotator module facilitates the sufficient use of all training data. Extensive experiments on RAF-DB, FERPlus, AffectNet, and CIFAR-10 datasets validate the effectiveness of our method. The stabilized enhancement based on different methods demonstrates the flexibility of our proposed GAAVE.
Jiawen Zheng, Bo Li 0115, Shengchuan Zhang, Shuang Wu 0001, Liujuan Cao, Shouhong Ding
AAAI2
2023 Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition
abstract
Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the imbalance of short- and long-term temporal relationships in DFER. Therefore, we introduce the Multi-3D Dynamic Facial Expression Learning (M3DFEL) framework, which utilizes Multi-Instance Learning (MIL) to handle inexact labels. M3DFEL generates 3D-instances to model the strong short-term temporal relationship and utilizes 3DCNNs for feature extraction. The Dynamic Long-term Instance Aggregation Module (DLIAM) is then utilized to learn the long-term temporal relationships and dynamically aggregate the instances. Our experiments on DFEW and FERV39K datasets show that M3DFEL outperforms existing state-of-the-art approaches with a vanilla R3D18 backbone. The source code is available at https://github.com/faceeyes/M3DFEL.
Hanyang Wang 0001, Bo Li 0115, Shuang Wu 0001, Feng Liu 0039, Shouhong Ding, Aimin Zhou
CVPR2
2023 Efficient Decision-based Black-box Patch Attacks on Video Recognition
abstract
Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have not been well investigated. Further, decision-based attacks, where attackers only access the predicted hard labels by querying threat models, have not been well explored on video models either, even if they are practical in real-world video recognition scenes. The absence of such studies leads to a huge gap in the robustness assessment for video models. To bridge this gap, this work first explores decision-based patch attacks on video models. We analyze that the huge parameter space brought by videos and the minimal information returned by decision-based models both greatly increase the attack difficulty and query burden. To achieve a query-efficient attack, we propose a spatial-temporal differential evolution (STDE) framework. First, STDE introduces target videos as patch textures and only adds patches on keyframes that are adaptively selected by temporal difference. Second, STDE takes minimizing the patch area as the optimization objective and adopts spatial-temporal mutation and crossover to search for the global optimum without falling into the local optimum. Experiments show STDE has demonstrated state-of-the-art performance in terms of threat, efficiency and imperceptibility. Hence, STDE has the potential to be a powerful tool for evaluating the robustness of video recognition models.
Kaixun Jiang, Zhaoyu Chen 0001, Dingkang Yang, Bo Li 0115, Yan Wang 0068
ICCV6
2023 Towards Decision-based Sparse Attacks on Video Recognition
abstract
Recent studies indicate that sparse attacks threaten the security of deep learning models, which modify only a small set of pixels in the input based on the l0 norm constraint. While existing research has primarily focused on sparse attacks against image models, there is a notable gap in evaluating the robustness of video recognition models. To bridge this gap, we are the first to study sparse video attacks and propose an attack framework named V-DSA in the most challenging decision-based setting, in which threat models only return the predicted hard label. Specifically, V-DSA comprises two modules: a Cross-Modal Generator (CMG) for query-free transfer attacks on each frame and an Optical flow Grouping Evolution algorithm (OGE) for query-efficient spatial-temporal attacks. CMG passes each frame to generate the transfer video as the starting point of the attack based on the feature similarity between image classification and video recognition models. OGE first initializes populations based on transfer video and then leverages optical flow to establish the temporal connection of the perturbed pixels in each frame, which can reduce the parameter space and break the temporal relationship between frames specifically. Finally, OGE complements the above optical flow modeling by grouping evolution which can realize the coarse-to-fine attack to avoid falling into the local optimum. In addition, OGE makes the perturbation with temporal coherence while balancing the number of perturbed pixels per frame, further increasing the imperceptibility of the attack. Extensive experiments demonstrate that V-DSA achieves state-of-the-art performance in terms of both threat effectiveness and imperceptibility. We hope V-DSA can provide valuable insights into the security of video recognition systems.
Kaixun Jiang, Zhaoyu Chen 0001, Xinyu Zhou 0006, Lingyi Hong, Bo Li 0115, Yan Wang 0068
ACM Multimedia7
2023 Content-based Unrestricted Adversarial Attack
abstract
Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. However, current works usually sacrifice unrestricted degrees and subjectively select some image content to guarantee the photorealism of unrestricted adversarial examples, which limits its attack performance. To ensure the photorealism of adversarial examples and boost attack performance, we propose a novel unrestricted attack framework called Content-based Unrestricted Adversarial Attack. By leveraging a low-dimensional manifold that represents natural images, we map the images onto the manifold and optimize them along its adversarial direction. Therefore, within this framework, we implement Adversarial Content Attack (ACA) based on Stable Diffusion and can generate high transferable unrestricted adversarial examples with various adversarial contents. Extensive experimentation and visualization demonstrate the efficacy of ACA, particularly in surpassing state-of-the-art attacks by an average of 13.3-50.4\% and 16.8-48.0\% in normally trained models and defense methods, respectively.
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Kaixun Jiang, Shouhong Ding
NeurIPS2
2023 BrightFormer: A transformer to brighten the image
Yong Wang 0053, Bo Li 0115, Xinlin Yuan
Comput. Graph.2
2023 Query-Efficient Decision-Based Black-Box Patch Attack
abstract
Deep neural networks (DNNs) have been showed to be highly vulnerable to imperceptible adversarial perturbations. As a complementary type of adversary, patch attacks that introduce perceptible perturbations to the images have attracted the interest of researchers. Existing patch attacks rely on the architecture of the model or the probabilities of predictions and perform poorly in the decision-based setting, which can still construct a perturbation with the minimal information exposed – the top-1 predicted label. In this work, we first explore the decision-based patch attack. To enhance the attack efficiency, we model the patches using paired key-points and use targeted images as the initialization of patches, and parameter optimizations are all performed on the integer domain. Then, we propose a differential evolutionary algorithm named DevoPatch for query-efficient decision-based patch attacks. Experiments demonstrate that DevoPatch outperforms the state-of-the-art black-box patch attacks in terms of patch area and attack success rate within a given query budget on image classification and face verification. Additionally, we conduct the vulnerability evaluation of ViT and MLP on image classification in the decision-based patch attack setting for the first time. Using DevoPatch, we can evaluate the robustness of models to black-box patch attacks. We believe this method could inspire the design and deployment of robust vision models based on various DNN architectures in the future.
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Shouhong Ding
IEEE Trans. Inf. Forensics Secur.2
2022 Towards Efficient Data Free Blackbox Adversarial Attack
abstract
Classic black-box adversarial attacks can take advantage of transferable adversarial examples generated by a similar substitute model to successfully fool the target model. However, these substitute models need to be trained by target models' training data, which is hard to acquire due to privacy or transmission reasons. Recognizing the limited availability of real data for adversarial queries, recent works proposed to train substitute models in a data-free black-box scenario. However, their generative adversarial networks (GANs) based framework suffers from the convergence failure and the model collapse, resulting in low efficiency. In this paper, by rethinking the collaborative relationship between the generator and the substitute model, we design a novel black-box attack framework. The proposed method can efficiently imitate the target model through a small number of queries and achieve high attack success rate. The comprehensive experiments over six datasets demonstrate the effectiveness of our method against the state-of-the-art attacks. Especially, we conduct both label-only and probability-only attacks on the Microsoft Azure online model, and achieve a 100% attack success rate with only 0.46% query budget of the SOTA method [49].
Jie Zhang 0081, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Lei Zhang 0197, Chao Wu 0001
CVPR2
2022 Towards Practical Certifiable Patch Defense with Vision Transformer
abstract
Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks. Existing certifiable patch defenses sacrifice the clean accuracy of classifiers and only obtain a low certified accuracy on toy datasets. Furthermore, the clean and certified accuracy of these methods is still significantly lower than the accuracy of normal classification networks, which limits their application in practice. To move towards a practical certifiable patch defense, we introduce Vision Transformer (ViT) into the framework of Derandomized Smoothing (DS). Specifically, we propose a progressive smoothed image modeling task to train Vision Transformer, which can capture the more discriminable local context of an image while preserving the global semantic information. For efficient inference and deployment in the real world, we innovatively reconstruct the global self-attention structure of the original ViT into isolated band unit self-attention. On ImageNet, under 2% area patch attacks our method achieves 41.70% certified accuracy, a nearly 1-fold increase over the previous best method (26.00%). Simultaneously, our method achieves 78.58% clean accuracy, which is quite close to the normal ResNet-101 accuracy. Extensive experiments show that our method obtains state-of-the-art clean and certified accuracy with inferring efficiently on CIFAR-10 and ImageNet.
Zhaoyu Chen 0001, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding
CVPR2
2022 Detecting Camouflaged Object in Frequency Domain
abstract
Camouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception ability of human eyes. Hence, we claim that the goal of COD task is not just to mimic the human visual ability in a single RGB domain, but to go beyond the human biological vision. We then introduce the frequency domain as an additional clue to better detect camouflaged objects from backgrounds. To well involve the frequency clues into the CNN models, we present a powerful network with two special components. We first design a novel frequency enhancement module (FEM) to dig clues of camouflaged objects in the frequency domain. It contains the offline discrete cosine transform followed by the learnable enhancement. Then we use a feature alignment to fuse the features from RGB domain and frequency domain. Moreover, to further make full use of the frequency information, we propose the high-order relation module (HOR) to handle the rich fusion feature. Comprehensive experiments on three widely-used COD datasets show the proposed method significantly outperforms other state-of-the-art methods by a large margin.
Yijie Zhong 0001, Bo Li 0115, Lv Tang, Senyun Kuang, Shuang Wu 0001, Shouhong Ding
CVPR2
2022 Shape Matters: Deformable Patch Attack
Zhaoyu Chen 0001, Bo Li 0115, Shuang Wu 0001, Jianghe Xu, Shouhong Ding
ECCV (4)2
2022 Rethinking Two-B-Real Net for Real-Time Salient Object Detection
abstract
Exploring a fast and accurate salient object detection (SOD) model is a promising research area. TBRS [1] has been proposed a two-branch network for real-time SOD. However, its principle of adding an extra path to encode spatial information is time-consuming. And its backbone is borrowed from image classification tasks, may be inefficient for SOD due to the deficiency of task-specific design. To handle these problems, we propose a novel and efficient structure named short-range concatenate module (SRCM) by removing structure redundancy. Specifically, we gradually reduce the dimension of feature maps and use the aggregation of them for image representation, which forms the basic module of SRCM network. Moreover, we propose an efficient detail guidance branch (DBG) to further encode detail structural information in low-level stages instead of the time-consuming perceptual branch used in TBRS. Finally, low-level features and high-level features are fused by the feature projection module (FPM). Extensive evaluations and analysis demonstrate that our proposed algorithm achieves the leading accuracy performance with real-time speed (216fps). We hope that our series of works can motivate future research for real-time SOD task.
Senyun Kuang, Shijin Meng, Lv Tang, Bo Li 0115
ICASSP5
2022 Federated Learning with Label Distribution Skew via Logits Calibration
abstract
Traditional federated optimization methods perform poorly with heterogeneous data (i.e. , accuracy reduction), especially for highly skewed data. In this paper, we investigate the label distribution skew in FL, where the distribution of labels varies across clients. First, we investigate the label distribution skew from a statistical view. We demonstrate both theoretically and empirically that previous methods based on softmax cross-entropy are not suitable, which can result in local models heavily overfitting to minority classes and missing classes. Additionally, we theoretically introduce a deviation bound to measure the deviation of the gradient after local update. At last, we propose FedLC (\textbf{Fed}erated learning via \textbf{L}ogits \textbf{C}alibration), which calibrates the logits before softmax cross-entropy according to the probability of occurrence of each class. FedLC applies a fine-grained calibrated cross-entropy loss to local update by adding a pairwise label margin. Extensive experiments on federated datasets and real-world datasets demonstrate that FedLC leads to a more accurate global model and much improved performance. Furthermore, integrating other FL methods into our approach can further enhance the performance of the global model.
Jie Zhang 0081, Zhiqi Li 0004, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001
ICML3
2022 DENSE: Data-Free One-Shot Federated Learning
abstract
One-shot Federated Learning (FL) has recently emerged as a promising approach, which allows the central server to learn a model in a single communication round. Despite the low communication cost, existing one-shot FL methods are mostly impractical or face inherent limitations, \eg a public dataset is required, clients' models are homogeneous, and additional data/model information need to be uploaded. To overcome these issues, we propose a novel two-stage \textbf{D}ata-fre\textbf{E} o\textbf{N}e-\textbf{S}hot federated l\textbf{E}arning (DENSE) framework, which trains the global model by a data generation stage and a model distillation stage. DENSE is a practical one-shot FL method that can be applied in reality due to the following advantages:(1) DENSE requires no additional information compared with other methods (except the model parameters) to be transferred between clients and the server;(2) DENSE does not require any auxiliary dataset for training;(3) DENSE considers model heterogeneity in FL, \ie different clients can have different model architectures.Experiments on a variety of real-world datasets demonstrate the superiority of our method.For example, DENSE outperforms the best baseline method Fed-ADI by 5.08\% on CIFAR10 dataset.
Jie Zhang 0081, Chen Chen 0043, Bo Li 0115, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chunhua Shen, Chao Wu 0001
NeurIPS3
2022 R2Net: Relight the restored low-light image based on complementarity of illumination and reflection
Yong Wang 0053, Bo Li 0115, Lijun Jiang, Wenming Yang
Signal Process. Image Commun.2
2022 Re-Thinking the Relations in Co-Saliency Detection
abstract
Co-salient object detection (CoSOD) aims to detect common salient objects sharing the same attributes in an image group. The key issue of CoSOD is how to model the inter-saliency relations within an image group. The major limitation of previous methods is that they pre-define the group-to-one relations within an image group. In this paper, we propose a new concept of structural inter-saliency relations and solve the CoSOD with deep reinforcement learning framework. Firstly, we design a semantic relation graph (SRG) to model the structural inter-saliency relations. Then the feature selecting agent (FS-agent) aims to select the informative features, which can help the SRG effectively model structural inter-saliency relations. Finally, relation updating agent (RU-agent) progressively updates the SRG to focus on the co-salient relations like human decision-making process. Extensive experiments on co-saliency datasets show that because of well modeling inter-saliency relations in image group, our proposed method achieves superior performance compared to the state-of-the-art methods. We hope that this paper can motivate future research for visual co-analysis tasks.
Lv Tang, Bo Li 0115, Senyun Kuang, Mofei Song, Shouhong Ding
IEEE Trans. Circuits Syst. Video Technol.2
2022 Toward Stable Co-Saliency Detection and Object Co-Segmentation
abstract
In this paper, we present a novel model for simultaneous stable co-saliency detection (CoSOD) and object co-segmentation (CoSEG). To detect co-saliency (segmentation) accurately, the core problem is to well model inter-image relations between an image group. Some methods design sophisticated modules, such as recurrent neural network (RNN), to address this problem. However, order-sensitive problem is the major drawback of RNN, which heavily affects the stability of proposed CoSOD (CoSEG) model. In this paper, inspired by RNN-based model, we first propose a multi-path stable recurrent unit (MSRU), containing dummy orders mechanisms (DOM) and recurrent unit (RU). Our proposed MSRU not only helps CoSOD (CoSEG) model captures robust inter-image relations, but also reduces order-sensitivity, resulting in a more stable inference and training process. Moreover, we design a cross-order contrastive loss (COCL) that can further address order-sensitive problem by pulling close the feature embedding generated from different input orders. We validate our model on five widely used CoSOD datasets (CoCA, CoSOD3k, Cosal2015, iCoseg and MSRC), and three widely used datasets (Internet, iCoseg and PASCAL-VOC) for object co-segmentation, the performance demonstrates the superiority of the proposed approach as compared to the state-of-the-art (SOTA) methods.
Bo Li 0115, Lv Tang, Senyun Kuang, Mofei Song, Shouhong Ding
IEEE Trans. Image Process.1
2021 Highly Efficient Natural Image Matting
Yijie Zhong 0001, Bo Li 0115, Lv Tang, Hao Tang 0005, Shouhong Ding
BMVC2
2021 Fast: Feature Aggregation for Detecting Salient Object in Real-Time
abstract
This paper introduces a method named FAST for real-time salient object detection with an extremely efficient CNN architecture. Our proposed network starts from a single lightweight backbone and aggregates discriminative features through network-level and phase-level respectively. Based on the multi-scale feature propagation, FAST substantially reduces the number of parameters, but still obtains sufficient receptive field and enhances the model learning ability, which strikes a balance between the speed and performance. To better preserve object boundaries, we also explore the complementary between salient object information and edge information within our lightweight architecture. Extensive evaluations and analysis demonstrate that the proposed algorithm achieves the leading accuracy performance with real-time speed (186fps) which is significantly faster than the existing state-of-the-art methods.
Lv Tang, Bo Li 0115, Yanliang Wu, Shouhong Ding
ICASSP2
2021 Disentangled High Quality Salient Object Detection
abstract
Aiming at discovering and locating most distinctive objects from visual scenes, salient object detection (SOD) plays an essential role in various computer vision systems. Coming to the era of high resolution, SOD methods are facing new challenges. The major limitation of previous methods is that they try to identify the salient regions and estimate the accurate objects boundaries simultaneously with a single regression task at low-resolution. This practice ignores the inherent difference between the two difficult problems, resulting in poor detection quality. In this paper, we propose a novel deep learning framework for high-resolution SOD task, which disentangles the task into a low-resolution saliency classification network (LRSCN) and a high-resolution refinement network (HRRN). As a pixel-wise classification task, LRSCN is designed to capture sufficient semantics at low-resolution to identify the definite salient, background and uncertain image regions. HRRN is a regression task, which aims at accurately refining the saliency value of pixels in the uncertain region to preserve a clear object boundary at high-resolution with limited GPU memory. It is worth noting that by introducing uncertainty into the training process, our HRRN can well address the high-resolution refinement task without using any high-resolution training data. Extensive experiments on high-resolution saliency datasets as well as some widely used saliency benchmarks show that the proposed method achieves superior performance compared to the state-of-the-art methods.
Lv Tang, Bo Li 0115, Yijie Zhong 0001, Shouhong Ding, Mofei Song
ICCV2