EDBT 2026 Demo / reviewers in the wild / expert
Jun Li 0072
dblp:116/1011-72
· DBLP profile ↗
16ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-5009-6536ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bullet-Screen-Emoji Attack With Temporal Difference Noise for Video Action RecognitionabstractRecent studies have shown that video action recognition models are also vulnerable to fooling by adversarial samples. However, currently existing video attack methods usually require high computational overhead (e.g., they generate adversarial perturbations for all frames by default), and most of them are difficult to implement printable attacks in the physical world. To address the above issues, we devise a novel efficient and effective framework for video action recognition attack: Bullet-Screen-Emoji Attack with Temporal Difference Noise (BSE), a reinforcement learning-based black-box attack method that fools the model by simply generating adversarial bullet screens for key frame and scrolling them on clean video. The agent is optimized to make the optimal actions, i.e., searching key frame. Moreover, we introduce a simple and effective temporal difference noise to enhance the attack capability of the adversarial bullet screen and accelerate the convergence speed. Most importantly, BSE enables printable physical attacks. Extensive experiments show that our proposed BSE achieves promising attack performance on mainstream datasets (HMDB51, UCF101 and Kinetics-400) and in the physical world with high efficiency. Yongkang Zhang 0001, Jun Li 0072, Zhi-Ping Shi 0002, Jian Yang 0030, Kaixin Yang, Qiuyan Liang, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Video Motion Blur Attack via Grad-Weighted and Discrete-Fusion Based Perturbation GenerationabstractRecent research has shown that deep learning networks are vulnerable to adversarial samples. Although there has been great progress in the study of adversarial attacks on images, there is relatively little research on adversarial attacks in the video domain, especially on intrinsic factors of videos, such as motion blur. In this paper, we devise a novel Grad-Weighted based One-step Motion Blur Attack (GWO-MBA) and a Discrete-Fusion based Progressive Motion Blur Attack (DFP-MBA) for video recognition, starting from the idea of integrating global adversarial attacks and adversarial patch attacks. Concretely, we use gradient maps to filter and weighted fusion motion blur (termed GWO-MBA) to achieve the attack that matches the motion information in the context of the video. In order to make the generated motion blur attack perturbations more natural and improve the attack success rate, we further introduce a progressive decomposition motion blur strategy (termed DFP-MBA) to progressively fuse more realistic discrete motion blurs. Besides, we propose an Aggressive Motion Blur Generation (AMBG), which generates natural motion blur based on the video context and has a better attack effect. The extensive experiments, on the HMDB-51 and UCF-101 datasets, demonstrate the effectiveness and superiority of our proposed attack method. In addition, the attack effectiveness of the mainstream denoising defense model and the deblur model further validates the robustness of our attack method. Guoming Wu, Jun Li 0072, Yangfan Xu, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Hard-Sample Style Guided Patch Attack With RL-Enhanced Motion Pattern for Video RecognitionabstractAdversarial attacks have been extensively studied in the image field. In recent years, research has shown that video recognition models are also vulnerable to adversarial examples. However, most studies about adversarial attacks for video models have focused on perturbation-based methods, while patch-based black-box attacks have received less attention. Despite the excellent performance of perturbation-based attacks, these attacks are impractical for real-world implementation. Most existing patch-based black-box attacks require occluding larger areas and performing more queries to the target model. In this paper, we propose a hard-sample style guided patch attack with reinforcement learning (RL) enhanced motion patterns for video recognition (HSPA). Specifically, we utilize the style features of video hard samples and transfer their multi-dimensional style features to images to obtain a texture patch set. Then we use reinforcement learning to locate the patch coordinates and obtain a specific adversarial motion pattern of the patch to successfully perform an effective attack on a video recognition model in both the spatial and temporal dimensions. Our experiments on three widely-used video action recognition models (C3D, LRCN, and TDN) and two mainstream datasets (UCF-101 and HMDB-51) demonstrate the superior performance of our method compared to other state-of-the-art approaches. Jian Yang 0030, Jun Li 0072, Yunong Cai, Guoming Wu, Zhi-Ping Shi 0002, Chaodong Tan, Xianglong Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Diffusion Patch Attack With Spatial-Temporal Cross-Evolution for Video RecognitionabstractDeep neural networks (DNNs) have demonstrated excellent performance across various domains. However, recent studies have shown that deep neural networks are vulnerable to adversarial examples, including DNN-based video action recognition models. While much of the existing research on adversarial attacks against video models focuses on perturbation-based attacks, there is limited research on patch-based black-box attacks. Existing patch-based attack algorithms suffer from the problem of a large search space of optimization algorithms and use patches with simple content, leading to suboptimal attack performance or requiring a large number of queries. To address these challenges, we propose the “Diffusion Patch Attack (DPA) with Spatial-Temporal Cross-Evolution (STCE) for Video Recognition,” a novel approach that integrates the excellent properties of the diffusion model into video black-box adversarial attacks for the first time. This integration significantly narrows the parameter search space while enhancing the adversarial content of patches. Moreover, we introduce the spatial-temporal cross-evolutionary algorithm to adapt to the narrowed search space. Specifically, we separate the spatial and temporal parameters and then employ an alternate evolutionary strategy for each parameter type. Extensive experiments conducted on three widely used video action recognition models (C3D, NL, and TPN) and two benchmark datasets (UCF-101 and HMDB-51) demonstrate the superior performance of our approach compared to other state-of-the-art black-box patch attack algorithms. Jian Yang 0030, Zhiyu Guan, Jun Li 0072, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Imperceptible Adversarial Attack With Multigranular Spatiotemporal Attention for Video Action RecognitionabstractIn recent years, the application of video Internet of Things (IoT) in various cities and public places has brought unprecedented opportunities to the security field and achieved great success. However, the latest research shows that video recognition models are also vulnerable to adversarial examples, but adversarial examples based on physical attacks are easily detected by humans, making it difficult to pass human review. To address this problem, in this article, we propose to introduce a novel multigranular spatiotemporal attention network (MSANet), which can attack the video action recognition models imperceptibly. Specifically, to exploit video motion information more effectively and to reduce the detectability of attack perturbations, we design a multiplexed spatiotemporal attention module to select and enhance spatial regions and temporal frames at coarse-grained and fine-grained levels, respectively, thus maintaining a certain degree of smoothness while reducing the perturbation size and avoiding attacking overfitting. In addition, our proposed MSANet achieves imperceptible perturbations to video sequences through alternate iterative optimization combined with the PGD attack mechanism. extended experimental results on two different models (e.g., TDN and TSM) and two widely used data sets [HMDB-51 (Kuehne et al., 2011) and UCF-101 (Soomro et al., 2012)], compared to the state-of-the-art model, demonstrate the effectiveness of our devised video action recognition attack approach. Guoming Wu, Yangfan Xu, Jun Li 0072, Zhi-Ping Shi 0002, Xianglong Liu 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Temporal Transformer Networks With Self-Supervision for Action RecognitionabstractIn recent years, Internet of Things (IoT) has made rapid development, and IoT devices are developing towards intelligence. IoT terminal devices represented by surveillance cameras play an irreplaceable role in modern society, most of them are integrated with video action recognition and other intelligent functions. However, their performance is somewhat affected by the limitation of computing resources of IoT terminal devices and the lack of long-range non-linear temporal relation modeling and reverse motion information modeling. To address this urgent problem, we introduce a startling Temporal Transformer Network with Self-supervision (TTSN). Our high-performance TTSN mainly consists of a temporal transformer module and a temporal sequence self-supervision module. Concisely speaking, we utilize the efficient temporal transformer module to model the non-linear temporal dependencies among non-local frames, which significantly enhances complex motion feature representations. The temporal sequence self-supervision module we employ unprecedentedly adopts the streamlined strategy of “random batch random channel” to reverse the sequence of video frames, allowing robust extractions of motion information representation from inversed temporal dimensions and improving the generalization capability of the model. Extensive experiments on three widely used datasets (HMDB51, UCF101, and Something-something V1) have conclusively demonstrated that our proposed TTSN is promising as it successfully achieves state-of-the-art performance for video action recognition. Our TTSN provides the possibility for its application in IoT scenarios due to its computational complexity and high performance. With the rapid development of the Internet of Things (IoT), more and more data is being disseminated in the form of video, which also puts new requirements on the understanding and modeling of video data. In recent years, 2D Convolutional Networks-based video action recognition has encouragingly gained wide popularity; However, constrained by the lack of long-range non-linear temporal relation modeling and reverse motion information modeling, the performance of existing models is, therefore, undercut seriously. To address this urgent problem, we introduce a startling Temporal Transformer Network with Self-supervision (TTSN). Our high-performance TTSN mainly consists of a temporal transformer module and a temporal sequence self-supervision module. Concisely speaking, we utilize the efficient temporal transformer module to model the non-linear temporal dependencies among non-local frames, which significantly enhances complex motion feature representations. The temporal sequence self-supervision module we employ unprecedentedly adopts the streamlined strategy of “random batch random channel” to reverse the sequence of video frames, allowing robust extractions of motion information representation from inversed temporal dimensions and improving the generalization capability of the model. Extensive experiments on three widely used datasets (HMDB51, UCF101, and Something-something V1) have conclusively demonstrated that our proposed TTSN is promising as it successfully achieves state-of-the-art performance for video action recognition. As a result, our work provides new attention and self-supervised algorithm for processing video data in IoT. Yongkang Zhang 0001, Jun Li 0072, Guoming Wu, Zhi-Ping Shi 0002, Zhaoxun Liu, Zizhang Wu, Xianglong Liu 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Spatio-Temporal Adaptive Network With Bidirectional Temporal Difference for Action RecognitionabstractAction Recognition is a fundamental task in computer vision field, with a wide range of applications in autonomous driving, security monitoring, etc. However, previous action recognition approaches usually suffer from the inappropriate spatio-temporal modeling or high computational consumption (e.g., 3D CNN). In this paper, we propose a novel Spatio-Temporal Adaptive Network (STANet) with bidirectional temporal difference, consisting of a Temporal Adaptive module (TA) and a Spatial Adaptive (SA) module, to sufficiently extract the crucial motion information and model the spatial pivotal appearance information from both forward and backward perspectives, respectively. Specifically, the Temporal Adaptive module uses bidirectional temporal differences to learn valuable motion trends and balance the static semantics and dynamic motion for a certain action during information fusion; while the Spatial Adaptive module uses the bidirectional temporal difference to obtain the spatio-channel attention to stress the discriminative position-relevant and semantic-relevant appearance features. Extensive experiments conducted on widely-used action recognition benchmarks UCF-101, HMDB-51, Something-Something V1, and Kinetics-400 prove the effectiveness of the proposed methods compared to other state-of-the-art approaches. Zhilei Li, Jun Li 0072, Yuqing Ma, Rui Wang 0024, Zhi-Ping Shi 0002, Yifu Ding 0001, Xianglong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Spatio-Temporal Self-Supervision Enhanced Transformer Networks for Action RecognitionabstractWith the development of deep neural networks, video action recognition has gradually become a research hotspot in recent years. However, the additional temporal dimension in video makes this task very challenging. In this paper, we propose a novel Spatio-Temporal Self-Supervision enhanced Transformer Networks (STTNet) for video action recognition, which mainly consists of Self-Supervised SpatioTemporal Representation Learning module and Transformer based Spatio-Temporal Aggregator module. Concretely, our STTNet can adaptively encode the spatial and temporal enhanced key features, which are respectively learned through the Temporal and Spatial Self-Supervised sub-module using the unlabeled video data, in a nonlinear and nonlocal manner via the Transformer based SpatioTemporal Aggrega-tor. The extensive experiments on three widely used datasets (HMDB51, UCF101 and Something-Something V1) demon-strate that our proposed STTNet can achieve the state-of-the-art performance. Code is available at https://github.com/ICME2022/STTNet. Yongkang Zhang 0001, Guoming Wu, Jun Li 0072 |
ICME | 4 |
| 2022 | TMN: Temporal-guided Multiattention Network for Action Recognitionabstract2D convolutional neural network, due to its low computational complexity and fast recognition speed, has attracted more and more attention from researchers in the field of video action recognition. Temporal shift and temporal differential, have made tremendous progress, but the lack of crucial spatiotemporal attention mechanism has led to huge performance loss. To address this issue, we propose a Temporal-guided Multiattention Network (TMN), which fully excavate and fuse spatio-temporal attention information for effective video action recognition. Concretely, the multi-attention module squeezes and expands spatio-temporal features to achieve weighting of corresponding regions for video in spatio-temporal dimensions, while the adaptive temporal guidance module imports temporal guiding signal to the spatial attention and re-weight the global temporal attention to accomplish the accurate temporal modeling. Extensive experiments and analyses show that our proposed temporal-guided multiattention network can achieve state-of-the-art promising video action recognition performance on the widely used benchmarks (HMDB51, UCF101 and Something-Something V1). Yongkang Zhang 0001, Guoming Wu, Yangfan Xu, Zhi-Ping Shi 0002, Jun Li 0072 |
ICPR | 6 |
| 2022 | Two-Branch Attention Network via Efficient Semantic Coupling for One-Shot LearningabstractOver the past few years, Convolutional Neural Networks (CNNs) have achieved remarkable advancement for the tasks of one-shot image classification. However, the lack of effective attention modeling has limited its performance. In this paper, we propose a Two-branch (Content-aware and Position-aware) Attention (CPA) Network via an Efficient Semantic Coupling module for attention modeling. Specifically, we harness content-aware attention to model the characteristic features (e.g., color, shape, texture) as well as position-aware attention to model the spatial position weights. In addition, we exploit support images to improve the learning of attention for the query images. Similarly, we also use query images to enhance the attention model of the support set. Furthermore, we design a local-global optimizing framework that further improves the recognition accuracy. The extensive experiments on four common datasets (miniImageNet, tieredImageNet, CUB-200-2011, CIFAR-FS) with three popular networks (DPGN, RelationNet and IFSL) demonstrate that our devised CPA module equipped with local-global Two-stream framework (CPAT) can achieve state-of-the-art performance, with a significant improvement in accuracy of 3.16% on CUB-200-2011 in particular. Jun Li 0072, Duorui Wang, Xianglong Liu 0001, Zhi-Ping Shi 0002, Meng Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Learning Attention Map for 3D Human Recovery from a Single RGB Image
Jun Li 0072, Zhi-Ping Shi 0002 |
BMVC | 3 |
| 2021 | Multi-Pretext Attention Network For Few-Shot Learning With Self-SupervisionabstractFew-shot learning is an interesting and challenging study, which enables machines to learn from few samples like humans. Existing studies rarely exploit auxiliary information from large amount of unlabeled data. Self-supervised learning is emerged as an efficient method to utilize unlabeled data. Existing self-supervised learning methods always rely on the combination of geometric transformations for the single sample by augmentation, while seriously neglect the endogenous correlation information among different samples that is the same important for the task. In this work, we propose a Graph-driven Clustering (GC), a novel augmentation-free method for self-supervised learning, which does not rely on any auxiliary sample and utilizes the endogenous correlation information among input samples. Besides, we propose Multi-pretext Attention Network (MAN), which exploits a specific attention mechanism to combine the traditional augmentation-relied methods and our GC, adaptively learning their optimized weights to improve the performance and enabling the feature extractor to obtain more universal representations. We evaluate our MAN extensively on miniImageNet and tieredImageNet datasets and the results demonstrate that the proposed method outperforms the state-of-the-art (SOTA) relevant methods.1 Hainan Li, Renshuai Tao, Jun Li 0072, Haotong Qin, Yifu Ding 0001, Shuo Wang 0008, Xianglong Liu 0001 |
ICME | 3 |
| 2020 | Candidate region correlation for video action detection
Yeguang Li, Liang Hu 0001, Jun Li 0072, Deqing Wang 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Spatio-temporal deformable 3D ConvNets with attention for action recognition
Jun Li 0072, Xianglong Liu 0001, Deqing Wang 0001 |
Pattern Recognit. | 1 |
| 2020 | Spatio-Temporal Attention Networks for Action Recognition and DetectionabstractRecently, 3D Convolutional Neural Network (3D CNN) models have been widely studied for video sequences and achieved satisfying performance in action recognition and detection tasks. However, most of the existing 3D CNNs treat all input video frames equally, thus ignoring the spatial and temporal differences across the video frames. To address the problem, we propose a spatio-temporal attention (STA) network that is able to learn the discriminative feature representation for actions, by respectively characterizing the beneficial information at both the frame level and the channel level. By simultaneously exploiting the differences in spatial and temporal dimensions, our STA module enhances the learning capability of the 3D convolutions when handling the complex videos. The proposed STA method can be wrapped as a generic module easily plugged into the state-of-the-art 3D CNN architectures for video action detection and recognition. We extensively evaluate our method on action recognition and detection tasks over three popular datasets (UCF-101, HMDB-51 and THUMOS 2014), and the experimental results demonstrate that adding our STA network module can obtain the state-of-the-art performance on UCF-101 and HMDB-51, which has the top-1 accuracies of 98.4% and 81.4% respectively, and achieve significant improvement on THUMOS 2014 dataset compared against original models. Jun Li 0072, Xianglong Liu 0001, Jingkuan Song, Nicu Sebe |
IEEE Trans. Multim. | 1 |
| 2018 | Progressive Generative Hashing for Image RetrievalabstractRecent years have witnessed the success of the emerging hashing techniques in large-scale image retrieval. Owing to the great learning capacity, deep hashing has become one of the most promising solutions, and achieved attractive performance in practice. However, without semantic label information, the unsupervised deep hashing still remains an open question. In this paper, we propose a novel progressive generative hashing (PGH) framework to help learn a discriminative hashing network in an unsupervised way. Very different from existing studies, it first treats the hash codes as a kind of semantic condition for the similar image generation, and simultaneously feeds the original image and its codes into the generative adversarial networks (GANs). The real images together with the synthetic ones can further help train a discriminative hashing network based on a triplet loss. By iteratively inputting the learnt codes into the hash conditioned GANs, we can progressively enable the hashing network to discover the semantic relations. Extensive experiments on the widely-used image datasets demonstrate that PGH can significantly outperforms state-of-the-art unsupervised hashing methods. Yuqing Ma, Yue He 0001, Jun Li 0072, Xianglong Liu 0001 |
IJCAI | 5 |