Shuohao Li

dblp:143/9657 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey of adversarial defense techniques in the visual domain
abstract
In recent years, deep neural networks (DNNs) have achieved widespread success in computer vision tasks, including face recognition, autonomous driving, and medical diagnosis. However, their vulnerability to adversarial attacks has also raised serious concerns regarding system security and reliability. This paper provides a comprehensive overview of recent advances in adversarial defense methods in computer vision. It explores defense strategies for Large Vision-Language Models (LVLMs) while covering traditional adversarial defense methods. The article first explains the basic concepts of adversarial samples, their generation principles, and their actual threats in white-box, black-box and physical worlds. It then systematically combs through the multi-dimensional defense strategies based on model architecture design, dynamic adversarial training, and input-output space purification. Meanwhile, for the security challenges of LVLM in the era of significant models, two types of strategies, input-output space defense and dynamic training defense, are discussed. Finally, this paper summarizes the key issues facing adversarial defense and provides an outlook on future research directions.
Yibin Dong, Jun Lei 0001, Shuohao Li, Jun Zhang 0067
Neurocomputing4
2026 Beyond single-source adversaries: Diversity-driven ensemble adversarial training for object detectors
Yibin Dong, Shuohao Li, Jun Lei 0001, Roel Leus, Jun Zhang 0067
Knowl. Based Syst.3
2026 DATR: Depth-aware transformer for hierarchical and fine-grained scene graph generation
Qinghao Meng, Yuwen Fu, Xuehu Duan, Shuohao Li, Jun Lei 0001, Jun Zhang 0067
Knowl. Based Syst.6
2025 Scene Graph-Based Semantic Enhancement for Multimodal Fake News Detection
Hongyun Ding, Shuohao Li
ICIC (7)2
2025 Hierarchical cross-modal interaction network for multimodal fake news detection
Xianghan Wang, Shuohao Li, Jun Zhang 0067
Neurocomputing5
2025 Uncertainty-aware disentangled representation learning for multimodal fake news detection
Xianghan Wang, Shuohao Li
Inf. Process. Manag.5
2025 Weighted Learnable Recursive Aggregation Network for Visible Remote Sensing Image Detection
abstract
With the continuous development of intelligent autonomous aerial vehicles (AAV) technology, efficient and accurate sensing of surrounding objects through onboard sensors has become an important research direction. Among these, the object detection is one of the most common perception techniques, but generic object detection methods have low-detection performance in remote sensing images. To address this, we propose the weighted learnable recursive aggregation detection framework, which aims to improve the detection performance in AAV remote sensing images. The network maintains efficiency while ensuring the accuracy of detection. First, to improve the fusion capability of multiscale characterization of small objects, we design a multichannel weights learnable recursive aggregation network. The network improves the multiscale representation fusion capability by dynamically fusing different layers of scale features while recursively aggregating different layers of features. In addition, we design a multichannel recursive residual fusion mechanism for this framework, which is capable of extracting object feature information at different spatial scales and enhances the ability of multiscale characterization of the object. Then, we migrate the multichannel weights learnable recursive aggregation network to different detection frameworks to verify its generalization. Finally, we perform experimental validation using VisDrone 2019, AI-TOD dataset and comparing it with existing methods. The proposed model improves the [email protected] value by 2.1% over the benchmark model on the VisDrone 2019 dataset. On the AI-TOD dataset, our proposed algorithm improves the [email protected] by 1.7% compared with the benchmark model. Also working on the LLVIP dataset, our algorithm achieves 89.8%.
Xuehu Duan, Jun Zhang 0067, Shuohao Li, Jun Lei 0001, Lixin Zhan
IEEE Trans. Geosci. Remote. Sens.4
2024 On the Convergence of an Adaptive Momentum Method for Adversarial Attacks
abstract
Adversarial examples are commonly created by solving a constrained optimization problem, typically using sign-based methods like Fast Gradient Sign Method (FGSM). These attacks can benefit from momentum with a constant parameter, such as Momentum Iterative FGSM (MI-FGSM), to enhance black-box transferability. However, the monotonic time-varying momentum parameter is required to guarantee convergence in theory, creating a theory-practice gap. Additionally, recent work shows that sign-based methods fail to converge to the optimum in several convex settings, exacerbating the issue. To address these concerns, we propose a novel method which incorporates both an innovative adaptive momentum parameter without monotonicity assumptions and an adaptive step-size scheme that replaces the sign operation. Furthermore, we derive a regret upper bound for general convex functions. Experiments on multiple models demonstrate the efficacy of our method in generating adversarial examples with human-imperceptible noise while achieving high attack success rates, indicating its superiority over previous adversarial example generation methods.
Wei Tao 0002, Shuohao Li
AAAI3
2024 Feature Balance Method for Multi-modal Entity Alignment
Shuohao Li
ICPR (8)5
2024 Multi-target Attention Dispersion Adversarial Attack Against Aerial Object Detector
Shujuan Wang, Zhichao Lian, Shuohao Li
ICPR (4)4
2024 RPID: Boosting Transferability of Adversarial Attacks on Vision Transformers
abstract
Vision Transformers (ViTs) have achieved excellent performance on many computer vision tasks, which has attracted attention of many researchers for their adversarial robustness. As a kind of black-box attack, transfer-based at-tacks usually use adversarial examples generated by a surrogate model to attack structurally different models. It is practical and poses a certain threat to the application of ViTs in critical security areas. Existing transfer-based attacks against ViTs suffer from weak adversarial transferability and noticeable perceptibility. In this work, we propose a method called Reduce Regional Perturbation Interaction and Differentiated (RPID) attack, which employs two strategies of reducing correlation between regional perturbations and adding differentiated perturbations to produce adversarial examples. Extensive experiments demonstrate that our proposed method improves the transferability of the baseline methods for adversarial attacks against ViTs while maintaining stealthiness.
Shujuan Wang, Zhichao Lian, Shuohao Li
SMC5
2024 FDML: Feature Disentangling and Multi-view Learning for face forgery detection
Hongying Li, Shuohao Li, Jun Zhang 0067
Neurocomputing5
2023 Feature Bias Correction: A Feature Augmentation Method for Long-tailed Recognition
abstract
The features extracted by the network trained on the long-tailed dataset have significant bias, and the existing methods exhibit poor performance. To alleviate this bias, we propose a Feature Bias Correction (FBC) method, which solves this problem by migrating the biased features back to their correct locations. FBC consists of two core components: Feature Saliency Rebalancing and Similarity Feature Distinguishing. Specifically, Feature Saliency Rebalancing encourages features to be more significant by giving less weight to the feature map, which allows the feature map to represent the sample better. The Similarity Feature Distinguishing module guides the model training by giving more accurate labels to the samples. Finally, our method can be easily combined with the existing long-tailed recognition methods. Experiments on multiple datasets show that our FBC achieves state-of-the-art performance on long-tailed recognition tasks.
Jun Zhang 0067, Shuohao Li
ICME4
2023 Locate, Refine and Restore: A Progressive Enhancement Network for Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to segment objects that blend in with their surroundings. Most existing methods mainly tackle this issue by a single-stage framework, which tends to degrade performance in the face of small objects, low-contrast objects and objects with diverse appearances. In this paper, we propose a novel Progressive Enhancement Network (PENet) for COD by imitating the human visual detection system, which follows a three-stage detection process: locate objects, refine textures and restore boundary. Specifically, our PENet contains three key modules, i.e., the object location module (OLM), the group attention module (GAM) and the context feature restoration module (CFRM). The OLM is designed to position the object globally, the GAM is developed to refine both high-level semantic and low-level texture feature representation, and the CFRM is leveraged to effectively aggregate multi-level features for progressively restoring the clear boundary. Extensive results demonstrate that our PENet significantly outperforms 32 state-of-the-art methods on four widely used benchmark datasets
Shuohao Li, Jun Lei 0001, Jun Zhang 0067, Dong Chen 0013
IJCAI3
2023 MSFRNet: Two-stream deep forgery detector via multi-scale feature extraction
abstract
Abstract Face forgery represented by DeepFake technique has raised severe societal concerns. Due to the different scales of tampering traces and the different resolutions of face images, adopting common processing pipelines and standard form of convolutional neural networks (CNNs) will lead to problems such as omission, redundancy, and bias when extracting key discriminative features. To solve the above issues, unlike most existing methods that treat face forensics as a vanilla binary classification task, the authors instead reformulate it as a multi‐scale object detection problem and propose a novel framework called MSFRNet based on multi‐scale feature extraction. Concretely, to alleviate the issues of features omission and redundancy, the authors construct a two‐stream prediction network, where the shallow branch discovers small‐scale objects such as tiny noise by capturing low‐level features with higher resolution and more details, while the deep stream exploits larger receptive fields to detect large‐scale blocky artefacts. Moreover, a multi‐scale feature extraction module is designed to enrich feature representations in each stream. To solve the problem of features bias and ensure that unbiased feature representations are learned, more appropriate data augmentation approaches are proposed by introducing counterfactual causal reasoning. Extensive experiments demonstrate that our framework outperforms most ordinary binary classifiers and achieves positive performance.
Jun Zhang 0067, Shuohao Li, Jun Lei 0001
IET Image Process.3
2023 Camouflaged object detection with counterfactual intervention
Hongying Li, Hao Zhou 0029, Dong Chen 0013, Shuohao Li, Jun Zhang 0067
Neurocomputing6
2022 Patch-DFD: Patch-based end-to-end DeepFake discriminator
Sigang Ju, Jun Zhang 0067, Shuohao Li, Jun Lei 0001
Neurocomputing4
2022 A unified deep sparse graph attention network for scene graph generation
Hao Zhou 0029, Yazhou Yang, Tingjin Luo, Jun Zhang 0067, Shuohao Li
Pattern Recognit.5
2021 Relationship-Aware Primal-Dual Graph Attention Network For Scene Graph Generation
abstract
The relationships and interactions between objects contain rich semantic information, which plays a crucial role in scene understanding. Existing methods do not attach great importance to the expression of relational features. To tackle this problem, we propose a novel Relationship-aware Primal-Dual Graph Attention Network (RPDGAT) to extract the comprehensive semantic features of objects and explore the sparse graph inference for scene graph generation. RPDGAT mines the inherent attributes and the relationships between objects by fusing multiple features, e.g. appearance, spatial, and category features. After feature extraction, we design a trainable relationship distance measure network to construct the robust and sparse graph structure for efficient graphical message passing. Moreover, it can preserve the contextual cues and neighboring dependency for objects and relationships from the interaction between primal and dual graphs. Extensive experimental results present the improved performance of our method over several state-of-the-art methods on the visual genome datasets.
Hao Zhou 0029, Tingjin Luo, Jun Zhang 0067, Jun Lei 0001, Shuohao Li
ICME5
2021 Deep forgery discriminator via image degradation analysis
abstract
Abstract Generative adversarial network‐based deep generative model is widely applied in creating hyper‐realistic face‐swapping images and videos. However, its malicious use has posed a great threat to online contents, thus making detecting the authenticity of images and videos a tricky task. Most of the existing detection methods are only suitable for one type of forgery and only work for low‐quality tampered images, restricting their applications. This paper concerns the construction of a novel discriminator with better comprehensive capabilities. Through analysis of the visual characteristics of manipulated images from the perspective of image quality, it is revealed that the synthesized face does have different degrees of quality degradation compared to the source content. Therefore, several kinds of image quality‐related handicraft features are extracted, including texture, sharpness, frequency domain features, and deep features, to unveil the inconsistent information and modification traces in the fake faces. In this way, a 1065‐dimensional vector of each image is obtained through multi‐feature fusion, and it is then fed into RF to train a targeted binary classification detector. Extensive experiments have shown that the proposed scheme is superior to the previous methods in recognition accuracy on multiple manipulation databases including the Celeb‐DF database with better visual quality.
Jun Zhang 0067, Shuohao Li, Jun Lei 0001, Fenglei Wang, Hao Zhou 0029
IET Image Process.3
2017 Truncating Wide Networks Using Binary Tree Architectures
abstract
In this paper, we propose a binary tree architecture to truncate architecture of wide networks by reducing the width of the networks. More precisely, in the proposed architecture, the width is incrementally reduced from lower layers to higher layers in order to increase the expressive capacity of networks with a less increase on parameter size. Also, in order to ease the gradient vanishing problem, features obtained at different layers are concatenated to form the output of our architecture. By employing the proposed architecture on a baseline wide network, we can construct and train a new network with same depth but considerably less number of parameters. In our experimental analyses, we observe that the proposed architecture enables us to obtain better parameter size and accuracy trade-off compared to baseline networks using various benchmark image classification datasets. The results show that our model can decrease the classification error of a baseline from 20:43% to 19:22% on Cifar-100 using only 28% of parameters that the baseline has. Code is available at https://github.com/ZhangVision/bitnet.
Yan Zhang 0055, Mete Ozay, Shuohao Li, Takayuki Okatani
ICCV3
2017 Deep neural network with attention model for scene text recognition
abstract
The authors present a deep neural network (DNN) with attention model for scene text recognition. The proposed model does not require any segmentation of the input text image. The framework is inspired by the attention model presented recently for speech recognition and image captioning. In the proposed framework, feature extraction, feature attention and sequence recognition are integrated in a jointly trainable network. Compared with previous approaches, the following contributions are mainly made. (i) The attention model is applied into DNN to recognise scene text, and it can effectively solve the sequence recognition problem caused by variable length labels. (ii) Rigorous experiments are performed across a number of challenging benchmarks, including IIIT5K, SVT, ICDAR2003 and ICDAR2013 datasets. Results in experiments show that the proposed model is comparable or better than the state‐of‐the‐art methods. (iii) This model only contains 6.5 million parameters. Compared with other DNN models for scene text recognition, this model has the least number of parameters so far.
Shuohao Li, Jun Lei 0001, Jun Zhang 0067
IET Comput. Vis.1
2017 Generating image descriptions with multidirectional 2D long short-term memory
abstract
Connecting visual imagery with descriptive language is a challenge for computer vision and machine translation. To approach this problem, the authors propose a novel end‐to‐end model to generate descriptions for images. Some early works used convolutional neural network‐long‐short‐term memory (CNN‐LSTM) model to describe the image, where a CNN encodes the input image into feature vector and an LSTM decodes the feature vector into a description. Since two‐dimensional LSTM (2DLSTM) has property of translation invariance and can encode the relationships between regions in an image, they not only apply a CNN to extract global features of an image, but also use a multidirectional 2DLSTM to encode the feature maps extracted by CNN into structural local features. Their model is trained through maximising the likelihood of the target description sentence from the training dataset. Experiments on two challenging datasets show the accuracy of the model and the fluency of the language which is learned by their model. They compare bilingual evaluation understudy score and retrieval metric of their results with current state‐of‐the‐art scores and show the improvements on Flickr30k and MS COCO.
Shuohao Li, Jun Zhang 0067, Jun Lei 0001, Dan Tu
IET Comput. Vis.1
2014 $L_{1/2}$-Regularized Deconvolution Network for the Representation and Restoration of Optical Remote Sensing Images
abstract
Optical remote sensing images of land cover are composed of many natural and man-made objects and thus exhibit rich image features ranging from low to high levels. Extracting such a wide range of features to represent the optical remote sensing images, beyond edge primitives, is a long-standing goal in the remote sensing and vision research community. The recently proposed deconvolution network (DN) can effectively learn and capture features in a variety of forms: low-level edges, midlevel edge junctions, high-level object parts, and complete objects. The approach is based on the convolutional decomposition of images under an L1sparsity constraint. Unfortunately, the L1regularizer cannot enforce further sparsity, hence limiting the practical efficacy of the DN in optical remote sensing representation and processing. In this paper, we extend the DN by incorporating the L1/2sparsity constraint, which we name the L1/2-DN. The L1/2regularizer not only induces sparsity but is also a better choice among Lq, (01/2-DN algorithm is more efficient, provides a sparser representation, and results in more accurate recovery than the DN. We illustrate the utility of our method on a wide range of optical remote sensing images and compare our results to those yielded by other state-of-the-art methods.
Jun Zhang 0067, Ping Zhong 0001, Yangtai Chen, Shuohao Li
IEEE Trans. Geosci. Remote. Sens.4