Jun Zhang 0067

dblp:29/4190-67 · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0003-1804-9198ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A comprehensive survey of adversarial defense techniques in the visual domain
abstract
In recent years, deep neural networks (DNNs) have achieved widespread success in computer vision tasks, including face recognition, autonomous driving, and medical diagnosis. However, their vulnerability to adversarial attacks has also raised serious concerns regarding system security and reliability. This paper provides a comprehensive overview of recent advances in adversarial defense methods in computer vision. It explores defense strategies for Large Vision-Language Models (LVLMs) while covering traditional adversarial defense methods. The article first explains the basic concepts of adversarial samples, their generation principles, and their actual threats in white-box, black-box and physical worlds. It then systematically combs through the multi-dimensional defense strategies based on model architecture design, dynamic adversarial training, and input-output space purification. Meanwhile, for the security challenges of LVLM in the era of significant models, two types of strategies, input-output space defense and dynamic training defense, are discussed. Finally, this paper summarizes the key issues facing adversarial defense and provides an outlook on future research directions.
Yibin Dong, Jun Lei 0001, Shuohao Li, Jun Zhang 0067
Neurocomputing5
2026 Beyond single-source adversaries: Diversity-driven ensemble adversarial training for object detectors
Yibin Dong, Shuohao Li, Jun Lei 0001, Roel Leus, Jun Zhang 0067
Knowl. Based Syst.6
2026 DATR: Depth-aware transformer for hierarchical and fine-grained scene graph generation
Qinghao Meng, Yuwen Fu, Xuehu Duan, Shuohao Li, Jun Lei 0001, Jun Zhang 0067
Knowl. Based Syst.8
2025 CRAVS-Net: Complex Residual Attention-Based Variable Splitting Network for Accelerated p-MRI
abstract
Magnetic Resonance Imaging (MRI) is widely used in medical diagnosis due to its excellent ability to image soft tissues. However, its scanning process suffers from issues such as long imaging times and slow reconstruction speeds, which limit its clinical efficiency. To address these limitations, we propose a novel accelerated MRI reconstruction method, termed CRAVS-Net. This method combines complex residual modeling with multihead attention mechanisms to effectively integrate amplitude and phase information from MRI images. Additionally, a multistage residual structure and a channel attention fusion module are introduced to enhance the modeling depth of multi-channel features and the utilization of cross-scale contextual information. Experiments on the NYU knee dataset demonstrate that CRAVSNet outperforms existing state-of-the-art approaches in terms of SSIM and PSNR, and exhibits robust reconstruction performance and strong generalization ability.
Yuwen Fu, Xuehu Duan, Qinghao Meng, Jun Zhang 0067
BIBM4
2025 Hierarchical cross-modal interaction network for multimodal fake news detection
Xianghan Wang, Shuohao Li, Jun Zhang 0067
Neurocomputing6
2025 Exploring the Essence of Relationships for Scene Graph Generation via Causal Features Enhancement Network
abstract
Scene graph generation (SGG) establishes a structured representation between multiple objects by exploring their relationship for visual perception and reasoning tasks. Existing SGG methods often fit the relationships' distribution by introducing language prior or statistical knowledge. However, the relationships should be the semantic reflection of the interaction between objects, rather than the statistical dependency between their categories. To solve this problem, we propose a novel Causal Features Enhancement Network (CFEN) to mine the essential semantic features between objects and relationships. Specifically, by decomposing the object features into class-generic and object-specific components, the causal graph framework is designed to analyze these existing SGG methods. To measure the influence of object-specific features for relationship recognition, we construct the counterfactual training framework for computing the difference between fact and counterfactual logits. Besides, to strengthen the role of object-specific features and learn the interaction between objects, a distribution matching loss is proposed to compute the KL divergence between counterfactual outputs and standard difference distributions and modulate the relations predictions. Finally, compared with the current state-of-the-art methods, the extensive experimental results on VG150 and VrR-VG datasets demonstrate the effectiveness and superiority of our proposed CFEN.
Hao Zhou 0029, Tingjin Luo, Jun Zhang 0067, Liguo Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Weighted Learnable Recursive Aggregation Network for Visible Remote Sensing Image Detection
abstract
With the continuous development of intelligent autonomous aerial vehicles (AAV) technology, efficient and accurate sensing of surrounding objects through onboard sensors has become an important research direction. Among these, the object detection is one of the most common perception techniques, but generic object detection methods have low-detection performance in remote sensing images. To address this, we propose the weighted learnable recursive aggregation detection framework, which aims to improve the detection performance in AAV remote sensing images. The network maintains efficiency while ensuring the accuracy of detection. First, to improve the fusion capability of multiscale characterization of small objects, we design a multichannel weights learnable recursive aggregation network. The network improves the multiscale representation fusion capability by dynamically fusing different layers of scale features while recursively aggregating different layers of features. In addition, we design a multichannel recursive residual fusion mechanism for this framework, which is capable of extracting object feature information at different spatial scales and enhances the ability of multiscale characterization of the object. Then, we migrate the multichannel weights learnable recursive aggregation network to different detection frameworks to verify its generalization. Finally, we perform experimental validation using VisDrone 2019, AI-TOD dataset and comparing it with existing methods. The proposed model improves the [email protected] value by 2.1% over the benchmark model on the VisDrone 2019 dataset. On the AI-TOD dataset, our proposed algorithm improves the [email protected] by 1.7% compared with the benchmark model. Also working on the LLVIP dataset, our algorithm achieves 89.8%.
Xuehu Duan, Jun Zhang 0067, Shuohao Li, Jun Lei 0001, Lixin Zhan
IEEE Trans. Geosci. Remote. Sens.3
2024 Dualswin-Ynet: A Novel Bimodal Fusion Network for Ship Detection in Remote Sensing Images
Rusheng Ju, Xiaoyang Liu 0010, Jiyuan Liu 0003, Jun Zhang 0067, Sihang Qiu
ICPR (5)5
2024 FDML: Feature Disentangling and Multi-view Learning for face forgery detection
Hongying Li, Shuohao Li, Jun Zhang 0067
Neurocomputing6
2023 Feature Bias Correction: A Feature Augmentation Method for Long-tailed Recognition
abstract
The features extracted by the network trained on the long-tailed dataset have significant bias, and the existing methods exhibit poor performance. To alleviate this bias, we propose a Feature Bias Correction (FBC) method, which solves this problem by migrating the biased features back to their correct locations. FBC consists of two core components: Feature Saliency Rebalancing and Similarity Feature Distinguishing. Specifically, Feature Saliency Rebalancing encourages features to be more significant by giving less weight to the feature map, which allows the feature map to represent the sample better. The Similarity Feature Distinguishing module guides the model training by giving more accurate labels to the samples. Finally, our method can be easily combined with the existing long-tailed recognition methods. Experiments on multiple datasets show that our FBC achieves state-of-the-art performance on long-tailed recognition tasks.
Jun Zhang 0067, Shuohao Li
ICME3
2023 Locate, Refine and Restore: A Progressive Enhancement Network for Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to segment objects that blend in with their surroundings. Most existing methods mainly tackle this issue by a single-stage framework, which tends to degrade performance in the face of small objects, low-contrast objects and objects with diverse appearances. In this paper, we propose a novel Progressive Enhancement Network (PENet) for COD by imitating the human visual detection system, which follows a three-stage detection process: locate objects, refine textures and restore boundary. Specifically, our PENet contains three key modules, i.e., the object location module (OLM), the group attention module (GAM) and the context feature restoration module (CFRM). The OLM is designed to position the object globally, the GAM is developed to refine both high-level semantic and low-level texture feature representation, and the CFRM is leveraged to effectively aggregate multi-level features for progressively restoring the clear boundary. Extensive results demonstrate that our PENet significantly outperforms 32 state-of-the-art methods on four widely used benchmark datasets
Shuohao Li, Jun Lei 0001, Jun Zhang 0067, Dong Chen 0013
IJCAI5
2023 MSFRNet: Two-stream deep forgery detector via multi-scale feature extraction
abstract
Abstract Face forgery represented by DeepFake technique has raised severe societal concerns. Due to the different scales of tampering traces and the different resolutions of face images, adopting common processing pipelines and standard form of convolutional neural networks (CNNs) will lead to problems such as omission, redundancy, and bias when extracting key discriminative features. To solve the above issues, unlike most existing methods that treat face forensics as a vanilla binary classification task, the authors instead reformulate it as a multi‐scale object detection problem and propose a novel framework called MSFRNet based on multi‐scale feature extraction. Concretely, to alleviate the issues of features omission and redundancy, the authors construct a two‐stream prediction network, where the shallow branch discovers small‐scale objects such as tiny noise by capturing low‐level features with higher resolution and more details, while the deep stream exploits larger receptive fields to detect large‐scale blocky artefacts. Moreover, a multi‐scale feature extraction module is designed to enrich feature representations in each stream. To solve the problem of features bias and ensure that unbiased feature representations are learned, more appropriate data augmentation approaches are proposed by introducing counterfactual causal reasoning. Extensive experiments demonstrate that our framework outperforms most ordinary binary classifiers and achieves positive performance.
Jun Zhang 0067, Shuohao Li, Jun Lei 0001
IET Image Process.2
2023 Camouflaged object detection with counterfactual intervention
Hongying Li, Hao Zhou 0029, Dong Chen 0013, Shuohao Li, Jun Zhang 0067
Neurocomputing7
2023 Debiased Scene Graph Generation for Dual Imbalance Learning
abstract
Scene graph generation (SGG) is one of the hottest topics in computer vision and has attracted many interests since it provides rich semantic information between objects. In practice, the SGG datasets are often dual imbalanced, presented as a large number of backgrounds and rarely few foregrounds, and highly skewed foreground relationships categories (i.e., the long-tailed distribution). How to tackle this dual imbalanced problem is crucial but rarely studied in literature. Existing methods only consider the long-tailed distribution of foregrounds classes and ignore the background-foreground imbalance in SGG, which results in a biased model and prevents it from being applied in the downstream tasks widely. To reduce its side effect and make the contributions of different categories equally, we propose a novel debiased SGG method (named DSDI) by incorporating biased resistance loss and causal intervention tree. We first deeply analyze the potential causes of dual imbalanced problem in SGG. Then, to learn more discriminate representation of the foreground by expanding the foreground features space, the biased resistance loss decouples the background classification from foreground relationship recognition. Meanwhile, a causal graph of content and context is designed to remove the context bias and learn unbiased relationship features via casual intervention tree. Extensive experimental results on two extremely imbalanced datasets: VG150 and VrR-VG, demonstrate our DSDI outperforms other state-of-the-art methods. All our models will be available in https://github.com/zhouhao0515/unbiasedSGG-DSDI.
Hao Zhou 0029, Jun Zhang 0067, Tingjin Luo, Yazhou Yang, Jun Lei 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Patch-DFD: Patch-based end-to-end DeepFake discriminator
Sigang Ju, Jun Zhang 0067, Shuohao Li, Jun Lei 0001
Neurocomputing3
2022 A unified deep sparse graph attention network for scene graph generation
Hao Zhou 0029, Yazhou Yang, Tingjin Luo, Jun Zhang 0067, Shuohao Li
Pattern Recognit.4
2021 Relationship-Aware Primal-Dual Graph Attention Network For Scene Graph Generation
abstract
The relationships and interactions between objects contain rich semantic information, which plays a crucial role in scene understanding. Existing methods do not attach great importance to the expression of relational features. To tackle this problem, we propose a novel Relationship-aware Primal-Dual Graph Attention Network (RPDGAT) to extract the comprehensive semantic features of objects and explore the sparse graph inference for scene graph generation. RPDGAT mines the inherent attributes and the relationships between objects by fusing multiple features, e.g. appearance, spatial, and category features. After feature extraction, we design a trainable relationship distance measure network to construct the robust and sparse graph structure for efficient graphical message passing. Moreover, it can preserve the contextual cues and neighboring dependency for objects and relationships from the interaction between primal and dual graphs. Extensive experimental results present the improved performance of our method over several state-of-the-art methods on the visual genome datasets.
Hao Zhou 0029, Tingjin Luo, Jun Zhang 0067, Jun Lei 0001, Shuohao Li
ICME3
2021 Deep forgery discriminator via image degradation analysis
abstract
Abstract Generative adversarial network‐based deep generative model is widely applied in creating hyper‐realistic face‐swapping images and videos. However, its malicious use has posed a great threat to online contents, thus making detecting the authenticity of images and videos a tricky task. Most of the existing detection methods are only suitable for one type of forgery and only work for low‐quality tampered images, restricting their applications. This paper concerns the construction of a novel discriminator with better comprehensive capabilities. Through analysis of the visual characteristics of manipulated images from the perspective of image quality, it is revealed that the synthesized face does have different degrees of quality degradation compared to the source content. Therefore, several kinds of image quality‐related handicraft features are extracted, including texture, sharpness, frequency domain features, and deep features, to unveil the inconsistent information and modification traces in the fake faces. In this way, a 1065‐dimensional vector of each image is obtained through multi‐feature fusion, and it is then fed into RF to train a targeted binary classification detector. Extensive experiments have shown that the proposed scheme is superior to the previous methods in recognition accuracy on multiple manipulation databases including the Celeb‐DF database with better visual quality.
Jun Zhang 0067, Shuohao Li, Jun Lei 0001, Fenglei Wang, Hao Zhou 0029
IET Image Process.2
2019 Edge gradient feature and long distance dependency for image semantic segmentation
abstract
Image semantic segmentation is a challenging problem for low‐level computer vision. Recently, deep convolutional neural networks (DCNNs) have been proved to achieve outstanding performance in image semantic segmentation. Most current methods still have some problems in segmenting the object edges and the integrity of objects. In this study, the authors first construct the difference‐pooling module in the DCNNs to extract the object edge gradients and get finer boundary in segmentation results. Then the combination of the pyramid pooling module and the atrous spatial pyramid pooling extracts the image global features and the context structure information by building long‐distance dependency between pixels, which is just like a simple fully connected conditional random field (CRF). Different from other methods, the proposed method does not need extra pre‐processing and post‐processing steps, such as extracting gradient features by the traditional algorithm and building context relationships by CRF. Finally, the experimental results on the PASCAL VOC2012 benchmark indicate that the proposed model can obtain the finer boundaries and more complete parts.
Hao Zhou 0029, Anqi Han, Haodong Yang, Jun Zhang 0067
IET Comput. Vis.4
2017 Deep neural network with attention model for scene text recognition
abstract
The authors present a deep neural network (DNN) with attention model for scene text recognition. The proposed model does not require any segmentation of the input text image. The framework is inspired by the attention model presented recently for speech recognition and image captioning. In the proposed framework, feature extraction, feature attention and sequence recognition are integrated in a jointly trainable network. Compared with previous approaches, the following contributions are mainly made. (i) The attention model is applied into DNN to recognise scene text, and it can effectively solve the sequence recognition problem caused by variable length labels. (ii) Rigorous experiments are performed across a number of challenging benchmarks, including IIIT5K, SVT, ICDAR2003 and ICDAR2013 datasets. Results in experiments show that the proposed model is comparable or better than the state‐of‐the‐art methods. (iii) This model only contains 6.5 million parameters. Compared with other DNN models for scene text recognition, this model has the least number of parameters so far.
Shuohao Li, Jun Lei 0001, Jun Zhang 0067
IET Comput. Vis.5
2017 Generating image descriptions with multidirectional 2D long short-term memory
abstract
Connecting visual imagery with descriptive language is a challenge for computer vision and machine translation. To approach this problem, the authors propose a novel end‐to‐end model to generate descriptions for images. Some early works used convolutional neural network‐long‐short‐term memory (CNN‐LSTM) model to describe the image, where a CNN encodes the input image into feature vector and an LSTM decodes the feature vector into a description. Since two‐dimensional LSTM (2DLSTM) has property of translation invariance and can encode the relationships between regions in an image, they not only apply a CNN to extract global features of an image, but also use a multidirectional 2DLSTM to encode the feature maps extracted by CNN into structural local features. Their model is trained through maximising the likelihood of the target description sentence from the training dataset. Experiments on two challenging datasets show the accuracy of the model and the fluency of the language which is learned by their model. They compare bilingual evaluation understudy score and retrieval metric of their results with current state‐of‐the‐art scores and show the improvements on Flickr30k and MS COCO.
Shuohao Li, Jun Zhang 0067, Jun Lei 0001, Dan Tu
IET Comput. Vis.2
2017 Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition
abstract
Text recognition in natural scene remains a challenging problem due to the highly variable appearance in unconstrained condition. The authors develop a system that directly transcribes scene text images to text without character segmentation. They formulate the problem as sequence labelling. They build a convolutional recurrent neural network (RNN) by using deep convolutional neural networks (CNN) for modelling text appearance and RNNs for sequence dynamics. The two models are complementary in modelling capabilities and so integrated together to form the segmentation free system. They train a Gaussian mixture model–hidden Markov model to supervise the training of the CNN model. The system is data driven and needs no hand labelled training data. Their method has several appealing properties: (i) It can recognise arbitrary length text images. (ii) The recognition process does not involve sophisticated character segmentation. (iii) It is trained on scene text images with only word‐level transcriptions. (iv) It can recognise both the lexicon‐based or lexicon‐free text. The proposed system achieves competitive performance comparison with the state of the art on several public scene text datasets, including both lexicon‐based and non‐lexicon ones.
Fenglei Wang, Jun Lei 0001, Jun Zhang 0067
IET Comput. Vis.4
2016 Continuous action segmentation and recognition using hybrid convolutional neural network-hidden Markov model model
abstract
Continuous action recognition in video is more complicated compared with traditional isolated action recognition. Besides the high variability of postures and appearances of each action, the complex temporal dynamics of continuous action makes this problem challenging. In this study, the authors propose a hierarchical framework combining convolutional neural network (CNN) and hidden Markov model (HMM), which recognises and segments continuous actions simultaneously. The authors utilise the CNN's powerful capacity of learning high level features directly from raw data, and use it to extract effective and robust action features. The HMM is used to model the statistical dependences over adjacent sub‐actions and infer the action sequences. In order to combine the advantages of these two models, the hybrid architecture of CNN‐HMM is built. The Gaussian mixture model is replaced by CNN to model the emission distribution of HMM. The CNN‐HMM model is trained using embedded Viterbi algorithm, and the data used to train CNN are labelled by forced alignment. The authors test their method on two public action dataset Weizmann and KTH. Experimental results show that the authors’ method achieves improved recognition and segmentation accuracy compared with several other methods. The superior property of features learnt by CNN is also illustrated.
Jun Lei 0001, Jun Zhang 0067, Dan Tu
IET Comput. Vis.3
2016 Discriminative orthogonal elastic preserving projections for classification
Tingjin Luo, Chenping Hou, Dongyun Yi, Jun Zhang 0067
Neurocomputing4
2014 $L_{1/2}$-Regularized Deconvolution Network for the Representation and Restoration of Optical Remote Sensing Images
abstract
Optical remote sensing images of land cover are composed of many natural and man-made objects and thus exhibit rich image features ranging from low to high levels. Extracting such a wide range of features to represent the optical remote sensing images, beyond edge primitives, is a long-standing goal in the remote sensing and vision research community. The recently proposed deconvolution network (DN) can effectively learn and capture features in a variety of forms: low-level edges, midlevel edge junctions, high-level object parts, and complete objects. The approach is based on the convolutional decomposition of images under an L1sparsity constraint. Unfortunately, the L1regularizer cannot enforce further sparsity, hence limiting the practical efficacy of the DN in optical remote sensing representation and processing. In this paper, we extend the DN by incorporating the L1/2sparsity constraint, which we name the L1/2-DN. The L1/2regularizer not only induces sparsity but is also a better choice among Lq, (01/2-DN algorithm is more efficient, provides a sparser representation, and results in more accurate recovery than the DN. We illustrate the utility of our method on a wide range of optical remote sensing images and compare our results to those yielded by other state-of-the-art methods.
Jun Zhang 0067, Ping Zhong 0001, Yangtai Chen, Shuohao Li
IEEE Trans. Geosci. Remote. Sens.1