EDBT 2026 Demo / reviewers in the wild / expert
Jiafeng Li 0001
dblp:144/1439-1
· DBLP profile ↗
38ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0001-6976-7275ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal driver behavior recognition based on frame-adaptive convolution and feature fusion
Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001 |
Comput. Vis. Image Underst. | 1 |
| 2026 | LRGFormer: A Multiscale Feature Fusion Transformer for Image RestorationabstractAdverse weather conditions can significantly degrade image quality and impair the capture of critical information. Existing restoration networks struggle to effectively combine local, regional, and global features, thereby limiting their ability to handle diverse impacts of such weather. This study proposes the local-region-global transformer (LRGFormer), a transformer-based image restoration model for multiscale feature perception. The model comprises a basic module composed of multi-scale fusion attention (MSFSA) and a channel-spatial dual-attention feed-forward network (CSDF). Specifically, this study designs an MSFSA module. For the first time, it combines rotation-equivariant convolution with local attention for local information extraction and introduces a frequency-domain adaptive attention mechanism. By incorporating a query-aware global adaptive sparse attention mechanism for global information extraction, the network gradually fuses along the channel dimension, enabling progressive capture of spatial and frequency-domain information from the local and regional to global scale. Secondly, a CSDF network structure was designed to enhance channel-spatial interaction and improve the representational capacity of the model. By constructing a basic U-Net framework, the excellent basic modules for image restoration proposed in recent years are compared on a unified framework. Experimental results demonstrated that the proposed basic module can not only better extracts multi-scale features of images and restores image distortion caused by various degradation factors, and also exhibits good universality and generalization. Jiafeng Li 0001, Wanying Hu, Tongyao Jia, Jiaqi Jin, Jing Zhang 0023, Li Zhuo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | BEV-CMHF: A Cross-Modality Hybrid Fusion Framework for BEV 3D Object Detection With Feature Interaction and Temporal FusionabstractAutonomous driving technology has garnered significant attention for its potential to reduce driver burden and enhance road safety. Modern autonomous driving systems rely on a variety of sensors to perceive complex driving environments. Many existing methods map heterogeneous data into the bird’s eye view (BEV) space for feature fusion. However, they often fail to fully exploit the cross-modal interactions between cameras and LiDAR, or incorporate temporal information, resulting in suboptimal performance. Furthermore, commonly used fusion strategies are often overly simplistic. This study proposes BEV-CMHF, a cross-modality hybrid fusion framework for BEV 3D object detection with feature interaction and temporal fusion. By introducing an interactive cross-attention module and a long-short-term temporal module, the proposed framework enhances the representational power of fused BEV features. Specifically, a feature-interaction attention module that facilitates effective interaction between the camera and LiDAR BEV features using deformable attention is designed, providing guidance and supervision for the camera BEV features. Subsequently, a historical feature temporal fusion module that integrates the long-short-term temporal module is introduced to incorporate additional critical temporal information into the BEV features. Moreover, a dynamic hybrid feature-fusion module is designed to fuse the BEV features of the camera and LiDAR effectively through a hybrid attention mechanism that combines coarse and fine attention. Extensive experiments conducted on the nuScenes benchmark validate the effectiveness of the proposed method, achieving 70.87% mAP and 74.00% NDS on the test set. Using a single NVIDIA GeForce RTX 4090, the method attained an inference speed of 5.79 images per second (5.79 img/s), corresponding to an inference time of 172.64 ms on the nuScenes dataset. The source code will be released athttps://github.com/BJUTsipl/BEV-CMHF Jiafeng Li 0001, Jinquan Xu, Mengxun Zhi, Jing Zhang 0023, Li Zhuo 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2026 | HDMDN: Hierarchical-Decoupling Based Meta-Knowledge Single Image Dehazing NetworkabstractImages suffer from color shift and detail distortion owing to limitations of contrast and visibility in hazy scenes, affecting their subjective perception. However, the performance of existing algorithms on real-world hazy images remains limited as scenes can be complex and haze degradation varies in outdoor visual systems. This study proposes a meta-knowledge single image dehazing algorithm based on hierarchical decoupling, combining the advantages of convolutional neural networks (CNNs) and transformers. We propose a novel dual-branch decoupling network that decouples low-level features from high level semantic information in images, leveraging the hierarchical properties of the network. It combines a CNN and cross dual branch transformer network (CDual transformer) in the encoder network to fully extract local and global features of images. To disentangle high-level semantic features, a style transfer module is designed to transform the style of hazy images while retaining the remaining semantic information. Afterward, the low-level features of images and the transformed high-level semantic features are used to reconstruct the dehazed images, fully capitalizing on the multilevel features. Furthermore, we built a meta-semi-supervised training strategy to improve the decoupling performance of the model and accumulated style knowledge of clear images from both synthetic and real-world hazy data, improving the generalizability of the model. Extensive experiments on both synthetic and real datasets show that the proposed algorithm effectively removes haze and offers better generalization abilities than similar methods. The project is publicly available at https://github.com/BJUTsipl/HDMDN. Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu |
IEEE Trans. Multim. | 2 |
| 2026 | RcFormer: Reconfigurable Self-Attention Transformer for Image RestorationabstractAdverse weather and imaging environments may degrade image quality and pose a significant challenge to the visual perception systems of multimedia. Various image restoration tasks necessitate the modeling of multiscale features, which is highly demanding on networks. To date, Vision Transformer has exhibited impressive image restoration performance. However, in this model, global self-attention is computationally expensive and local self-attention typically limits the interaction domain of each token. To solve this problem, we propose a novel reconfigurable self-attention transformer called RcFormer, which is designed to adequately model multiscale image features. This is achieved through a cross-grouped transformer (CGTransformer) block that uses convolution, area self-attention, and row-column self-attention for different head groups. CGTransformer is combined with an intragroup operation interaction structure. Moreover, an intergroup reconfigurable mechanism is implemented based on CGTransformer and channel circulation. The combination of multiple operations effectively enhances the modeling capability in the spatial and channel dimensions for various image recovery tasks. The performance of the proposed RcFormer is compared with low-level vision modules in a unified framework. Extensive experiments demonstrated that RcFormer exhibited a superior performance for the following image restoration tasks: image dehazing, rain streak removal, raindrop removal, snow removal, and single image deblurring. The source code is publicly available athttps://github.com/dehazing/RcFormer. Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023, Tianjian Yu |
IEEE Trans. Multim. | 2 |
| 2025 | Position Guided Dynamic Receptive Field Network: A Small Object Detection Friendly to Optical and SAR ImagesabstractObject detection in remote sensing images (RSIs), including optical and SAR images, has emerged as a rapidly advancing field. However, the abundance of small objects in RSIs poses a significant challenge in designing a network structure with effective receptive fields to support accurate localization and classification. In this paper, we propose a position guided dynamic receptive field network (PG-DRFNet) for small object detection friendly to optical and SAR images. Specifically, PG-DRFNet overcomes the problem of small objects vanishing or being submerged in features by establishing a positional guidance relationship of small objects between different feature layers. Then, we design a combination head structure that utilizes additional supervised information extracted from small objects to make the model more effective and flexible. Moreover, a dynamic perception algorithm based on feature construction is developed to dynamically optimize the perception regions and feature hierarchies of the model, while seeking the optimal tradeoff between model accuracy and inference speed. Without bells and whistles, our model is robust to two modalities of remote sensing data, and our experiments are conducted on four benchmark RSI datasets, including DOTA-v2.0, VEDAI, SSDD, and HRSID. The experimental results achieve competitive performance with 59.01%, 84.06%, 90.06%, and 80.59% mAP, respectively. Code and models are released athttps://github.com/BJUT-AIVBD/PG-DRFNet. Liuqian Wang, Jiafeng Li 0001, Jing Zhang 0023, Li Zhuo 0001, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hybrid-MambaCD: Hybrid Mamba-CNN Network for Remote Sensing Image Change Detection With Region-Channel Attention Mechanism and Iterative Global-Local Feature FusionabstractMamba has gained significant attention for its outstanding long-range context modeling capability while maintaining linear complexity, compared with Transformer. In this article, a hybrid Mamba and convolutional neural network (CNN) architecture is proposed for remote sensing image change detection (RSICD), named Hybrid-MambaCD, which leverages the advantages of CNN for local detail information extraction and Mamba for global context information extraction, providing an efficient solution for RSICD tasks. First, the region-channel attention mechanism (RCAM) is designed to enhance the CNN features from both channel and region dimensions, enabling the network to focus more on change regions while suppressing interference from background areas. Second, an iterative global-local feature fusion (IGLFF) strategy is proposed, which performs an adaptive weighted fusion of global and local features across multiple scales in a progressive manner, enhancing the representation ability of the features. Experimental results on three public datasets of LEVIR-CD, WHU-CD, and DSIFN-CD show that compared to the existing RSICD methods, the proposed Hybrid-MambaCD achieves the state-of-the-art (SOTA) detection performance. Li Zhuo 0001, Hui Zhang 0049, Jiafeng Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | TSTrack: A Lightweight Transformer-Based Spatiotemporal Feature Refinement Tracking AlgorithmabstractSingle-object tracking is a fundamental enabling technology in the field of remote sensing observation. It plays a crucial role in tasks such as unmanned aerial vehicle route surveillance and maritime vessel trajectory prediction. However, because of challenges such as the weak discriminative power of target features, interference from complex environments, and frequent viewpoint changes, existing trackers often suffer from insufficient temporal modeling capabilities and low computational efficiency, which limit their practical deployment. To address these challenges, we propose TSTrack, a novel lightweight single-object tracking framework that integrates Transformer and Mamba-based spatiotemporal modeling. First, we propose the target-aware feature purification preprocessor (TAFPP) , designed to dynamically enhance target representation through a synergistic combination of the dynamic position acuity module (DPAM) and spectral channel recalibrator (SCR). Second, we introduce the recurrent Mamba interaction pyramid (RM-IP) to replace traditional recurrent neural network-based structures, leveraging a state-space model for efficient and expressive temporal modeling with significantly reduced parameter overhead. Finally, we propose the elastic reconstructive multi-scale fusion (ERMSF) module, which adopts a four-branch parallel architecture to achieve effective multiscale feature fusion and dynamic shape adaptation, thereby enhancing robustness against target deformations and scale variations. Extensive experiments conducted on benchmark datasets, including LaSOT, TrackingNet, and GOT-10k, demonstrate the effectiveness of TSTrack. The results show that TSTrack achieves a superior tracking accuracy while maintaining a lightweight design, significantly outperforming existing state-of-the-art methods. The source code is publicly available at https://github.com/BJUTsipl/TSTrack. Jiafeng Li 0001, Shengyao Sun, Yang Wang 0023, Jing Zhang 0023, Li Zhuo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | CycFormer: Unsupervised Rain Removal Network Based on CycleGAN and TransformerabstractRainy weather presents significant challenges for applications relying on visual perception in intelligent transportation systems. The scarcity of real paired training data complicates single-image rain removal tasks, prompting an increasing interest in unsupervised methods capable of handling real-world rainy images without paired data. At present, most unsupervised rain removal methods are based on the CycleGAN framework; however, the combination of this framework and transformer is not satisfactory owing to most Transformers’ insufficient ability to model real rain features with global inhomogeneous distributions, which prevents them from being fully applicable to unsupervised tasks. This study devised an unsupervised rain removal network based on CycleGAN and the DerainFormer transformer. First, a deformable sparse attention mechanism was developed to improve the Transformer’s suitability for unsupervised tasks in CycleGAN architectures. Subsequently, a two-stage alternating transformer structure was designed to enhance its global non-uniform modeling capabilities for real rain images, In addition, a dual-channel parallel feed-forward network was used to establish the correlation between multiscale rain stripes. Finally, since rain removal is considered a decomposition task, a rain layer unsupervised training method for joint positional contrastive learning was proposed to separate the rain streaks effectively. We conducted several experiments on different real and synthetic rain datasets and the results confirmed that our unsupervised rain removal method performed well. The source code will be released athttps://github.com/derainsipl/CycFormer. Jiafeng Li 0001, Shuhao Yan, Jing Zhang 0023, Li Zhuo 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Unpaved road segmentation of UAV imagery via a global vision transformer with dilated cross window self-attention for dynamic map
Jing Zhang 0023, Jiafeng Li 0001, Li Zhuo 0001 |
Vis. Comput. | 3 |
| 2024 | HDUD-Net: heterogeneous decoupling unsupervised dehaze network
Jiafeng Li 0001, Lingyan Kuang, Jiaqi Jin, Li Zhuo 0001, Jing Zhang 0023 |
Neural Comput. Appl. | 1 |
| 2024 | Self-guided disentangled representation learning for single image dehazing
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Jing Zhang 0023 |
Neural Networks | 2 |
| 2024 | MPLA-Net: Multiple Pseudo Label Aggregation Network for Weakly Supervised Video Salient Object DetectionabstractWeakly Supervised Video Salient Object Detection (WSVSOD) only requires coarse-grained manual annotations, which can achieve a good trade-off between labeling efficiency and detection performance. In this paper, a Multiple Pseudo Label Aggregation Network (MPLA-Net) is proposed for WSVSOD. Firstly, the video frames that can obtain high-quality pseudo labels are selected to generate multiple pseudo labels, so as to avoid the prejudice of the single label. Moreover, the pseudo label with fine edge information is used to generate the Edge Information Map (EIM). Secondly, MPLA-Net is designed to adequately excavate and utilize the comprehensive saliency cues in multiple pseudo labels to improve the detection accuracy, in which ResNet-50 is adopted as the backbone network. Edge loss, pseudo label loss, self-supervised loss and fusion loss are exploited to jointly supervise and optimize the network training to obtain a robust detection model. Experimental results on five benchmark datasets demonstrate that, compared with existing weakly supervised methods, the proposed method can achieve state-of-the-art detection accuracy with less model parameters and higher detection speed. And the detected salient objects have fine boundaries. Chunjie Ma, Lina Du, Li Zhuo 0001, Jiafeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | GCFormer: Global Context-Aware Transformer for Remote Sensing Image Change DetectionabstractIn recent years, Transformer-based Change Detection (CD) in Remote Sensing Images has achieved significant advances, making it an emerging hot research topic. However, the current CD methods suffer from some problems, such as incomplete detection of change regions and missed detection of small change regions. In this paper, a Global Context-aware Transformer is proposed for CD tasks, named GCFormer, to address above issues by efficiently enhancing the global context information. It is fulfilled from two aspects based on the hybrid Convolutional Neural Network (CNN)+Transformer framework. Firstly, a Multi-Receptive Field Conv-Attention (MRFCA) mechanism is designed, which combines dilated convolutions with multiple rates and Conv-Attention, fully leveraging the advantages of convolution operation and self-attention mechanism. It is embedded at the highest layer of CNN to extract multi-receptive-field global context information. Secondly, a Context-aware Relative Position Encoding (CRPE) mode is proposed to replace the Absolute Position Encoding (APE) mode of Transformer. As a result, it can capture long-range dependency more efficiently and further enhance the global context information extraction and representation ability of the network. Experimental results on three public benchmark datasets of LEVIR-CD, WHU-CD and DSIFN-CD show that, the proposed GCFormer achieves superior detection performance with lower model complexity than the state-of-the-art Transformer-based CD methods. The source code is available at: https://github.com/yuwanting828/yuwanting828.github.io. Wanting Yu, Li Zhuo 0001, Jiafeng Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Semi-Supervised Single-Image Dehazing Network via Disentangled Meta-KnowledgeabstractCaptured outdoor scene images are easily affected by haze. Most image dehazing methods have limited generalization capabilities for real-world hazy images owing to the complexities of real-world environments and domain gaps in the training datasets. This article proposes a semi-supervised single-image dehazing network based on disentangled meta-knowledge. The symmetric and heterogeneous design of the disentangled network is conducive to the separation of the content and mask features of hazy images and these features are used as meta-knowledge to guide feature fusion in the dehazing network. Moreover, functions describing constant-color and disentangled-reconstruction-checking losses are designed to ensure the subjective qualities of the generated dehazed images. The results of extensive experiments conducted on synthetic datasets and real-world images indicate that the proposed algorithm outperforms state-of-the-art single-image dehazing algorithms. In addition, the algorithm effectively improves the performance of object-detection tasks. Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Tianjian Yu |
IEEE Trans. Multim. | 2 |
| 2023 | Graph Disentangled Representation Based Semi-supervised Single Image Dehazing Network
Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001 |
ICIC (2) | 2 |
| 2023 | Occluded prohibited object detection in X-ray images with global Context-aware Multi-Scale feature Aggregation
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023 |
Neurocomputing | 3 |
| 2023 | Few-Shot Remote Sensing Scene Classification With Spatial Affinity Attention and Class Surrogate-Based Supervised Contrastive LearningabstractLearning powerful and discriminative representation is critical for boosting the performance of Few-Shot Remote Sensing Scene Classification (FSRSSC). The remote sensing images have unique characteristics, such as, complex background and co-occurrence of multiple objects, making FSRSSC challenging. To address this problem, in this paper, a novel FSRSSC method is proposed. Firstly, a Spatial Affinity Attention (SAA) mechanism is designed to encourage the network model to focus on critical regions. The SAA infers attention maps from both channel and spatial dimensions, and encodes the mean values and affinities of feature nodes in each channel along the vertical and horizontal directions. Secondly, a Class Surrogate-based Supervised Contrastive Learning (CSSCL) strategy is proposed to promote intra-class compactness and inter-class dispersion. Different from Supervised Contrastive Learning (SCL), the CSSCL learns a surrogate for each class in the contrastive space and select the corresponding class surrogate rather than the samples from the same class for the anchor to form positive pairs. It can alleviate the impact of too hard or too simple positive sample pairs on model generalization that exists when SCL is introduced into FSRSSC. The model is trained under the joint supervision of Cross-Entropy (CE) loss and CSSCL loss on the merged base class data instead of using a meta-learning strategy to learn a robust base feature extractor. Extensive experimentations on three public remote sensing benchmark datasets show that our proposed method can achieve a competitive or state-of-the-art classification performance. Li Zhuo 0001, Jiafeng Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Efficient Fine-Grained Object Recognition in High-Resolution Remote Sensing Images From Knowledge Distillation to Filter GraftingabstractWith the development of high-resolution remote sensing images (HR-RSIs) and the escalating demand for intelligent analysis, fine-grained recognition of geospatial objects has become a more practical and challenging task. Although deep learning-based object recognition has achieved superior performance, it is inflexible to be directly utilized to the fine-grained object recognition tasks of HR-RSIs under the limitation of the size of geospatial objects. An efficient fine-grained object recognition method in HR-RSIs from knowledge distillation to filter grafting is proposed. Specifically, fine-grained object recognition consists of two stages: Stage 1 utilizes oriented region convolutional neural network (oriented R-CNN) to accurately locate and preliminarily classify geospatial objects. At the same time, it serves as a teacher network to guide students’ effective learning of fine-grained object recognition; in Stage 2, we design a coarse-to-fine object recognition network (CF-ORNet), as the second teacher network, which realizes fine-grained recognition through feature learning and category correction. After that, we propose a lightweight model from knowledge distillation to filter grafting on two teacher networks to achieve efficient fine-grained object recognition. The experimental results on VEDAI and HRSC2016 datasets achieve competitive performance. Liuqian Wang, Jing Zhang 0023, Jimiao Tian, Jiafeng Li 0001, Li Zhuo 0001, Qi Tian 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | BARRN: A Blind Image Compression Artifact Reduction Network for Industrial IoT SystemsabstractMost industrial Internet of Things (IoT) devices reduce the capture image size using high-ratio joint photographic experts group (JPEG) compression, saving storage space, and transmission bandwidth consumption. However, the resulting compression artifacts considerably affect the accuracy of subsequent tasks. Most artifact reduction algorithms do not consider the limitations of storage space and computing power of edge devices. In this study, a blind artifact reduction recurrent network (BARRN), which can reduce compression artifacts when the quality factors are unknown, is proposed. First, a structure based on recurrent convolution is designed for the specific requirements of industrial IoT image acquisition devices; the network can be scaled according to system resource constraints. Second, a more efficient convolution group, capable of adaptively processing different degradation levels, is proposed for optimal use of the limited computational resources. The experimental results demonstrate that the proposed BARRN can meet the needs of industrial systems with high computational efficiency. Jiafeng Li 0001, Yuqi Gao, Li Zhuo 0001, Jing Zhang 0023 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | USID-Net: Unsupervised Single Image Dehazing Network via Disentangled RepresentationsabstractCaptured images of outdoor scenes usually exhibit low visibility in cases of severe haze, which interferes with optical imaging and degrades image quality. Most of the existing methods solve the single-image dehazing problem by applying supervised training on paired images; however, in practice, the pairing of real-world images is not viable. Additionally, the processing speed of individual dehazing models is important in practical applications. In this study, a novel unsupervised single image dehazing network (USID-Net) based on disentangled representations without paired training images is explored. Furthermore, considering the trade-off between performance and memory storage, a compact multi-scale feature attention (MFA) module is developed, integrating multi-scale feature representation and attention mechanism to facilitate feature representation. To effectively extract haze information, a mechanism referred to as OctEncoder is designed to include multi-frequency representations that can capture more global information. Extensive experiments show that USID-Net achieves competitive dehazing results and a relatively high processing speed compared to state-of-the-art methods. The source code is available athttps://github.com/dehazing/USID-Net. Jiafeng Li 0001, Yaopeng Li, Li Zhuo 0001, Lingyan Kuang, Tianjian Yu |
IEEE Trans. Multim. | 1 |
| 2023 | Cascade Transformer Decoder Based Occluded Pedestrian Detection With Dynamic Deformable Convolution and Gaussian Projection Channel Attention MechanismabstractOccluded pedestrian detection is very challenging in computer vision, because the pedestrians are frequently occluded by various obstacles or persons, especially in crowded scenarios. In this article, an occluded pedestrian detection method is proposed under a basic DEtection TRansformer (DETR) framework. Firstly, Dynamic Deformable Convolution (DyDC) and Gaussian Projection Channel Attention (GPCA) mechanism are proposed and embedded into the low layer and high layer of ResNet50 respectively, to improve the representation capability of features. Secondly, Cascade Transformer Decoder (CTD) is proposed, which aims to generate high-score queries, avoiding the influence of low-score queries in the decoder stage, further improving the detection accuracy. The proposed method is verified on three challenging datasets, namely CrowdHuman, WiderPerson, and TJU-DHD-pedestrian. The experimental results show that, compared with the state-of-the-art methods, it can obtain a superior detection performance. Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023 |
IEEE Trans. Multim. | 3 |
| 2022 | Prohibited Object Detection in X-ray Images with Dynamic Deformable Convolution and Adaptive IoUabstractDue to the variety and complexity of objects in X-ray images, how to detect the prohibited items automatically and accurately is a challenging problem. In this paper, an X-ray image prohibited object detection method based on Dynamic Deformable Convolution (DyDC) and adaptive Intersection over Union (IoU) is proposed based on Cascade R-CNN framework. The main contributions are as follows. First, DyDC is proposed to cope with the diversity of the prohibited objects in X-ray images and to improve the feature representation capability. Then, adaptive IoU mechanism is proposed, which can dynamically adjust the IoU threshold during the training process to generate high quality proposals. The proposed method is extensively evaluated on two publicly available benchmark datasets, namely SIXray and OPIXray, and the experimental results show that it can achieve the state-of-the-art detection accuracy, compared with other existing methods. Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023 |
ICIP | 3 |
| 2022 | EAOD-Net: Effective anomaly object detection networks for X-ray imagesabstractAbstract Anomaly object detection is the core technology in the application for X‐ray images. However, the accuracy of current X‐ray anomaly object detection method still needs to be improved. In this paper, an effective anomaly object detection network is proposed to improve the detection accuracy of anomaly object for X‐ray images. Firstly, learnable Gabor convolution layer, deformable convolution, and spatial attention mechanism are introduced to enhance the representative ability of features in ResNeXt. Then, dense local regression is applied to predict the offset of multiple dense boxes in region proposal to locate the object accurately. At last, bigger discriminative RoI pooling is proposed to classify the candidate boxes to improve the accuracy of object classification. Experimental results on the SIXray and OPIXray datasets show that compared with the state‐of‐the‐art methods, the proposed EAOD‐Net can achieve the competitive detection performance. Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023 |
IET Image Process. | 3 |
| 2022 | VCFNet: video clarity-fluency network for quality of experience evaluation model of HTTP adaptive video streaming services
Lina Du, Jiafeng Li 0001, Li Zhuo 0001 |
Multim. Tools Appl. | 2 |
| 2022 | Effective Meta-Attention Dehazing Networks for Vision-Based Outdoor Industrial SystemsabstractHaze seriously affects the reliability of industrial systems, especially vision-based outdoor industrial systems such as autopilot systems. A majority of existing dehazing methods are not specifically designed for industrial systems and do not consider the reliability and resource cost of industrial system implementation. In this article, a novel meta-attention dehazing network (MADN) is proposed for direct restoration of clear images from hazy images without using the physical scattering model. Combined with parallel operation and enhancement modules, the meta-network automatically selects the most suitable dehazing network structure based on the current input hazy image by a meta-attention module. In addition, a novel feature loss calculated by the meta-network is proposed, which can accelerate the convergence of the dehazing network to meet the application requirements of practical industrial systems. A large number of experimental results on synthetic and real-world datasets show that the proposed MADN satisfies the needs of industrial systems. Tongyao Jia, Jiafeng Li 0001, Li Zhuo 0001, Guoqiang Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | A discriminative self-attention cycle GAN for face super-resolution and recognitionabstractAbstract Face image captured via surveillance videos in an open environment is usually of low quality, which seriously affects the visual quality and recognition accuracy. Most image super‐resolution methods adopt paired high‐quality and its interpolated low‐resolution version to train the super‐resolution network. It is difficult to achieve contented visual quality and restoring discriminative features in real scenarios. A discriminative self‐attention cycle generative adversarial network is proposed for real‐world face image super‐resolution. Based on the cycle GAN framework, unpaired samples are adopted to train a degradation network and a reconstruction network simultaneously. A self‐attention mechanism is employed to capture the contextual information for details restoring. A Siamese face recognition network is introduced to provide a constraint on identify consistency. In addition, an asymmetric perceptual loss is introduced to handle the imbalance between the degradation model and the reconstruction model. Experimental results show that the observation model achieved more realistic low‐quality face images, and the super‐resolved face images have shown better subjective quality and higher face recognition performance. Jianglu Huang, Li Zhuo 0001, Jiafeng Li 0001 |
IET Image Process. | 5 |
| 2021 | Blind image quality assessment with channel attention based deep residual network and extended LargeVis dimensionality reduction
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Meng Wang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | 3D Hand Pose Estimation with Disentangled Cross-Modal Latent SpaceabstractEstimating 3D hand pose from a single RGB image is a challenging task because of its ill-posed nature (i.e., depth ambiguity). Recently, various generative approaches have been proposed to predict the 3D joints of an RGB hand image by learning a unified latent space between two modalities (i.e., RGB image and 3D joints). However, projecting multi-modal data (i.e., RGB images and 3D joints) into a unified latent space is difficult as the modality-specific features usually interfere the learning of the optimal latent space. Hence in this paper, we propose to disentangle the latent space into two sub-latent spaces: modality- specific latent space and pose-specific latent space for 3D hand pose estimation. Our proposed method, namely Disentangled Cross-Modal Latent Space (DCMLS), consists of two variational autoencoder networks and auxiliary components which connect the two VAEs to align underlying hand poses and transfer modality-specific context from RGB to 3D. For the hand pose latent space, we align it with the two modalities by using a cross-modal discriminator with an adversarial learning strategy. For the context latent space, we learn a context translator to gain access to the cross-modal context. Experimental results on two widely used public benchmark datasets RHD and STB demonstrate that our proposed DCMLS method is able to clearly outperform the state-of-the-art ones on single image based 3D hand pose estimation. Jiajun Gu, Zhiyong Wang 0001, Wanli Ouyang, Jiafeng Li 0001, Li Zhuo 0001 |
WACV | 5 |
| 2020 | Effective Data-Driven Technology for Efficient Vision-Based Outdoor Industrial SystemsabstractVision systems are the core information collection module in outdoor industrial systems such as factory inspection robots. However, haze greatly reduces working efficiency. Existing dehazing methods have two problems-first, they are not specifically designed for the industrial systems; second, these methods include several assumptions in their design processes and imaging models, leading to unsatisfactory results. In this article, an approach for single image dehazing is proposed to improve the efficiency of outdoor vision-based systems. First, a novel haze imaging model is proposed based on the dichromatic atmospheric scattering model. It considers the effects of multiple scattering and involves fewer assumptions. Then a data-driven technique called sparse representation is used to solve this model. Considering a haze image, a distorted and blurred version of a fine image, every patch is presented using dedicatedly prepared over-complete dictionaries and is traced back to a haze-free image. Quantitative and qualitative comparisons on a number of real-world haze images demonstrate that the proposed approach not only is more stable but also leads to better dehazing results. Jiafeng Li 0001, Li Zhuo 0001, Hong Zhang 0018, Guoqiang Li 0001, Naixue Xiong |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Face Super-Resolution via Discriminative-Attributes
Jiafeng Li 0001, Li Zhuo 0001 |
PRCV (2) | 3 |
| 2019 | Deep-network based method for joint image deblocking and super-resolutionabstractMany pieces of research have been conducted on image‐restoration techniques to recover high‐quality images from their low‐quality versions, but they usually aim to handle a single degraded factor. However, captured images usually suffer from various degradation factors, such as low resolution and compression distortion, in the procedures of image acquisition, compression, and transmission simultaneously. Ignoring the correlation of different degraded factors may result in the limited efficiency of the existing image‐restoration methods for captured images. A joint deep‐network‐based image‐restoration algorithm is proposed to establish a restoration framework for image deblocking and super‐resolution. The proposed convolutional neural network is made up of two stages. A deblocking network is constructed with two cascade deblocking subnets first, then, super‐resolution is performed by a very deep network with skipping links. Cascading these two stages forms a novel deep network. An end‐to‐end training scheme is developed, which makes the two stages be trained jointly so as to achieve better performance. Intensive evaluations have been conducted to measure the performance of the authors’ method both in general images and face images. Experimental results on several datasets demonstrate that the proposed method outperforms other state‐of‐the‐art methods, in terms of both subjective and objective performances. Kin-Man Lam 0001, Li Zhuo 0001, Jiafeng Li 0001 |
IET Image Process. | 5 |
| 2018 | An efficient method of content-targeted online video advertising
Guanyao Wang, Li Zhuo 0001, Jiafeng Li 0001, Dongyue Ren, Jing Zhang 0023 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Vehicle color recognition using Multiple-Layer Feature Representations of lightweight convolutional neural network
Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023, Hui Zhang 0049 |
Signal Process. | 3 |
| 2017 | Tag tree creation of social image for personalized recommendationabstractThe tags are usually tagged by different users in social image sharing websites, which can indicate image semantic information and imply user's preference. Therefore, the tags can contribute to personalized recommendation of social image. However, the present social image tags models only consider single tag, resulting in the relationships among tags are ignored. In this paper, we propose a novel method to create tag tree of social image for personalized recommendation. Firstly, the tag ranking is realized to remove noisy tags. Then, the first layer tags are selected from re-ranked tags lists. To sufficiently express tag's significances, the tag subtrees can be created based on different image categories and combined with first layer tags to create tag tree. Finally, the personalized recommendation of social image is achieved by using tag tree. Experimental results show that our tag tree can effectively express the relationships among tags as well as obtain satisfactory results in personalized recommendation of social image. Ying Yang 0018, Jing Zhang 0023, Jihong Liu, Jiafeng Li 0001, Li Zhuo 0001 |
ICIP | 4 |
| 2017 | A joint deep-network-based image restoration algorithm for multi-degradationsabstractIn the procedures of image acquisition, compression, and transmission, captured images usually suffer from various degradations, such as low-resolution and compression distortion. Although there have been a lot of research done on image restoration, they usually aim to deal with a single degraded factor, ignoring the correlation of different degradations. To establish a restoration framework for multiple degradations, a joint deep-network-based image restoration algorithm is proposed in this paper. The proposed convolutional neural network is composed of two stages. Firstly, a de-blocking subnet is constructed, using two cascaded neural network. Then, super-resolution is carried out by a 20-layer very deep network with skipping links. Cascading these two stages forms a novel deep network. Experimental results on the Set5, Setl4 and BSD100 benchmarks demonstrate that the proposed method can achieve better results, in terms of both the subjective and objective performances. Li Zhuo 0001, Kin-Man Lam 0001, Jiafeng Li 0001 |
ICME | 5 |
| 2017 | Vehicle classification for large-scale traffic surveillance videos using Convolutional Neural Networks
Li Zhuo 0001, Liying Jiang, Jiafeng Li 0001, Jing Zhang 0023 |
Mach. Vis. Appl. | 4 |
| 2015 | Single image dehazing using the change of detail prior
Jiafeng Li 0001, Hong Zhang 0018, Ding Yuan 0001, Mingui Sun |
Neurocomputing | 1 |