Zhenxue Chen

dblp:122/2640 · DBLP profile ↗
← Back
55ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0001-9637-5170ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 UCGM: Enhancing pseudo labels via uncertainty and cross-image Gaussian Mixture Model for semi-supervised semantic segmentation
Zhenyan Wang, Zhenxue Chen, Chengyun Liu, Jiazheng Wu
Expert Syst. Appl.2
2026 Facial sketch synthesis with multi-level guided latent diffusion model
Dan Lu 0006, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu
Neurocomputing2
2026 ASNet: An adaptive scene-aware network for RGB-thermal urban scene semantic segmentation
Zhenxue Chen, Xuewen Rong, Chengyun Liu, Lili Song, Yidi Li 0001
J. Vis. Commun. Image Represent.2
2026 GGCN: Gait Recognition with Generate Network and Convolutional Neural Network
Hao Qin 0006, Zhenxue Chen, Qingqiang Guo, Q. M. Jonathan Wu, Mengxu Lu
J. Vis. Commun. Image Represent.2
2026 TSNUNet: Two-Stage Nested U-Network for salient object detection
Luna Sun, Zhenxue Chen, Xinming Zhu, Yu Bi, Chengyun Liu, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.2
2026 TAWNet: Three-dimensional adaptive weighted network for RGB-D salient object detection
Jiazheng Wu, Zhenxue Chen, Qingqiang Guo, Chengyun Liu, Zhenyan Wang, Qinggang Meng
Knowl. Based Syst.2
2025 Diversity augmentation and multi-fuzzy label for semi-supervised semantic segmentation
abstract
Semantic segmentation aims to provide pixel-wise accurate predictions for images. Semi-supervised semantic segmentation aims to learn a semantic segmentation model using a limited number of labeled images and a large fraction of unlabeled images. Existing methods primarily focus on introducing additional models or complex training procedures but overlook the model itself and such complex strategies tend to discard many usable pixels, exacerbating the class imbalance problem . In this paper, we propose DAM for semi-supervised semantic segmentation, a simple yet effective method that mainly focuses on the inputs and outputs of the model itself. For the input component, we posit that diverse data augmentations can provide more semantic information . Therefore, we propose a method called Random Diversity Augmentations. Given an unlabeled image, we apply different triple-level data augmentations to provide more semantic information. For the output component, our approach is inspired by the fact that many unreliable predictions are confused only among the top classes rather than all classes, so we contend that fuzzy pixels can still provide valuable guidance to the model. Specifically, we select fuzzy pixels based on confidence and assign multi-fuzzy labels to these pixels for training the model, which allows us to leverage the information more effectively. Our straightforward DAM achieves new state-of-the-art performance on SSS different benchmarks. Code is available at https://github.com/Wang-zhenyan/DAM .
Zhenyan Wang, Zhenxue Chen, Chengyun Liu, Xinming Zhu, Q. M. Jonathan Wu
Neurocomputing2
2025 DSINet: Dual semantic interaction and edge refinement network for salient object detection
Jiazheng Wu, Zhenxue Chen, Qingqiang Guo, Chengyun Liu, Hanxiao Zhai
Neurocomputing2
2025 Synergy-driven multi-modal prompting for weakly supervised semantic segmentation
Chengyun Liu, Zhenyan Wang, Xiaona Peng, Zhenxue Chen, Q. M. Jonathan Wu
Neurocomputing5
2025 3CNet: Cross-modal cooperative correction network for RGB-T semantic segmentation
Zhenxue Chen, Xuewen Rong, Chengyun Liu, Lili Song, Yidi Li 0001
Image Vis. Comput.2
2025 Few-Shot Facial Sketch Synthesis via Progressive Domain Gap Reduction
abstract
Facial sketch synthesis (FSS) has advanced significantly in recent years, but challenges remain in few-shot settings. Some few-shot learning methods can convert photos (source domain) into sketches of a specified style (target sketch domain). However, they overlook the available samples of other sketch styles (non-target sketch domains). We argue that the information in these samples can help the model enhance its mapping ability from the source domain to the target domain. This paper proposes a progressive domain gap reduction (PDGR) method for few-shot facial sketch synthesis, which consists of three stages: teacher training, knowledge distillation, and intra-domain few-shot adaptation. In the first stage, we adapt a pretrained StyleGAN to a non-target sketch domain with more available samples than the target sketch domain. To generate diverse and high-quality sketches, we employ a dual-discriminator adversarial mechanism to guide the model in focusing on the overall structure and style, as well as multi-scale details and textures. In the second stage, the knowledge from StyleGAN is transferred to a U-Net for more efficient image translation. In the third stage, we adapt the output of the U-Net from the non-target sketch domain to the target sketch domain in few-shot settings. To alleviate overfitting, preserve individual characteristics, and enhance detail representation, we leverage the FFHQ dataset to construct dual training paths and design a domain-directional triple loss. Experiments show that PDGR significantly outperforms previous few-shot learning methods and even outperforms the state-of-the-art FSS methods trained on the full dataset.
Dan Lu 0006, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu
IEEE Trans. Inf. Forensics Secur.2
2024 SAFLFusionGait: Gait recognition network with separate attention and different granularity feature learnability fusion
Zhenxue Chen, Chengyun Liu, Dan Lu 0006
J. Vis. Commun. Image Represent.2
2024 BNDCNet: Bilateral nonlocal decoupled convergence network for semantic segmentation
Mengting Ye, Zhenxue Chen, Kaili Yu, Longcheng Liu, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.2
2024 Category-based depth incorporation for salient object ranking
Hanxiao Zhai, Zhenxue Chen, Chengyun Liu, Huibin Bai, Q. M. Jonathan Wu
J. Vis. Commun. Image Represent.2
2024 AdaptiveGait: adaptive feature fusion network for gait recognition
Zhenxue Chen, Chengyun Liu, Jiyang Chen, Q. M. Jonathan Wu
Multim. Tools Appl.2
2024 Supervised contrastive learning with multi-scale interaction and integrity learning for salient object detection
Yu Bi, Zhenxue Chen, Chengyun Liu
Mach. Vis. Appl.2
2023 OIPNet: Multimodal Network with Orthogonal Information Processing for Semantic Segmentation in Indoor Scenes
abstract
Semantic segmentation in indoor environments is a crucial task for artificial intelligence-driven visual robotics, enabling pixel-level classification results to facilitate robot path planning. Inspired by the success of multimodal models, we propose an end-to-end multimodal semantic segmentation model for image segmentation tasks in indoor scenes, which we call OIPNet. We design the OIP module to enhance the network’s ability to extract global information and enable information interaction in different directions. We have validated on NYUv2 and Sun RGB-D datasets, and the experiments show the generality and effectiveness of the proposed model. Our code is available at https://github.com/Mantee0810/OIP.
Mengting Ye, Kaili Yu, Zhenxue Chen, Longcheng Liu
Int. J. Pattern Recognit. Artif. Intell.3
2023 Unsupervised self-attention lightweight photo-to-sketch synthesis with feature maps
Kunru Zhong, Zhenxue Chen, Chengyun Liu, Q. M. Jonathan Wu, Shuchao Duan
J. Vis. Commun. Image Represent.2
2023 Low-light image enhancement based on virtual exposure
Wencheng Wang 0002, Dongliang Yan, Xiaojin Wu, Weikai He, Zhenxue Chen, Xiaohui Yuan 0001
Signal Process. Image Commun.5
2022 M-PFGMNet: multi-pose feature generation mapping network for visual object tracking
Pei'en Luo, Tao Xu 0003, Zhenxue Chen
Multim. Tools Appl.4
2022 Improved edge-guided network for single image super-resolution
Zhenxue Chen, Q. M. Jonathan Wu, Xianming Li
Multim. Tools Appl.2
2022 Simple low-light image enhancement based on Weber-Fechner law in logarithmic space
Wencheng Wang 0002, Zhenxue Chen, Xiaohui Yuan 0001
Signal Process. Image Commun.2
2022 CATFPN: Adaptive Feature Pyramid With Scale-Wise Concatenation and Self-Attention
abstract
It is a typical problem in the field of object detection to simultaneously detect objects with large scale variation in one image. Recently proposed state-of-the-art object detectors generally learn pyramidal feature representation to deal with the scale variation, which has been proved effective via various feature pyramid networks. However, the majority of the feature pyramid networks based on heuristic feature fusion strategies may be suboptimal, as excess human guidance will restrict the self-learning of deep neural networks. An adaptive feature pyramid is bound to provide a significant performance boost. In this paper, we propose a novel feature pyramid network named CATFPN that consists of Scale-Wise Feature Concatenation (SWFC) module and Global Context (GC) block. The SWFC module evenly distributes semantic features for each feature layer and the GC block introduces a self-attention mechanism. As a feature pyramid network, the CATFPN can be applied to any detector based on multi-scale features. We adopt the CATFPN in typical RetinaNet and Faster R-CNN detector models, without bells and whistles, achieving 1.1% AP and 0.7% AP improvements over FPN on the MS COCO benchmark, respectively. Our competitive performance reported on the test-dev subset of COCO achieves 42.3% AP.
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu, Hui Yuan 0001, Weikai He
IEEE Trans. Circuits Syst. Video Technol.2
2022 RPNet: Gait Recognition With Relationships Between Each Body-Parts
abstract
At present, many studies have shown that partitioning the gait sequence and its feature map can improve the accuracy of gait recognition. However, most models just cut the feature map at a fixed single scale, which loses the dependence between various parts. So, our paper proposes a structure called Part Feature Relationship Extractor (PFRE) to discover all of the relationships between each parts for gait recognition. The paper uses PFRE and a Convolutional Neural Network (CNN) to form the RPNet. PFRE is divided into two parts. One part that we call the Total-Partial Feature Extractor (TPFE) is used to extract the features of different scale blocks, and the other part, called the Adjacent Feature Relation Extractor (AFRE), is used to find the relationships between each block. At the same time, the paper adjusts the number of input frames during training to perform quantitative experiments and finds the rule between the number of input frames and the performance of the model. Our model is tested on three public gait datasets, CASIA-B, OU-LP and OU-MVLP. It exhibits a significant level of robustness to occlusion situations, and achieves accuracies of 92.82% and 80.26% on CASIA-B under BG # and CL # conditions, respectively. The results show that our method reaches the top level among state-of-the-art methods.
Hao Qin 0006, Zhenxue Chen, Qingqiang Guo, Q. M. Jonathan Wu, Mengxu Lu
IEEE Trans. Circuits Syst. Video Technol.2
2022 MFNet: Multi-Feature Fusion Network for Real-Time Semantic Segmentation in Road Scenes
abstract
Although high-accuracy networks have been applied to semantic segmentation at present, their inference speeds remain slow. A trade-off between accuracy and speed is demanded for real-time applications. To approach this problem, we propose Multi-Feature Fusion Network (MFNet) with real-time efficient prediction capacity. MFNet adopts three branches (attention, semantic and spatial information) to capture low-level and high-level features. Additionally, MFNet exerts asymmetric factorized (AF) blocks to extract local and long-range features. As a result, without any pre-training or post-processing, MFNet using only 1.34 M parameters, achieves 72.1% mean intersection over union (mIoU) on the Cityscapes test set at a speed of 116 frames per second (FPS), with$512\times 1024$high resolution on a single Titan Xp graphics card. Our network’s performance stands out from other state-of-the-art networks on four datasets (Cityscapes, CamVid, KITTI, and Gatech).
Mengxu Lu, Zhenxue Chen, Chengyun Liu, Sile Ma, Hao Qin 0006
IEEE Trans. Intell. Transp. Syst.2
2022 FRNet: Factorized and Regular Blocks Network for Semantic Segmentation in Road Scene
abstract
Nowadays, semantic segmentation methods for systems in road scene have a great demand. Most existing methods focus on high accuracy with low inference speed. And some approaches emphasize on speed, significantly sacrificing model accuracy. To make a trade-off between accuracy and inference speed, we propose a real-time network for semantic segmentation titled Factorized and Regular Network (FRNet), which employs an asymmetric encoder-decoder architecture with Factorized and Regular (FR) blocks. Our method achieves 70.4% mIoU on the Cityscapes test set with 1 million parameters at a speed of 127 frames per second (FPS) on a single Titan Xp at a resolution of$512\times 1024$. We evaluate FRNet on Cityscapes, Camvid, Kitti, and Gatech datasets to identify that our network stands out from other state-of-the-art networks.
Mengxu Lu, Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001, Xuewen Rong, Xinghe Yan
IEEE Trans. Intell. Transp. Syst.2
2021 FSFN: feature separation and fusion network for single image super-resolution
Zhenxue Chen, Q. M. Jonathan Wu, Nannan Wang 0001
Multim. Tools Appl.2
2021 3MNet: Multi-task, multi-level and multi-channel feature aggregation network for salient object detection
Xinghe Yan, Zhenxue Chen, Q. M. Jonathan Wu, Mengxu Lu, Luna Sun
Mach. Vis. Appl.2
2021 AMPNet: Average- and Max-Pool Networks for Salient Object Detection
abstract
Salient Object Detection aims to detect the most visually distinctive objects in an image. We solve this problem by introducing the average pool to explore the multi-level deep average pool convolution features different from the max pool information. Based on the U-net structure, we propose an Average- and Max-Pool Network (AMPNet) that leverages the average- and max-pool modules to integrate the multi-level complementary contextual features in the spatial and channel-wise dimensions, respectively. The complementary contextual features generated by our network can improve the completeness of detected objects. It has been observed that the non-salient regions are misrecognized as the salient objects because of the redundant information contained in the multi-level convolution features. To address the problem, two top-down feedback paths are introduced based on the above two modules, and their top-level semantic guidance information is fully utilized to improve the accuracy of salient objects detection. Finally, we apply the Feature Fusion Module and Deep Supervision Mechanism to further improve the performance of the network over different datasets. Experimental results on six benchmark datasets show that our network is on par with state-of-the-art approaches. Our method runs at more than 45 FPS (based on VGG) and 35 FPS (based on ResNet) on a single GPU and meets real-time requirements.
Luna Sun, Zhenxue Chen, Q. M. Jonathan Wu, Hongjian Zhao, Weikai He, Xinghe Yan
IEEE Trans. Circuits Syst. Video Technol.2
2021 Multi-Scale Gradients Self-Attention Residual Learning for Face Photo-Sketch Transformation
abstract
Face sketch synthesis, as a key technique for solving face sketch recognition, has made considerable progress in recent years. Due to the difference of modality between face photo and face sketch, traditional exemplar-based methods often lead to missed texture details and deformation while synthesizing sketches. And limited to the local receptive field, Convolutional Neural Networks-based methods cannot deal with the interdependence between features well, which makes the constraint of facial features insufficient; as such, it cannot retain some details in the synthetic image. Moreover, the deeper the network layer is, the more obvious the problems of gradient disappearance and explosion will be, which will lead to instability in the training process. Therefore, in this paper, we propose a multi-scale gradients self-attention residual learning framework for face photo-sketch transformation that embeds a self-attention mechanism in the residual block, making full use of the relationship between features to selectively enhance the characteristics of specific information through self-attention distribution. Simultaneously, residual learning can keep the characteristics of the original features from being destroyed. In addition, the problem of instability in GAN training is alleviated by allowing discriminator to become a function of multi-scale outputs of the generator in the training process. Based on cycle framework, the matching between the target domain image and the source domain image can be constrained while the mapping relationship between the two domains is established so that the tasks of face photo-to-sketch synthesis (FP2S) and face sketch-to-photo synthesis (FS2P) can be achieved simultaneously. Both Image Quality Assessment (IQA) and experiments related to face recognition show that our method can achieve state-of-the-art performance on the public benchmarks, whether using FP2S or FS2P.
Shuchao Duan, Zhenxue Chen, Q. M. Jonathan Wu, Dan Lu 0006
IEEE Trans. Inf. Forensics Secur.2
2020 Face hallucination with K-means++ dictionary learning
Zhenxue Chen, Jiadi Li, Chengyun Liu
Multim. Tools Appl.1
2020 Deep mutual learning network for gait recognition
Zhenxue Chen, Q. M. Jonathan Wu, Xuewen Rong
Multim. Tools Appl.2
2020 Improved face super-resolution generative adversarial networks
Mengxue Wang, Zhenxue Chen, Q. M. Jonathan Wu, Muwei Jian
Mach. Vis. Appl.2
2020 3D video semantic segmentation for wildfire smoke
Guodong Zhu, Zhenxue Chen, Chengyun Liu, Xuewen Rong, Weikai He
Mach. Vis. Appl.2
2020 Pedestrian detection via deep segmentation and context network
Zhaoqing Li, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neural Comput. Appl.2
2020 3D Parallel Fully Convolutional Networks for Real-Time Video Wildfire Smoke Detection
abstract
Wildfires have devastating consequences on ecological systems and human lives. Accurate and fast wildfire detection is crucial to reduce damage. The existing smoke detection algorithms using convolution neural network are mostly based on the classification of smoke images or patches, whereas the traditional smoke detection algorithms are often necessary to extract multiple features for integration. With the methods mentioned above, false positive is always an insurmountable problem in wildfire smoke detection. Moreover, there are few studies on the detection of wildfire smoke. Thus, to detect the wildfire smoke more intelligent, a 3D parallel fully convolutional network for wildfire smoke detection is proposed to segment the smoke regions in video sequences. Wildfire smoke detection is considered as a segmentation problem in this paper. There are more than 90 videos including various scenes used for training and test. Experiments have demonstrated that our architecture can segment smoke regions accurately and eliminate the interference of natural scenes. Smoke targets in multiple scenes can be detected accurately and quickly.
Xiuqing Li, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
IEEE Trans. Circuits Syst. Video Technol.2
2019 FCN based preprocessing for exemplar-based face sketch synthesis
Dan Lu 0006, Zhenxue Chen, Q. M. Jonathan Wu, Xuetao Zhang 0003
Neurocomputing2
2019 Sketch Face Recognition: P-HOG Multi-Features Fusion
abstract
With the development of biometric recognition technology, sketch face recognition has been widely applied to assist the police to confirm the identity of the criminal suspect. Most of the present recognition methods use the image features directly, in which the key parts can’t be used sufficiently. This paper presents a sketch face recognition method based on P-HOG multi-features weighted fusion. Firstly, the global face image and the local face image which contains key components of the face are divided into patches based on spatial scale pyramid, and then the global P-HOG features and local P-HOG features are extracted, respectively. After that, the dimensions of global and local features are reduced using PCA and NLDA. Finally, the features are weighted based on sensitivity and fused. The nearest neighbor classifier is used to complete the final recognition. The experimental results on different databases show that the proposed method outperforms state-of-the-art methods.
Zhenxue Chen, Saisai Yao, Chengyun Liu
Int. J. Pattern Recognit. Artif. Intell.1
2019 Adaptive image enhancement method for correcting low-illumination images
Wencheng Wang 0002, Zhenxue Chen, Xiaohui Yuan 0001, Xiaojin Wu
Inf. Sci.2
2019 Saliency object detection: integrating reconstruction and prior
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Mach. Vis. Appl.2
2019 Two-stage local details restoration framework for face hallucination
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Mach. Vis. Appl.2
2019 Face recognition using AMVP and WSRC under variable illumination and pose
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neural Comput. Appl.2
2019 Fast Semantic Segmentation for Scene Perception
abstract
Semantic segmentation is a challenging problem in computer vision. Many applications, such as autonomous driving and robot navigation with urban road scene, need accurate and efficient segmentation. Most state-of-the-art methods focus on accuracy, rather than efficiency. In this paper, we propose a more efficient neural network architecture, which has fewer parameters, for semantic segmentation in the urban road scene. An asymmetric encoder-decoder structure based on ResNet is used in our model. In the first stage of encoder, we use continuous factorized block to extract low-level features. Continuous dilated block is applied in the second stage, which ensures that the model has a larger view field, while keeping the model small-scale and shallow. The down sampled features from encoder are up sampled with decoder to the same-size output as the input image and the details refined. Our model can achieve end-to-end and pixel-to-pixel training without pretraining from scratch. The parameters of our model are only 0.2M, 100× less than those of others such as SegNet, etc. Experiments are conducted on five public road scene datasets (CamVid, CityScapes, Gatech, KITTI Road Detection, and KITTI Semantic Segmentation), and the results demonstrate that our model can achieve better performance.
Xuetao Zhang 0003, Zhenxue Chen, Q. M. Jonathan Wu, Dan Lu 0006, Xianming Li
IEEE Trans. Ind. Informatics2
2019 Deep Saliency With Channel-Wise Hierarchical Feature Responses for Traffic Sign Detection
abstract
Traffic sign detection is challenging in cases of a complex background, occlusions, distortions, and so on. To overcome the above-mentioned challenges, this paper pays close attention to channel-wise feature responses to propose an end-to-end deep learning-based saliency traffic sign detection method. Our model contains three main components: channel-wise coarse feature extraction (CCFE), channel-wise hierarchical feature refinement (CHFR), and hierarchical feature map fusion (HFMF). In addition, it is based on the squeeze-and-excitation-residual network to explicitly model the inter dependences between the channels of its convolution features at a slight computational cost. We first apply CCFE to produce coarse feature maps with much information loss. To make full use of spatial information and fine details, CHFR is executed to refine hierarchical features. After that, HFMF is used to fuse hierarchical feature maps to generate the final traffic sign saliency map. Compared with other five traffic sign detection methods, the experimental results demonstrate the efficiency (a real-time speed) and superior performance of the proposed method according to comprehensive evaluations over three benchmark data sets.
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
IEEE Trans. Intell. Transp. Syst.2
2018 Deep saliency detection via channel-wise hierarchical feature responses
Cuiping Li 0003, Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Neurocomputing2
2018 Denoising and 3D Reconstruction of CT Images in Extracted Tooth via Wavelet and Bilateral Filtering
abstract
Three-dimensional reconstruction of teeth plays an important role in the operation of living dental implants. However, the tissue around teeth and the noise generated in the process of image acquisition bring a serious impact on the reconstruction results, which must be reduced or eliminated. Combined with the advantages of wavelet transform and bilateral filtering, this paper proposes an image denoising method based on the above methods. The method proposed in this paper not only removes the noise but also preserves the image edge details. The noise in high frequency subbands is denoised using a locally adaptive thresholding and the noise in low frequency subbands is filtered by the bilateral filtering. Peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM) and 3D reconstruction using the iso-surface extraction method are used to evaluate the denoising effect. The experimental results show that the proposed method is better than the wavelet denoising and bilateral filtering, and the reconstruction results meet the requirements of clinical diagnosis.
Cuizhen Wang, Zhenxue Chen
Int. J. Pattern Recognit. Artif. Intell.2
2018 Cascade heterogeneous face sketch-photo synthesis via dual-scale Markov Network
abstract
Heterogeneous face sketch-photo synthesis is an important and challenging task in computer vision, which has widely applied in law enforcement and digital entertainment. According to the different synthesis results based on different scales, this paper proposes a cascade sketch-photo synthesis method via dual-scale Markov Network. Firstly, Markov Network with larger scale is used to synthesise the initial sketches and the local vertical and horizontal neighbour search (LVHNS) method is used to search for the neighbour patches of test patches in training set. Then, the initial sketches and test photos are jointly entered into smaller scale Markov Network. Finally, the fine sketches are obtained after cascade synthesis process. Extensive experimental results on various databases demonstrate the superiority of the proposed method compared with several state-of-the-art methods.
Saisai Yao, Zhenxue Chen, Yunyi Jia, Chengyun Liu
J. Exp. Theor. Artif. Intell.2
2018 Face sketch-photo synthesis and recognition: Dual-scale Markov Network and multi-information fusion
Zhenxue Chen, Saisai Yao, Yunyi Jia, Chengyun Liu
J. Vis. Commun. Image Represent.1
2017 Illumination and pose variable face recognition via adaptively weighted ULBP_MHOG and WSRC
Zhenxue Chen, Q. M. Jonathan Wu, Chengyun Liu
Signal Process. Image Commun.2
2016 Fast Face Sketch-Photo Image Synthesis and Recognition
abstract
Face sketch recognition has great practical value in the criminal detection, security and other fields. Especially, it can help the police narrow down potential suspects in criminal detection effectively. Face sketch represents the original photos in a simple and recognizable form, so sketch and photo are images of two different modes. In order to identify the corresponding sketch face image in a lot of photo face images, this paper presents an improved sketch–photo transformation algorithm, and it uses the effective characteristics of the photo image more reasonably during transforming a photo image into sketch. In this way, it can reduce the difference between the sketch and photo image to improve the matching effect, and save the recognition time. Many experiments on CUHK Face Sketch database including 188 sketch–photos prove the effectiveness of the method in this paper.
Zhenxue Chen, Kaifang Wang, Chengyun Liu
Int. J. Pattern Recognit. Artif. Intell.1
2016 Low-Resolution Face Recognition of Multi-Scale Blocking CS-LBP and Weighted PCA
abstract
A novel method is proposed in this paper to improve the recognition accuracy of Local Binary Pattern (LBP) on low-resolution face recognition. More precise descriptors and effectively face features can be extracted by combining multi-scale blocking center symmetric local binary pattern (CS-LBP) based on Gaussian pyramids and weighted principal component analysis (PCA) on low-resolution condition. Firstly, the features statistical histograms of face images are calculated by multi-scale blocking CS-LBP operator. Secondly, the stronger classification and lower dimension features can be got by applying weighted PCA algorithm. Finally, the different classifiers are used to select the optimal classification categories of low-resolution face set and calculate the recognition rate. The results in the ORL human face databases show that recognition rate can get 89.38% when the resolution of face image drops to 12[Formula: see text]10 pixel and basically satisfy the practical requirements of recognition. The further comparison of other descriptors and experiments from videos proved that the novel algorithm can improve recognition accuracy.
Jiadi Li, Zhenxue Chen, Chengyun Liu
Int. J. Pattern Recognit. Artif. Intell.2
2016 Fast Traffic Sign Recognition via High-Contrast Region Extraction and Extended Sparse Representation
abstract
In this paper, we propose a high-performance traffic sign recognition (TSR) framework to rapidly detect and recognize multiclass traffic signs in high-resolution images. This framework includes three parts: a novel region-of-interest (ROI) extraction method called the high-contrast region extraction (HCRE), the split-flow cascade tree detector (SFC-tree detector), and a rapid occlusion-robust traffic sign classification method based on the extended sparse representation classification (ESRC). Unlike the color-thresholding or extreme region extraction methods used by previous ROI methods, the ROI extraction method of the HCRE is designed to extract ROI with high local contrast, which can keep a good balance of the detection rate and the extraction rate. The SFC-tree detector can detect a large number of different types of traffic signs in high-resolution images quickly. The traffic sign classification method based on the ESRC is designed to classify traffic signs with partial occlusion. Instead of solving the sparse representation problem using an overcomplete dictionary, the classification method based on the ESRC utilizes a content dictionary and an occlusion dictionary to sparsely represent traffic signs, which can largely reduce the dictionary size in the occlusion-robust dictionaries and achieve high accuracy. The experiments demonstrate the advantage of the proposed approach, and our TSR framework can rapidly detect and recognize multiclass traffic signs with high accuracy.
Chunsheng Liu 0001, Faliang Chang, Zhenxue Chen, Dongmei Liu 0007
IEEE Trans. Intell. Transp. Syst.3
2014 Illumination Processing in Face Recognition
abstract
Changes in light intensity and angle present a major challenge to the creation of reliable face recognition systems. The existence of bright regions and dark regions has been shown to have a serious negative impact on the performance of face recognition systems. This paper proposes a solution to this problem based on self-quotient image (SQI) processing method. In this method, bright and dark areas are processed separately without changing the essential characteristics of the image of the face. The dark and light areas are processed separately by SQI. Experimental results indicate that this Single-Light-Region and Single-Dark-Region SQI method removes the adverse effect of multi-bright and multi-dark areas better than competing methods.
Zhenxue Chen, Chengyun Liu, Faliang Chang, Xuzhen Han, Kaifang Wang
Int. J. Pattern Recognit. Artif. Intell.1
2014 Rapid Multiclass Traffic Sign Detection in High-Resolution Images
abstract
This paper describes a traffic sign detection (TSD) framework that is capable of rapidly detecting multiclass traffic signs in high-resolution images while achieving a high detection rate. There are three key contributions. The first is the introduction of two features called multiblock normalization local binary pattern (MN-LBP) and tilted MN-LBP (TMN-LBP), which are able to express multiclass traffic signs effectively. The second is a tree structure called split-flow cascade, which utilizes common features of multiclass traffic signs to construct a coarse-to-fine TSD detector. The third contribution is the Common-Finder AdaBoost (CF.AdaBoost) algorithm, which is designed to find common features of different training sets to develop an efficient Split-Flow Cascade tree (SFC-tree) for multiclass TSD. Through experiments with an evaluation data set of high-resolution images, we show that the proposed framework is able to detect multiclass traffic signs with high detection accuracy in real time and that it outperforms the state-of-the-art approaches at detecting a large number of different types of traffic signs rapidly without using any color information.
Chunsheng Liu 0001, Faliang Chang, Zhenxue Chen
IEEE Trans. Intell. Transp. Syst.3
2013 Chinese License Plate Recognition Based on Human Vision Attention Mechanism
abstract
License plate recognition (LPR) is one of the most important elements affecting intelligent transportation systems. A number of LPR techniques have been proposed. Humans are good target recognition systems. In other words, humans easily recognize common objects. In this paper, the researchers present a novel method of recognizing Chinese license plates. The method is based on the Human Vision Attention Mechanism (HVAM) and uses Chinese license plates as the targets. The research consists of three stages. The first stage involved finding and identifying license plates in videos of moving vehicles. The second stage separated each license plate into the seven characters. In the third stage, the character recognizer extracted some salient features of Chinese characters and used a multi-stage classifier to recognize each character on the license plate. In the experiment locating license plates, 1176 images taken from various scenes and conditions were employed. The method failed to identify the license plates in only 27 of the images; resulting in a license plate location rate of success of 97.7%. In the experiment for identifying license characters, 1149 images were used, from which license plates had been successfully located. The method failed to identify the characters in 45 of these images giving a success rate of 96.1%. Combining the above two rates, the overall rate of success for our LPR is 93.9%.
Zhenxue Chen, Faliang Chang, Chunsheng Liu 0001
Int. J. Pattern Recognit. Artif. Intell.1