Kangli Zeng

dblp:226/2532 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-8592-053XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 An angle-guided bidirectional feature transformation network for multi-frame tilt-angle face recognition
Wenqin Song, Xihao Wang, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001
Expert Syst. Appl.4
2026 MSAW-Net: Dual-Path Fusion Network With Multi-Scale Spatiotemporal Attention Weighting for Human Action Recognition
abstract
Human action recognition methods based on multimodal channels, especially the strategy combining RGB and skeleton information, have achieved remarkable success. However, the existing action recognition methods still face challenges in terms of insufficient multimodal fusion, imperfect multi-scale modeling, and inadequate preservation of spatial information. This paper proposes an innovative framework: by combining early bidirectional lateral connections with late fusion, deep interaction between modalities and complementarity at the prediction layer are achieved; the Multi-scale Spatiotemporal Attention Weighting (MSAW) mechanism is designed to adaptively fuse spatiotemporal features across different scales, selecting key frames in spatiotemporal sequences; the Pixel Sub-sampling Downsampling (PSD) module is proposed, which effectively retains spatial details during the strong downsampling process. A large number of experiments on the NTU RGB+D 60, NTU RGB+D 120, and Kinetics-400 datasets verified the effectiveness of the proposed method. The results show that our method demonstrates stronger competitiveness compared with the current state-of-the-art (SOTA) methods. Our code is available at:https://github.com/course-prog/MASW
Hui Wang 0091, Kangli Zeng, Chunhui Zhao 0001, Hui Yang 0005
IEEE Signal Process. Lett.3
2025 CrowdFPN: crowd counting via scale-enhanced and location-aware feature pyramid network
Hamido Fujita, Jiamao Yu, Kangli Zeng, Enhong Chen
Appl. Intell.6
2025 Multi-Stage Statistical Texture-Guided GAN for Tilted Face Frontalization
abstract
Existing pose-invariant face recognition mainly focuses on frontal or profile, whereas high-pitch angle face recognition, prevalent under surveillance videos, has yet to be investigated. More importantly, tilted faces significantly differ from frontal or profile faces in the potential feature space due to self-occlusion, thus seriously affecting key feature extraction for face recognition. In this paper, we asymptotically reshape challenging high-pitch angle faces into a series of small-angle approximate frontal faces and exploit a statistical approach to learn texture features to ensure accurate facial component generation. In particular, we design a statistical texture-guided GAN for tilted face frontalization (STG-GAN) consisting of three main components. First, the face encoder extracts shallow features, followed by the face statistical texture modeling module that learns multi-scale face texture features based on the statistical distributions of the shallow features. Then, the face decoder performs feature deformation guided by the face statistical texture features while highlighting the pose-invariant face discriminative information. With the addition of multi-scale content loss, identity loss and adversarial loss, we further develop a pose contrastive loss of potential spatial features to constrain pose consistency and make its face frontalization process more reliable. On this basis, we propose a divide-and-conquer strategy, using STG-GAN to progressively synthesize faces with small pitch angles in multiple stages to achieve frontalization gradually. A unified end-to-end training across multiple stages facilitates the generation of numerous intermediate results to achieve a reasonable approximation of the ground truth. Extensive qualitative and quantitative experiments on multiple-face datasets demonstrate the superiority of our approach.
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Chao Liang 0001, Zhen Han 0002
IEEE Trans. Image Process.1
2023 LSA3D: Lightweight Separate Asynchronous 3D Convolutional Neural Network for Gait Recognition
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Jinsheng Xiao, Zhen Han 0002
ICANN (10)3
2023 Multi-frame Tilt-angle Face Recognition Using Fusion Re-ranking
Wenqin Song, Zhen Han 0002, Kangli Zeng, Zhongyuan Wang 0001
ICANN (2)3
2023 Structure-Aware Multi-Feature Co-Learning for Dual Branch Face Super Resolution
abstract
Recently, face super-resolution has achieved pleasing performance. Numerous works have shown that texture features and structural information play a crucial role for super-resolution reconstruction. However, effective co-learning of both has been limiting the performance improvement of existing state-of-the-art methods. Therefore, we focus on the texture and structure of images in this paper, and design a two-branch network containing a texture network (T-Net) and a structure network (S-Net) to jointly explore texture and structure information for co-learning. T-Net serves as the backbone network to learn both texture and structure information for reconstruction, while the S-Net serves as the auxiliary network that can effectively exploit the multi-scale information of the T-Net encoder to recover the structure. To better facilitate the co-learning of the two branches, two co-learning modules deal with the information flow interaction between the encoder and decoder of the two branches, respectively, thus explicitly guiding the structure-aware image reconstruction. Additionally, a dense feature enhancement module investigates the channel and spatial correlation of features and enhances the representation capability of the network. Extensive experiments on the CelebA and Helen datasets show that our proposed approach outperforms state-of-the-art methods.
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008
ICASSP1
2023 Efficient face image super-resolution with convenient alternating projection network
abstract
Abstract The existing deep learning‐based face super‐resolution techniques can achieve satisfactory performance. However, these methods often incur large computational costs, and deeper networks generate redundant features. Some lightweight reconstruction networks also present limited representation ability because they ignore the entire contour and fine texture of the face for the sake of efficiency. Here, the authors propose a convenient alternating projection network (CAPN) for efficient face super‐resolution. First, the authors design a novel alternating projection block cascaded convolutional neural network to alternately achieve content consistency and learn detailed facial feature differences between super‐resolution and ground‐truth face images. Second, the self‐correction mechanism enabled the convolutional layer to capture faithful features that facilitate adaptive reconstruction. Moreover, a convenient connection operation can reduce the generation of redundant facial features while maintaining accurate reconstruction information. Extensive experiments demonstrated that the proposed CAPN can effectively reduce the computational cost while achieving competitive qualitative and quantitative results compared to state‐of‐the‐art super‐resolution methods.
Xitong Chen, Yuntao Wu, Jiangchuan Chen, Kangli Zeng
IET Signal Process.5
2023 GaitAMR: Cross-view gait recognition via aggregated multi-feature representation
Jianyu Chen 0008, Zhongyuan Wang 0001, Caixia Zheng, Kangli Zeng, Qin Zou 0001, Laizhong Cui
Inf. Sci.4
2023 Self-attention learning network for face super-resolution
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Jiaming Wang 0001, Zixiang Xiong
Neural Networks1
2023 Implicit space pose consistent transfer network for deep face verification
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Zhen Han 0002
Pattern Recognit. Lett.1
2022 Video Face Recognition Using Neural Aggregation Networks with Mutual Relational Learning
abstract
Video face recognition benefits profoundly from deep convolutional neural networks (CNNs), which learn robust feature embeddings. However, due to their fixed geometric structures, CNNs are inherently limited in modeling the significant variations from the angle, pose, occlusion and other factors of face images. In this paper, a neural aggregation network based on mutual relation learning is proposed for video face recognition. First, Intra-frame Relational Learning network (Intra-Net) is introduced, which models the interdependencies between the re-gional components of individual features and develops relevance between fine-grained features. Such processing can determine the region of interest adaptively according to the quality of the input face image to achieve the extraction of valuable information. Secondly, we introduce Inter-frame Relational Learning Network (Inter-Net), which considers the most significant appearance representation in the overall structure of the face image to cor-relate the complementarity of features between frames. Finally, information aggregation is performed by combining Inter-Net and Intra-Net. Joint optimization of the two branches allows our model to effectively exploit the complementary information between them to improve the aggregation capability. We validate the effectiveness of our model for video face recognition, proving its superiority over state-of-the-art methods on two benchmark datasets.
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008
ICTAI1
2022 Realistic frontal face reconstruction using coupled complementarity of far-near-sighted face images
Kangli Zeng, Zhongyuan Wang 0001, Tao Lu 0001, Jianyu Chen 0008, Baojin Huang, Zhen Han 0002, Xin Tian 0006
Pattern Recognit.1
2022 Rethinking Lightweight: Multiple Angle Strategy for Efficient Video Action Recognition
abstract
Video action recognition task involves modeling spatiotemporal information, and efficiency is critical to capture spatiotemporal dependencies in the video. Most existing models rely on optical flow information to capture the dynamic visual tempos between consecutive video frames. Although impressive performance can be achieved by combining optical flow with RGB, the time-consuming nature of optical flow computation cannot be ignored. Moreover, 3D CNN has successfully modeled spatiotemporal information, yet the enormous computational volume is unsuitable for real-time action recognition. In this letter, we propose a novel lightweight video feature extraction strategy that achieves better recognition performance with lower FLOPs. In particular, we perform convolution on the video cube from three orthogonal angles to learn its appearance and motion features. Compared with the computational volume of 3D CNN, our proposed method is more economical and thus meets the lightweight requirements. Extensive experimental results on public Something Something-V1$\&$V2 and Diving48 datasets show our approach achieves the state-of-the-art performance.
Jianyu Chen 0008, Zhongyuan Wang 0001, Kangli Zeng, Zheng He 0001, Zixiang Xiong
IEEE Signal Process. Lett.3
2021 When Face Recognition Meets Occlusion: A New Benchmark
abstract
The existing face recognition datasets usually lack occlusion samples, which hinders the development of face recognition. Especially during the COVID-19 coronavirus epidemic, wearing a mask has become an effective means of preventing the virus spread. Traditional CNN-based face recognition models trained on existing datasets are almost ineffective for heavy occlusion. To this end, we pioneer a simulated occlusion face recognition dataset. In particular, we first collect a variety of glasses and masks as occlusion, and randomly combine the occlusion attributes (occlusion objects, textures,and colors) to achieve a large number of more realistic occlusion types. We then cover them in the proper position of the face image with the normal occlusion habit. Furthermore, we reasonably combine original normal face images and occluded face images to form our final dataset, termed as Webface-OCC. It covers 804,704 face images of 10,575 subjects, with diverse occlusion types to ensure its diversity and stability. Extensive experiments on public datasets show that the ArcFace retrained by our dataset significantly outperforms the state-of-the-arts. Webface-OCC is available at https://github.com/Baojin-Huang/Webface-OCC.
Baojin Huang, Zhongyuan Wang 0001, Guangcheng Wang, Kui Jiang, Kangli Zeng, Zhen Han 0002, Xin Tian 0006, Yuhong Yang 0001
ICASSP5
2020 Low-quality watermarked face inpainting with discriminative residual learning
abstract
Most existing image inpainting methods assume that the location of the repair area (watermark) is known, but this assumption does not always hold. In addition, the actual watermarked face is in a compressed low-quality form, which is very disadvantageous to the repair due to compression distortion effects. To address these issues, this paper proposes a low-quality watermarked face inpainting method based on joint residual learning with cooperative discriminant network. We first employ residual learning based global inpainting and facial features based local inpainting to render clean and clear faces under unknown watermark positions. Because the repair process may distort the genuine face, we further propose a discriminative constraint network to maintain the fidelity of repaired faces. Experimentally, the average PSNR of inpainted face images is increased by 4.16dB, and the average SSIM is increased by 0.08. TPR is improved by 16.96% when FPR is 10% in face verification.
Zheng He 0001, Xueli Wei, Kangli Zeng, Zhen Han 0002, Qin Zou 0001, Zhongyuan Wang 0001
MMAsia3
2020 Face super-resolution via nonlinear adaptive representation
Tao Lu 0001, Kangli Zeng, Shenming Qu, Yanduo Zhang
Neural Comput. Appl.2
2019 Face super-resolution via bilayer contextual representation
Kangli Zeng, Tao Lu 0001, Xuefeng Liang, Kai Li 0005, Yanduo Zhang
Signal Process. Image Commun.1
2018 Contextual-Field Supported Iterative Representation for Face Hallucination
Kangli Zeng, Tao Lu 0001, Yanduo Zhang, Li Peng 0003, Shenming Qu
ICA3PP (3)1
2018 Face Hallucination Using Manifold-Regularized Group Locality-Constrained Representation
abstract
Sparsity and locality regularizations are successfully applied to face hallucination algorithms to ameliorate their ill-posed nature. However, most of patch-based face hallucination approaches only consider the manifold structure of single patch, thus resulting in unstable solution for image reconstruction. In this paper, we propose a novel face hallucination, termed manifold-regularized group locality-constrained representation (MGLR), in order to exploit the multiple manifold structures rooted in grouped self-similarly patches. Specifically, we first group similar patches to form a matrix which contains the recurrent non-local patches. Then graph regularization term is formulated to represent the group manifolds for better reconstruction quality. Taking advantages of grouped self-similar patches, MGLR can offer stable sparse solution to take advantage of the the accurate prior for super-resolution reconstruction. Experimental results on LFW database and CMU real-world images demonstrate the superiority of the proposed method over some state-of-the-art face methods both in terms of subjective and objective qualities.
Tao Lu 0001, Kangli Zeng, Junjun Jiang, Yanduo Zhang, Zhongyuan Wang 0001, Huabing Zhou
ICIP2