Wenzhe Zhai

dblp:301/3470 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0003-0996-6832ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distributed quantum model learning for traffic density estimation
Kewen Wang 0014, Wenzhe Zhai, Jingan Cheng
Image Vis. Comput.3
2026 Federated Learning for Edge Computing Enabled Artificial Intelligence of Things: A comprehensive survey
abstract
Among contemporary AI computing paradigms, Federated Learning (FL) stands out as an innovative method and has shown great potential in conjunction with edge computing. The two techniques combined serve as a building block forthe development of the Artificial Intelligence of Things (AIoT). This paper sheds light on the synergistic integration of FL with edge computing to propel AIoT’s capabilities in decentralized environments. By executing computing tasks closer to the data, FL at the edge not only alleviates latency and bandwidth limitations inherent in cloud-centric architectures, but also presents a robust solution to privacy concerns—a crucial obstacle in traditional centralized training setups. This paper delves into how FL tackles these privacy issues, providing an intricate explanation of its operational principles, applications, and the resultant benefits for AIoT systems. Through this scrutiny, we highlight FL’s potential in bolstering the efficiency and privacy of AIoT deployments while also delineating future research directions and the expected impact across various domains. This study aims to comprehensively comprehend FL for Edge Computing-enabled AIoT and foster developments in intelligent technologies and applications in an interconnected world.
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Wentai Wu, Chen Wang 0011, Ahmed M. Abdelmoniem
Knowl. Based Syst.3
2025 SCANSleepNet: A spatial-channel attention network for sleep stage classification
Yuyun Liu, Qilei Li, Mingliang Gao 0001, Wenzhe Zhai
Appl. Intell.5
2025 Privacy-Preserving Crowd Counting via Quantum-Enhanced Federated Learning
abstract
ABSTRACT Crowd counting plays a crucial role in analyzing group behavior in smart cities. Traditional crowd‐counting models rely on large datasets gathered from diverse individuals for training while ignoring the privacy protection for each training client. Meanwhile, the scale variation has long been a difficult problem in crowd counting and has greatly reduced model accuracy. Therefore, it is essential to achieve privacy‐aware crowd counting and to solve the problem of scale variation in dense scenes. To this end, we propose a Privacy‐preserving Quantum‐enhanced Network (PQNet). The PQNet uses federated learning to share parameters rather than data, which ensures the privacy of each client. Subsequently, a multi‐scale quantum‐driven calibration module is designed to capture multi‐scale information via quantum states. It enhances counting accuracy in dense crowd environments where scale varies. Experiments on four crowd counting and two vehicle counting benchmarks demonstrate that PQNet outperforms state‐of‐the‐art methods subjectively and objectively. The code will be available at: https://github.com/sdutzhangchen/PQNet .
Jingan Cheng, Wenzhe Zhai, Mingliang Gao 0001
Expert Syst. J. Knowl. Eng.4
2025 Cross-Camera Discriminative Person Association by Unsupervised Frame Clustering and Selection
abstract
The objective of cross-camera persson association is to identify individuals captured across disjoint cameras. It is achieved by Re-Identification (ReID) models, which extract unique identity representations from the visual input, dominated by video sequence. Most ReID methods primarily focus on modifying the backbone network architecture to learn more representative features of people. However, these methods often overlook the impact of low-quality frames on the training process. Several studies have confirmed that low-quality data not only hinders the model from learning meaningful content but also diminishes its performance. One possible solution is to manually label the quality of each frames, but this is time-consuming and inefficient. In this paper, we propose a Unsupervised Frame Clustering and Selection framework called UFCS to address this problem by applying the unsupervised clustering to automatically select high-quality frames. Specifically, we applied three unsupervised clustering solutions for high-quality frame selection, namely K-means, Deep K-means, and DBSCAN. These clustering techniques integrate both appearance and, indirectly, temporal consistency by operating within tracklets. These schemes perform clustering in the image or deep feature space to select highquality frames for network training. This straightforward yet effective approach enables the ReID network to generate a more discriminative representations, thereby improving recognition performance. Experimental results obtained on the challenging video-based person ReID datasets MARS indicate that our proposed scheme can outperform related state-of-the-art methods by a large margin.
Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai, Gwanggil Jeon
IEEE Internet Things J.4
2025 Composed image retrieval by Multimodal Mixture-of-Expert Synergy
Wenzhe Zhai, Mingliang Gao 0001, Gwanggil Jeon, David Camacho
Image Vis. Comput.1
2025 Towards text-refereed multi-modal image fusion by cross-modality interaction
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai
Signal Process.4
2025 Training-Free 3-D Face Avatars Generation by Knowledge Discovering in Foundational Models
abstract
An informative 3-D avatar, closely mirroring real-world traits, plays a pivotal role in accessing the metaverse. Traditional methods for creating 3-D avatars usually employ one-to-one training, which restricts avatar diversity. To enhance style diversity in generated 3-D avatars, we utilize synthesized images derived with prompts from ChatGPT in a conversational manner, ultimately resulting in a broader range of 3-D variations. Rather than creating models from scratch, we devise a training-free framework that utilizes established large-scale foundation models. Specifically, we employ a real-world image synthesis technique guided by text prompts that are generated by ChatGPT in a conversational manner, to describe the desired characteristics of the synthesized image. As a result, these informative latent representations can accurately reflect the distinct style of the synthesized image, and further lead to the creation of photorealistic and diverse 3D avatars. Our training-free design allows this proposed method to achieve competitive performance compared to existing generation models, while requiring minimal computational resources.
Qilei Li, Wenzhe Zhai, David Camacho, Gwanggil Jeon
IEEE Trans. Comput. Soc. Syst.3
2025 Multi-View Gait Recognition With Joint Local Multi-Scale and Global Contextual Spatio-Temporal Features
abstract
Existing gait recognition methods are capable of extracting rich spatial gait information but often overlook fine-grained temporal features within local regions and temporal contextual information across different sub-regions. Considering gait recognition as a fine-grained recognition task and each individual exhibits uniqueness in their movements across different temporal sequences, we propose a local multi-scale and global contextual spatio-temporal (LMGCS) network for gait recognition. It divides the whole gait sequence into sub-sequences with multiple spatio resolutions and extracts multi-scale temporal features. We extract the temporal context information of different sub-sequences with the transformer, and all sub-sequences are fused to form global features. Furthermore, the loss function that combines the triplet loss function and cross-entropy loss function is utilized to prompt the proposed model to fulfill the gait recognition. The proposed method achieved state-of-the-art results on two popular public datasets. It achieved rank-1 accuracy of 98.0%, 95.4%, and 85.0% on the three walk states of the CASIA-B dataset and 90.9% on the OU-MVLP dataset.
Wenzhe Zhai, Haomiao Li, Chaoqun Zheng, Xianglei Xing
IEEE Trans. Circuits Syst. Video Technol.1
2025 Zero-Shot Object Counting With Vision-Language Prior Guidance Network
abstract
The majority of existing counting models are designed to operate on a singular object category, such as crowds or vehicles. The emergence of multi-modal foundational models, e.g., Contrastive Language-Image Pre-training (CLIP), has paved the way for class-agnostic counting. This approach facilitates the counting of objects across diverse classes within a single image based on textual indications. However, class-agnostic counting models based on CLIP confront two primary challenges. Firstly, the CLIP model exhibits limited sensitivity towards location information, which prioritizes global content over the precise localization of objects. Therefore, directly employing the CLIP model is regarded as suboptimal. Secondly, these models commonly employ frozen pre-trained vision and language encoders while disregarding potential misalignment within the constructed hypothesis space. In this paper, we propose a unified framework, named the Vision-Language Prior Guidance (VLPG) Network, to tackle these two challenges. The VLPG consists of three key components, namely the Grounding DINO module, Spatial Prior Calibration (SPC) module, and Object-Centric Alignment (OCA) module. The Grounding DINO module utilizes the spatial-awareness capability of extensive pre-trained object grounding models to incorporate the spatial position as an additional prior for a particular query class. This adaptation enables the network to concentrate more precisely on the exact location of the objects. Meanwhile, the SPC module is built to extract the long-range dependencies and local regions of the spatial position. Additionally, to align the feature space across different modalities, we design an OCA module that condenses textual information into an object query which serves as an instruction for cross-modality matching. Through the collaborative efforts of these three modules, multimodal representations are aligned while maintaining their discriminative nature. Comprehensive experiments conducted on various benchmarks validate the effectiveness of the proposed model.
Wenzhe Zhai, Xianglei Xing, Mingliang Gao 0001, Qilei Li
IEEE Trans. Circuits Syst. Video Technol.1
2025 Generic Representation Learning for Vehicle Association Guided by Foundational Models
abstract
Vehicle association is a vital yet complex task to retrieve specific vehicles across various camera angles, time frames, and geographical locations. In environments supported by autonomous driving and 6G networks, this task plays a vital role in urban surveillance and traffic management by enabling the real-time sharing of vehicle location and status information through ultra-high-speed, low-latency 6G communication. The success of a retrieval model largely depends on the quality of the extracted representations, which can be influenced by factors such as background diversity and occlusions. This study proposes a method to extract representations that remain consistent across different domains while retaining the discriminative power necessary to determine a vehicle’s spatial location, regardless of background or environmental variations. To achieve this, we introduce a framework called Generic Representation Learning (GRL). Within GRL, we leverage large-scale pre-trained foundational models to provide spatial priors of vehicles, specifically the Grounding DINO model for object detection and the SAM model for object segmentation. These modules collaborate to help the network understand the spatial context of the object, enabling the feature extractor to focus on discriminative areas while minimizing interference. Additionally, we introduce a complementary feature alignment mechanism based on a memory bank to explore globally applicable knowledge within the learned representation of the object. These constituent elements collectively form SRP, to enhance its capability for outstanding performance in vehicle retrieval. Extensive experimentation demonstrates that SRP significantly outperforms existing models on widely recognized benchmarks.
Qilei Li, Mingliang Gao 0001, Wenzhe Zhai, Gwanggil Jeon, Ahmed M. Abdelmoniem
IEEE Trans. Intell. Transp. Syst.4
2024 Dual-branch and triple-attention network for pan-sharpening
Mingliang Gao 0001, Abdellah Chehri, Wenzhe Zhai, Qilei Li, Gwanggil Jeon
Appl. Intell.4
2024 Multiscale aggregation and illumination-aware attention network for infrared and visible image fusion
abstract
Abstract Image fusion plays a significant role in computer vision since numerous applications benefit from the fusion results. The existing image fusion methods are incapable of perceiving the most discriminative regions under varying illumination circumstances and thus fail to emphasize the salient targets and ignore the abundant texture details of the infrared and visible images. To address this problem, a multiscale aggregation and illumination‐aware attention network (MAIANet) is proposed for infrared and visible image fusion. Specifically, the MAIANet consists of four modules, namely multiscale feature extraction module, lightweight channel attention module, image reconstruction module, and illumination‐aware module. The multiscale feature extraction module attempts to extract multiscale features in the images. The role of the lightweight channel attention module is to assign different weights to each channel so as to focus on the essential regions in the infrared and visible images. An illumination‐aware module is employed to assess the probability distribution regarding the illumination factor. Meanwhile, an illumination perception loss is formulated by the illumination probabilities to enable the proposed MAIANet to better adjust to the changes in illumination. Experimental results on three datasets, that is, MSRS, TNO, and RoadSence, verify the effectiveness of the MAIANet in both qualitative and quantitative evaluations.
Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Abdellah Chehri, Gwanggil Jeon
Concurr. Comput. Pract. Exp.2
2024 SaReGAN: a salient regional generative adversarial network for visible and infrared image fusion
Mingliang Gao 0001, Yi'nan Zhou, Wenzhe Zhai, Qilei Li
Multim. Tools Appl.3
2024 Multiscale aggregation network via smooth inverse map for crowd counting
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou
Multim. Tools Appl.3
2024 Object counting in remote sensing via selective spatial-frequency pyramid network
abstract
Abstract The integration of remote sensing object counting in the Mobile Edge Computing (MEC) environment is of crucial significance and practical value. However, the presence of significant background interference in remote sensing images poses a challenge to accurate object counting, as the results are easily affected by background noise. Additionally, scale variation within remote sensing images presents a further difficulty, as traditional counting methods face challenges in adapting to objects of different scales. To address these challenges, we propose a selective spatial‐frequency pyramid network (SSFPNet). Specifically, the SSFPNet consists of two core modules, namely the pyramid attention (PA) module and the hybrid feature pyramid (HFP) module. The PA module accurately extracts target regions and eliminates background interference by operating on four parallel branches. This enables more precise object counting. The HFP module is introduced to fuse spatial and frequency domain information, leveraging scale information from different domains for object counting, so as to improve the accuracy and robustness of counting. Experimental results on RSOC, CARPK, and PUCPR+ benchmark datasets demonstrate that the SSFPNet achieves state‐of‐the‐art performance in terms of accuracy and robustness.
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon
Softw. Pract. Exp.4
2024 Defending Deepfakes by Saliency-Aware Attack
abstract
With the rapid development of deep learning, especially the generative adversarial network (GAN), face modification has been substantially advanced and enables the generated images to look more realistic. Given an image or a video frame of a person, such a system can create fake images, which manipulates the movement, expression, and even appearance, e.g., hair color, eye color, and age. Such a system is termed Deepfake, which has raised significant ethical issues, especially for celebrities. With the pretrained Deepfake models being widely available on the Internet, its negative applications, such as face manipulation and pornographic generation, have exposed the dark side of the Deepfake technology to the sociocyber world. In this article, we aim to defend a well-trained Deepfake model by manipulating the raw image with unperceived perturbation. To minimize the alterations to the original image while effectively fooling the Deepfake model, we propose to selectively perturb only the foreground person region and maintain the irrelevant background. This is based on the observation that the salient object in a person’s image is always the foreground face region. Such a strategy introduces negligible alterations to the original image, which makes the attack remain effective. We experimentally demonstrate the superiority of the proposed attacking framework over the existing models and show our approach is ready to be applied for out-of-the-box development.
Qilei Li, Mingliang Gao 0001, Guisheng Zhang, Wenzhe Zhai
IEEE Trans. Comput. Soc. Syst.4
2023 FPANet: feature pyramid attention network for crowd counting
Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon, Marco Anisetti
Appl. Intell.1
2023 Crowd counting in smart city via lightweight Ghost Attention Pyramid Network
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon
Future Gener. Comput. Syst.4
2023 Scale-Context Perceptive Network for Crowd Counting and Localization in Smart City System
abstract
The task of crowd counting and localization is to predict the count and position of people in a crowd, which is a practical and essential sub-task in crowd analysis and smart city systems. However, the inherent problems of scale variation and background disturbance restrain their performance. While recent researches focus on studying counting and localization independently, a few works are capable of executing both tasks simultaneously. To this end, we propose a Scale-Context Perceptive Network (SCPNet) to jointly tackle the crowd counting and localization tasks in a unified framework. Specifically, a scale perceptive (SP) module with a local-global branch schema is designed to capture multiscale information. Meanwhile, a context perceptive (CP) module, by the channel-spatial self-attention mechanism, is derived to suppress the background disturbance. Furthermore, a novel hierarchical scale loss function that combines the Euclidean loss function and structural similarity loss function is designed to prompt the proposed model to fulfill the counting and localization simultaneously. Extensive experiments on challenging crowd datasets prove the superiority of the proposed SCPNet compared with the state-of-the-art competitors in both objective and subjective evaluations.
Wenzhe Zhai, Mingliang Gao 0001, Qilei Li, Gwanggil Jeon
IEEE Internet Things J.1
2023 $\hbox {DA}^2$Net: a dual attention-aware network for robust crowd counting
Wenzhe Zhai, Qilei Li, Jinfeng Pan, Guofeng Zou, Mingliang Gao 0001
Multim. Syst.1
2023 Dense Attention Fusion Network for Object Counting in IoT System
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Kyu Hyung Kim, Gwanggil Jeon
Mob. Networks Appl.3
2023 Scale Region Recognition Network for Object Counting in Intelligent Transportation System
abstract
Self-driving technology and safety monitoring devices in intelligent transportation systems require superb capacity for context awareness. Accurately inferring the counts of crowds and vehicles are the two practical and fundamental tasks in the transportation system. However, the scale variation and background interference in the traffic image hinder the counting performance. To solve the aforementioned problems, a scale region recognition network (SRRNet) is proposed in this paper. It has two key components, termed scale level awareness (SLA) module and object region recognition (ORR) module. The SLA module aims to encode the representations at multiple scales, which are beneficial to address the scale variation. The ORR module is designed to suppress background interference through the visual attention mechanism. Extensive experimental results on four crowd counting datasets and five vehicle counting datasets have demonstrated the superiority of the proposed SRRNet in both counting accuracy and robustness compared with the mainstream competitors. Meanwhile, substantial ablation studies have proved the effectiveness of the proposed SLA and ORS modules.
Mingliang Gao 0001, Wenzhe Zhai, Qilei Li, Gwanggil Jeon
IEEE Trans. Intell. Transp. Syst.3