Zhicong Chen

dblp:156/0704 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-3471-6395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Few-Shot Class-Incremental Continual Learning-Based Fault Diagnosis of Photovoltaic Arrays Using Prompt-Guided Knowledge Distillation and I-V Curves
Haoxin Zheng, Lijun Wu 0002, Zhicong Chen, Shuying Cheng, Peijie Lin, Lin Jiang 0001
IEEE Internet Things J.3
2025 Advancing AI-Enhanced Financial Security: A Review of Facial, Voice, and Medical Biometrics for Identity Verification
abstract
Identity verification is a critical component of financial data and systems security and privacy preservation, and is required by regulatory guidelines to ensure compliance with regulatory requirements and to aid in fraud prevention. With advances in artificial intelligence (AI), deep learning, and statistical methods, financial institutions are increasingly adopting multifactor authentication (MFA) that incorporates biometric-based approaches to perform authentication of identities for access to accounts, records or data and for verification of decision making and confirmation of actions.This paper presents a systematic review of current methodologies that utilize facial recognition, voice biometrics, and other data for identity verification in financial institutions. We explore the effectiveness and challenges associated with these approaches, highlighting recent developments in AI-driven models, deep learning architectures, and statistical techniques. In addition, we discuss the integration of multimodal biometric data and the decision and access systems that are developed for MFA approaches to improve security and accuracy. This review offers insights into the future of biometric identity verification in the financial sector.Our findings suggest that the integration of multi-modal data in financial applications could serve as a valuable avenue for future research and practical applications. In addition, investigating the role of biometric authentication in back-end systems is an important area worth further exploration.
Zhicong Chen, Alice Xiaodan Dong, Gareth W. Peters, Jennifer Chan, Weidong Huang 0001, Chun-Cheng Lin
SMC1
2025 Lightweight deep learning based solar cells defect detection using electroluminescence images
Zhicong Chen, Haoxin Zheng, Lijun Wu 0002, Shuying Cheng, Peijie Lin
Eng. Appl. Artif. Intell.1
2025 Deep-Transfer-Learning-Based Intelligent Gunshot Detection and Firearm Recognition Using Tri-Axial Acceleration
abstract
Reliable identification of gunshot events is crucial for reducing gun violence and enhancing public safety. However, current gunshot detection and recognition methods are still affected by complex shooting scenarios, various nongunshot events, diverse firearm types, and scarce gunshot datasets. To address these issues, based on triaxial acceleration of guns, a novel general deep transfer learning approach is proposed for gunshot detection and recognition, which combines a temporal deep learning model with transfer learning and automated machine learning (AutoML) to improve the accuracy, reliability and generalization performance. First, a new gunshot recognition model named as MobileNetTime is proposed for the two-class gunshot event detection, three-class coarse firearm recognition, and 15-class fine firearm recognition, which utilizes 1-D convolution and inverted residual modules to autonomously extract higher-level features from the time series acceleration data. Second, considering the impact of nongunshot events, the AutoML is employed for model fine tuning, to transfer the pretrained MobileNetTime from the handgun to various firearm types. In addition, we propose a low-power versatile gunshot recognition system framework employing a triaxial accelerometer for both of wrist-worn and gun-embedded scenarios, which adopts a two-stage wake-up mechanism that selectively monitors gunshot events using temporal and spectral energy features. The experimental results on the two gunshot datasets DGUWA and GRD show that the proposed model can achieve up to 100% accuracy on the DGUWA dataset and 98.98% accuracy on the GRD dataset for the two-class gunshot detection. Moreover, the proposed deep transfer learning approach achieves a 98.98% accuracy for 16-class firearm classification, which is 6.21% higher than the model without transfer learning.
Zhicong Chen, Haoxin Zheng, Lijun Wu 0002, Jingchang Huang, Yang Yang 0001
IEEE Internet Things J.1
2024 SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification
abstract
Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly models speaker information — Speaker Related HuBERT, abbreviated as SR-HuBERT. This framework aims to further explore speaker-related information inherent in speech universal representations. The proposed SR-HuBERT utilizes an unsupervised clustering algorithm based on graph structures to generate speaker pseudo-labels and promotes the learning of segment-level speaker-related representations through a multi-task pre-training framework. Experimental results on VoxCeleb1 test set demonstrate the effectiveness of the proposed SR-HuBERT. Even in the scenarios of limited fine-tuning data, SR-HuBERT outperforms the other existing PTMs on SV tasks. Additionally, SR-HuBERT also performs well on speaker-related tasks of SUPERB benchmark.
Yishuang Li, Hukai Huang, Zhicong Chen, Wenhao Guan, Jiayan Lin, Qingyang Hong
ICASSP3
2023 Unsupervised Speaker Verification Using Pre-Trained Model and Label Correction
abstract
Recently, the fine-tuning pre-trained model framework has emerged as a promising paradigm for speech-processing tasks. In this study, we present a novel strategy for unsupervised speaker verification using the Sub-structure of Pre-Trained Model (Sub-PTM), which consists of a CNN-based feature extractor and several Transformer blocks. To obtain the initial pseudo labels, we utilize Infomap to perform clustering on the representations extracted from the Sub-PTM. The generated pseudo labels are then leveraged to train a speaker verification model containing a Sub-PTM and a downstream network. We also propose an Online and Offline Label Correction (OAO-LC) method to alleviate the effects of incorrect pseudo labels. By incorporating these techniques, our system achieves competitive results compared to the supervised baseline.
Zhicong Chen, Wenxuan Hu, Lin Li 0032, Qingyang Hong
ICASSP1
2023 Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization
abstract
The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering approach called Community Detection Graph Convolutional Network (CDGCN) to improve the performance of the speaker diarization system. The CDGCN-based clustering method consists of graph generation, sub-graph detection, and Graph-based Overlapped Speech Detection (Graph-OSD). Firstly, the graph generation refines the local linkages among speech segments. Secondly the sub-graph detection finds the optimal global partition of the speaker graph. Finally, we view speaker clustering for overlap-aware speaker diarization as an overlapped community detection task and design a Graph-OSD component to output overlap-aware labels. By capturing local and global information, the speaker diarization system with CDGCN clustering outperforms the traditional Clustering-based Speaker Diarization (CSD) systems on the DIHARD III corpus.
Zhicong Chen, Haodong Zhou, Qingyang Hong
ICASSP2
2023 CTHPose: An Efficient and Effective CNN-Transformer Hybrid Network for Human Pose Estimation
Danya Chen, Lijun Wu 0002, Zhicong Chen, Xufeng Lin
PRCV (5)3
2023 Teacher-Student Cross-Domain Object Detection Model Combining Style Transfer and Adversarial Learning
Lijun Wu 0002, Zhicong Chen
PRCV (10)3
2023 A Wireless Gunshot Recognition System Based on Tri-Axis Accelerometer and Lightweight Deep Learning
abstract
Gun violence and misuse pose great threat to the public safety. Real-time monitoring of gun usage and gunshot events are very promising for effective gun control. However, most available monitoring systems are installed in a fixed location instead of the guns, which greatly limits the flexibility and coverage. In this study, we propose a wireless gun monitoring and gunshot recognition system based on a low-cost triaxial acceleration sensor, which can monitor the gun in real time and accurately recognize gunshot events. Addressing the limited resources of the embedded systems, we further propose an efficient gunshot recognition algorithm EfficientNetTime that combines the lightweight neural network and knowledge distillation, so as to enable the deployment on embedded devices. First, a novel lightweight deep learning model is proposed as the basic model, which combines the advantages of 1-D convolution and depthwise separable convolution to effectively characterize the gunshot signal while decreasing the computing cost of convolution. Second, using the knowledge distillation, EfficientNetTime is used as the teacher model to generate a compressed student model that maintains accuracy and greatly reducing model size. Finally, the EfficientNetTime student model can be deployed on resource-limited embedded systems. The proposed method can automatically extract features for end-to-end recognition and is robust to temporal transformations of input signals. Using a publicly available gunshot data set, the proposed EfficientNetTime model is verified and compared against the state-of-the-art models. Experimental results demonstrate that the EfficientNetTime model surpasses other gunshot recognition methods in terms of the accuracy and model size.
Zhicong Chen, Haoxin Zheng, Jingchang Huang, Lijun Wu 0002, Shuying Cheng, Qianwei Zhou, Yang Yang 0001
IEEE Internet Things J.1
2021 An Efficient Binary Convolutional Neural Network With Numerous Skip Connections for Fog Computing
abstract
Fog computing is promising to solve the challenge caused by an extremely large amount of data on cloud computing. In this study, an efficient binary convolutional neural network with numerous skip connections (BNSC-Net) is proposed for fog computing to enable real-time smart industrial applications. This network features decomposition convolution kernels and concatenated feature maps. Moreover, the network performance is further improved through expanding the update interval of the straight-through estimator. To verify the performance, BNSC-Net is tested on two broadly used public data sets: 1) ImageNet and 2) CIFAR-10. An ablation study is first conducted to verify the effectiveness of the proposed improved operations, and results demonstrate that BNSC-Net can obviously increase the classification accuracy for both data sets. ImageNet-based classification results indicate that BNSC-Net can achieve 59.9% TOP-1 accuracy that is 2.6% higher than the state-of-the-art binary neural networks, such as projection convolutional neural networks (PCNNs). Finally, a subset with ten classes is selected from ImageNet to simulate the data collected in the smart industry with limited categories, based on which BNSC-Net also demonstrates an impressive classification performance with friendly memory and calculation requirements. Particularly, the receiver operating characteristic curves of BNSC-Net surpass that of the state-of-the-art algorithm DeepIns. Therefore, the proposed BNSC-Net is effective and efficient for building deep learning-enabled industrial applications on fog nodes.
Lijun Wu 0002, Zhicong Chen, Jingchang Huang, Yang Yang 0001
IEEE Internet Things J.3
2021 Improved Magnetic Guidance Approach for Automated Guided Vehicles by Error Analysis and Prior Knowledge
abstract
Navigation accuracy and robustness are key performance indexes of automated guided vehicles (AGVs). In our previous study, the magnetic guidance approach based on magnetic dipole model and non-linear optimization algorithm was proposed, which has high positioning accuracy and could estimate the yaw angle of AGV directly. However, the localization accuracy of the magnetic guidance approach will deteriorate if the magnetic nails (MNs) buried in the ground have installation errors or the magnetic moments between the MNs are inconsistent. To overcome this problem, we propose an improved method based on error analysis and prior knowledge for the magnetic guidance approach. Firstly, the factors that affect the localization accuracy are analyzed, and the parameters ($B_{\mathrm {T}}$,$p$,$c$), whose errors will deteriorate the localization accuracy, are combined with MN pose ($a$,$b$,$m$,$n$) as optimization variables. Then, the prior knowledge regarding ($B_{\mathbf {T}}$,$p$,$c$) is employed to construct the constraint conditions for the magnetic guidance approach. Finally, the global convergence probability and convergence speed of the improved magnetic guidance approach are analyzed. Experimental results demonstrate the adaptability and robustness of the improved magnetic tracking approach, which diminishes the impact of MN installation errors and magnetic moment deviation. The parking accuracy of AGV is improved to 1.42±0.85 mm and 1.10±0.38°.
Shijian Su, Houde Dai, Shuying Cheng, Zhicong Chen
IEEE Trans. Intell. Transp. Syst.4
2015 Orderless and Blurred Visual Tracking via Spatio-temporal Context
Manna Dai, Peijie Lin, Lijun Wu 0002, Zhicong Chen, Songlin Lai, Shuying Cheng, Xiangjian He
MMM (1)4