Zhidong Zhao

dblp:47/10426 · DBLP profile ↗
← Back
37ranked-venue papers
0as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-annotation agreement and prediction consistency networks: Improving semi-supervised segmentation of medical images with ambiguous boundaries
Shuai Wang 0003, Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Yixiu Liu, Pengfei Jiao, Zhiming Cheng, Yaoqi Sun, Yaqi Wang 0002
Artif. Intell. Medicine6
2026 Open-set electrocardiogram identity authentication via multi-modal pretraining and structured representation learning
Mingyu Dong, Zhidong Zhao, Yefei Zhang, Yanjun Deng, Xianfei Zhang
Eng. Appl. Artif. Intell.2
2026 Expert consensus-driven spatial-temporal graph neural network for enhanced diagnosis of chronic fetal distress
Yefei Zhang, Yanjun Deng, Bingxin Ruan, Zhidong Zhao
Eng. Appl. Artif. Intell.5
2026 Dual-gated multi-behavior recommendation with variational graph autoencoder
Nana Huang, Pengfei Jiao, Zhidong Zhao, Chao Liang 0001, Fengjun Xiao
Expert Syst. Appl.4
2026 Secrecy Rate Optimization Based on GNN for RIS-Assisted ISAC System
abstract
Integrated Sensing and Communication (ISAC) systems are playing an increasingly crucial role in modern wireless networks. However, in the ISAC scenario, the high transmission power required for communicating and sensing signals poses an increased risk of signal interception by eavesdroppers. To address this issue and enhance the physical layer security (PLS) of ISAC, we utilize Reconfigurable Intelligent Surfaces (RIS) to optimize the secrecy rate in the ISAC context, improving link security and effectively preventing eavesdropping. In the ISAC scenario, the location of the eavesdropper can be obtained through sensing. Leveraging this advantage, we employ Graph Neural Networks (GNN) to aggregate the node information of users and eavesdroppers, which iteratively passes messages and updates node states, thereby adaptively optimizing the transmitting beamforming vector and the RIS phase shift matrix. Simulation results show that this method is superior to the benchmark algorithm.
Jieling Zhang, Huijun Tang, Pengfei Jiao, Huaming Wu, Zhidong Zhao, Ruidong Li 0001
IEEE Internet Things J.5
2026 Hypergraph-enhanced cascades prediction via text-informed learning
Fangfang Su, Yingliang Wu, Dongdong Xie 0003, Zhidong Zhao, Pengfei Jiao
Knowl. Based Syst.4
2026 Mining user features with hyperbolic representations for diffusion prediction
Pengfei Jiao, Zhidong Zhao, Fangfang Su, Wang Zhang 0001
Neural Networks4
2026 MEDiT: A mask-enhanced diffusion transformer model for fetal heart rate signal generation
Yefei Zhang, Pengfei Jiao, Yanjun Deng, Yehui Chen, Zhidong Zhao
Neural Networks6
2026 TGFormer: Towards temporal graph transformer with auto-correlation mechanism
Hongjiang Chen 0001, Pengfei Jiao, Ming Du 0003, Xuan Guo 0005, Zhidong Zhao, Di Jin 0001
Pattern Recognit.5
2026 SGM-Net: 3D Point Cloud Class-Incremental Segmentation via Semantic-Aware Global Modeling
Jinshuo Liu, Bingtao Ma, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
IEEE Signal Process. Lett.3
2026 Unified Network Embedding via Mutual Fusion of Communities and Roles
abstract
Most network embedding (NE) methods are either based on the proximity for community-guided tasks or on the structural similarity for role-oriented tasks. While being prevalent and effective, there still exists some potential issues that need further attention: 1) community and role are always regarded as orthogonal problems. They have rarely been combined to model the latent structures within the complex network. However, the generation of the network is usually jointly driven by these two mechanisms; and 2) few works study the interaction between roles or communities, which leads to the generation process of the network cannot being effectively modeled. To solve these problems, we propose a unified network embedding framework via mutual fusion of community and role (UMFCR). We combine the Gaussian mixture model (GMM) with a variational graph auto-encoder to generate node embeddings and discover the membership distribution of each node. An elaborate fusion pattern is then designed to produce the generation process for each link from the perspective of both community and role. The promising experimental results on real-world data demonstrate the necessity of fusing these two mechanisms and the superior performance of the model on different network tasks.
Pengfei Jiao, Wang Zhang 0001, Xuan Guo 0005, Huan Liu 0001, Yanxian Bi, Yefei Zhang, Zhidong Zhao
IEEE Trans. Comput. Soc. Syst.7
2025 RIS-Assisted Beamforming Optimization Based on DNN in the UAV-ISAC System
Jieling Zhang, Bin Yang 0034, Pinlong Zhao, Pengfei Jiao, Zhidong Zhao, Huaming Wu
ICA3PP (5)6
2025 Knowledge-Guided Prompt Learning for Deepfake Facial Image Detection
abstract
Recent generative models demonstrate impressive performance on synthesizing photographic images, which makes humans hardly to distinguish them from pristine ones, especially on realistic-looking synthetic facial images. Previous works mostly focus on mining discriminative artifacts from vast amount of visual data. However, they usually lack the exploration of prior knowledge and rarely pay attention to the domain shift between training categories (e.g., natural and indoor objects) and testing ones (e.g., fine-grained human facial images), resulting in unsatisfactory detection performance. To address these issues, we propose a novel knowledge-guided prompt learning method for deepfake facial image detection. Specifically, we retrieve forgery-related prompts from large language models as expert knowledge to guide the optimization of learnable prompts. Besides, we elaborate test-time prompt tuning to alleviate the domain shift, achieving significant performance improvement and boosting the application in real-world scenarios. Extensive experiments on DeepFakeFaceForensics dataset show that our proposed approach notably outperforms state-of-the-art methods.
Hao Wang 0062, Cheng Deng 0002, Zhidong Zhao
ICASSP3
2025 Heartid: A Deep Learning Model for Biometrics Using Multimodal ECG and PPG Signals
abstract
Biometrics is one of the most prominent sources of knowledge for human identification in the rapidly advancing field of cybersecurity. Among various types of biometric information, electrocardiogram (ECG) and photoplethysmography (PPG) signals gained considerable attention due to its inherent resistance to spoofing. However, existing researches are predominantly conducted in controlled environments, where ECG or PPG signals are susceptible to interference from motion and noise in real-world scenarios, resulting in suboptimal model robustness. Additionally, the few multimodal ECG identification studies mainly based on different conversion methods of ECG signal, resulting in information redundancy between modalities. To address these, we propose a lightweight identity recognition model based on cross-domain multimodal feature fusion of ECG and PPG signals. It incorporates an enhanced MultiRes block that forwards upsampled data for multi-resolution analysis of signal features, and employs spatial pyramid pooling for multi-scale feature extraction. Using this network structure, we separately extract and fuse features from ECG and PPG signals, then perform identity recognition based on the fused features. Extensive experiments were conducted to evaluate the proposed algorithm. The proposed fusion model reaches a high accuracy of 97.26%, improving by 17.06% and 40.48% compared to single-modal ECG and PPG identification models, respectively. By visualizing the fused features alongside single-modal ECG and PPG features, we intuitively explain the effectiveness of multimodal feature fusion for feature extraction, offering a new approach for multimodal biometric recognition.
Bingxin Ruan, Zhidong Zhao, Yefei Zhang, Yanjun Deng, Pengfei Jiao, Xianfei Zhang
ICC2
2025 A Survey on Temporal Interaction Graph Representation Learning: Progress, Challenges, and Opportunities
abstract
Temporal interaction graphs (TIGs), defined by sequences of timestamped interaction events, have become ubiquitous in real-world applications due to their capability to model complex dynamic system behaviors. As a result, temporal interaction graph representation learning (TIGRL) has garnered significant attention in recent years. TIGRL aims to embed nodes in TIGs into low-dimensional representations that effectively preserve both structural and temporal information, thereby enhancing the performance of downstream tasks such as classification, prediction, and clustering within constantly evolving data environments. In this paper, we begin by introducing the foundational concepts of TIGs and emphasizing the critical role of temporal dependencies. We then propose a comprehensive taxonomy of state-of-the-art TIGRL methods, systematically categorizing them based on the types of information utilized during the learning process to address the unique challenges inherent to TIGs. To facilitate further research and practical applications, we curate the source of datasets and benchmarks, providing valuable resources for empirical investigations. Finally, we examine key open challenges and explore promising research directions in TIGRL, laying the groundwork for future advancements that have the potential to shape the evolution of this field.
Pengfei Jiao, Hongjiang Chen 0001, Xuan Guo 0005, Zhidong Zhao, Dongxiao He, Di Jin 0001
IJCAI4
2025 FedCCH: Automatic Personalized Graph Federated Learning for Inter-Client and Intra-Client Heterogeneity
abstract
Graph federated learning (GFL) is increasingly utilized in domains such as social network analysis and recommendation systems, where non-IID data exist extensively and necessitate a strong emphasis on personalized learning. However, existing methods focus only on the personality among different clients instead of the personality within a client which widely exists in the real social networks, where intra-client personality addresses the heterogeneity of known data, while inter-client personality always tackle client heterogeneity under privacy constraint. In this paper, we propose a novel automatic personalized graph federated learning (PGFL) scheme named FedCCH to capture both inter-client and intra-client heterogeneity. For intra-client heterogeneity, we innovatively propose the learnable Personalized Factor (PF) to automatically normalize each graph representation within clients by learnable parameters, which weakens the impact of non-IID data distribution. For inter-client heterogeneity, we propose a novel hash-based similarity clustering method to generate the hash signature for each client, and then group similar clients for joint training among different clients. Ultimately, we collaboratively train intra-client and inter-client modules to improve the effectiveness of capturing the heterogeneity of the graph data of clients. Experiment results demonstrate that FedCCH outperforms other state-of-the-art baseline methods.
Pengfei Jiao, Zian Zhou, Meiting Xue, Huijun Tang, Zhidong Zhao, Huaming Wu
IJCAI5
2025 HyperRole: Hyperbolic Graph Transformer for Role Discovery in Online Social Networks
abstract
Role discovery assist in various applications of online social networks, such as water army detection, shopping recommendation, rumor tracing, etc. However, existing studies often overlook the significance of hierarchical structures in online social networks, which are crucial for understanding the roles played by different users. To address this gap, we propose a novel approach based on hyperbolic graph learning, called HyperRole, which effectively leverages the hierarchical structure of online social networks for role discovery. HyperRole first extracts structural features from users and constructs user sequences based on feature similarity, capturing the relationships between users across different scales. Then, we learn role information from structural features by hyperbolic graph Transformer to embed users into the hyperbolic space, preserving the hierarchical structure between users and enabling interactions between users of the same level that are far away from each other. Additionally, we leverage the hierarchical distance between the target user and other users within the same sequence to guide and modify the role information of the target user. Based on the generated user role embeddings, we train a multi-class classifier to classify roles. Extensive experiments on several real-world network datasets demonstrate that our model outperforms existing baseline methods, showcasing its superior performance.
Huijun Tang, Ming Du 0003, Pengfei Jiao, Huaming Wu, Zhidong Zhao
INFOCOM5
2025 A time-series progressive generative adversarial network for improving imbalanced fetal heart rate signal classification
Yanjun Deng, Yefei Zhang, Hao Wang 0062, Pengfei Jiao, Zhidong Zhao
Appl. Intell.6
2025 A two-stage model for unified sentence- and document-level biomedical event extraction
abstract
Biomedical event extraction, a cornerstone of information extraction, has increasingly attracted attention within the biomedical research community. Moreover, it is a highly complex task, which not only deals with many sub-tasks but also involves nested events. Currently, the research on biomedical event extraction, whether pipelined model or joint method, needs to be processed for each sub-task. The process of processing each sub-task one by one lead to the degradation of event extraction performance. In addition, most studies focus on extracting sentence-level events and ignore cross-sentence event information. To solve these problems, we simplify the process of event extraction, reduce the processing steps, and combine the two sub-tasks of relation extraction and argument combination as one sub-task. In addition, we consider document-level event extraction, which not only extracts cross-sentence events but also considers broader context information. Experimental results indicate that our novel approach outperforms prior studies. Additionally, the document-level event extraction model attains the top performance on the BioNLP’11 test data and achieves near-leading performance on the BioNLP’13 test data.
Fangfang Su, Yue Zhang 0004, Pengfei Jiao, Zhidong Zhao, Bobo Li 0001, Fei Li 0021, Donghong Ji
Eng. Appl. Artif. Intell.4
2025 Domain generalization for image classification with dynamic decision boundary
Zhiming Cheng, Mingxia Liu 0001, Defu Yang, Zhidong Zhao, Chenggang Yan 0001, Shuai Wang 0003
Pattern Recognit.4
2025 Zero-Shot Interpretable Image Steganalysis for Invertible Image Hiding
abstract
Image steganalysis, which aims at detecting secret information concealed within images, has become a critical coun termeasure for assessing the security of steganography methods, especially the emerging invertible image hiding approaches. However, prior studies merely classify input images into two categories (i.e., stego or cover) and typically conduct steganalysis under the constraint that training and testing data must follow similar distribution, thereby hindering their application in real-world scenarios. To overcome these shortcomings, we propose a novel interpretable image steganalysis framework tailored for invertible image hiding schemes under a challenging zero-shot setting. Specifically, we integrate image hiding, revealing, and steganalysis into a unified framework, endowing the steganalysis component with the capability to recover the secret information embedded in stego images. Additionally, we elaborate a simple yet effective residual augmentation strategy for generating stego images to further enhance the generalizability of the steganalyzer in cross dataset and cross-architecture scenarios. Extensive experiments on benchmark datasets demonstrate that our proposed approach significantly outperforms the existing steganalysis techniques for invertible image hiding schemes.
Hao Wang 0062, Yaguang Xie, Zhidong Zhao
IEEE Signal Process. Lett.5
2025 Multi-time-scale with clockwork recurrent neural network modeling for sequential recommendation
Nana Huang, Ruimin Hu, Pengfei Jiao, Zhidong Zhao, Bin Yang 0034
J. Supercomput.5
2025 DVGMAE: Self-Supervised Dynamic Variational Graph Masked Autoencoder
abstract
Although contrastive self-supervised learning (SSL) on dynamic graphs has made significant success, the issue of heavy reliance on data augmentation and training tricks has been a persistent pain point. Generative SSL, especially masked autoencoders (MAEs) have recently produced promising results and can avoid these issues. However, the research on MAE in dynamic graphs remains largely unexplored due to the following challenges: 1) how to design an effective masking strategy for dynamic graphs? and 2) how to design a decoder to retain temporal dependency when graphs are perturbed? In this article, we propose DVGMAE, a novel dynamic variational graph masked autoencoder model to solve these challenges. DVGMAE simultaneously captures the evolving behaviors and topological features via an innovative masking strategy and an elaborate decoder. Specifically, we first implement a temporal-aware masking strategy on the edges of each snapshot based on the updated probabilities derived from historical mask information. This strategy mitigates potential masking bias in dynamic graphs. We then design a globally enhanced decoder to recover the temporal and spatial information of each snapshot. Extensive experiments demonstrate that DVGMAE outperforms the existing state-of-the-art on various tasks across different datasets.
Mengzhou Gao 0001, Xinxun Zhang, Pengfei Jiao, Tianpeng Li, Zhidong Zhao
IEEE Trans. Neural Networks Learn. Syst.5
2025 Interactive Graph Learning for Multilevel Network Alignment
abstract
The task of network alignment aims to identify corresponding nodes across multiple networks, with applications in various fields such as social network analysis and bioinformatics. Traditional methods typically focus on the topological structure of networks at a specific level, but they may overlook important properties exhibited by many networks, such as scale-free properties and specific power-law structures often found in social networks. Consequently, these methods fail to effectively capture and utilize such information, leading to misalignment. In this article, we propose a network alignment framework that incorporates both topological and attribute information from multiple levels in the network, including homogeneity, power-law, and higher order structures. We introduce a Euclidean hyperbolic interactive graph learning method specifically designed for modeling power-law structures in networks, aiming to improve the accuracy of network alignment. To evaluate the effectiveness of our proposed method, we conduct experiments on several real-world datasets. The results demonstrate that our approach achieves higher accuracy compared to other advanced baselines.
Pengfei Jiao, Yuanqi Liu, Yinghui Wang 0005, Huijun Tang, Zhidong Zhao, Shirui Pan
IEEE Trans. Neural Networks Learn. Syst.5
2025 Joint Optimization Based on Two-Phase GNN in RIS- and DF-Assisted MISO Systems With Fine-Grained Rate Demands
abstract
Reconfigurable intelligent Surfaces (RIS) and half-duplex decoded and forwarded (DF) relays can collaborate to optimize wireless signal propagation in communication systems. Users typically have different rate demands and are clustered into groups in practice based on their requirements, where the former results in the trade-off between maximizing the rate and satisfying fine-grained rate demands, while the latter causes a trade-off between inter-group competition and intra-group cooperation when maximizing the sum rate. However, traditional approaches often overlook the joint optimization encompassing both of these trade-offs, disregarding potential optimal solutions and leaving some users even consistently at low date rates. To address this issue, we propose a novel joint optimization model for a RIS- and DF-assisted multiple-input single-output (MISO) system where a base station (BS) is with multiple antennas transmits data by multiple RISs and DF relays to serve grouped users with fine-grained rate demands. We design a new loss function to not only optimize the sum rate of all groups but also adjust the satisfaction ratio of fine-grained rate demands by modifying the penalty parameter. We further propose a two-phase graph neural network (GNN) based approach that inputs channel state information (CSI) to simultaneously and autonomously learn efficient phase shifts, beamforming, and relay selection. The experimental results demonstrate that the proposed method significantly improves system performance.
Huijun Tang, Jieling Zhang, Zhidong Zhao, Huaming Wu, Hongjian Sun 0001, Pengfei Jiao
IEEE Trans. Wirel. Commun.3
2024 Simple yet Effective ECG Identity Authentication with Low EER & without Retraining
abstract
Recently, electrocardiogram (ECG) signals have garnered significant attention in the field of identity authentication due to its biological uniqueness. For identity authentication, the ECG signals that collected by wearing smart devices need to be determined whether the signal belongs to an enrolled one. Constrained by the computational efficiency of smart devices in practical scenarios, it is essential to reduce the complexity of the method to lower the computational load. To maintain the accuracy of identity authentication, most research efforts rely on both R-wave extraction and segmentation for subsequent authentication. Moreover, many methods constantly require the model to be retrained during the user enrollment stage, leading to performance degradation and waste of training resources. Hence, we propose a simple yet effective ECG identity authentication method that applies blind segmentation and is free from retraining, which greatly simplifies the authentication process. To mitigate the Equal Error Rate (EER) during the verification phase, a combination of AAM-softmax and triplet losses is employed, along with the incorporation of the hard negative mining within batch samples. Extensive experiments demonstrate that our method outperforms competitors by a large margin, e.g., achieving 0.40% EER on the large-scale Autonomic dataset.
Mingyu Dong, Zhidong Zhao, Yefei Zhang, Yanjun Deng, Hao Wang 0062, Bingxin Ruan
BIBM2
2024 Accurate Segmentation of Optic Disc and Cup from Multiple Pseudo-labels by Noise-aware Learning
abstract
Optic disc and cup segmentation plays a crucial role in automating the screening and diagnosis of optic glaucoma. While data-driven convolutional neural networks (CNNs) show promise in this area, the inherent ambiguity of segmenting objects and background boundaries in the task of optic disc and cup segmentation leads to noisy annotations that impact model performance. To address this, we propose an innovative label-denoising method of Multiple Pseudo-labels Noise-aware Network (MPNN) for accurate optic disc and cup segmentation. Specifically, the Multiple Pseudo-labels Generation and Guided Denoising (MPGGD) module generates pseudo-labels by multiple different initialization networks trained on true labels, and the pixel-level consensus information extracted from these pseudo-labels guides to differentiate clean pixels from noisy pixels. The training framework of the MPNN is constructed by a teacher-student architecture to learn segmentation from clean pixels and noisy pixels. Particularly, such a framework adeptly leverages (i) reliable and fundamental insight from clean pixels and (ii) the supplementary knowledge within noisy pixels via multiple perturbation-based unsupervised consistency. Compared to other label-denoising methods, comprehensive experimental results on the RIGA dataset demonstrate our method’s excellent performance. The code is available at https://github.com/wwwtttjjj/MPNN.
Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Zhiming Cheng, Shuai Wang 0003
CSCWD3
2024 Informative Subgraphs Aware Masked Auto-Encoder in Dynamic Graphs
abstract
Generative self-supervised learning (SSL), especially masked autoencoders (MAE), has greatly succeeded and garnered substantial research interest in graph machine learning. However, the research of MAE in dynamic graphs is still scant. This gap is primarily due to the dynamic graph not only possessing topological structure information but also encapsulating temporal evolution dependency. Applying a random masking strategy which most MAE methods adopt to dynamic graphs will remove the crucial subgraph that guides the evolution of dynamic graphs, resulting in the loss of crucial spatio-temporal information in node representations. To bridge this gap, in this paper, we propose a novel Informative Subgraphs Aware Masked Auto-Encoder in Dynamic Graph, namely DyGIS. Specifically, we introduce a constrained probabilistic generative model to generate informative subgraphs that guide the evolution of dynamic graphs, successfully alleviating the issue of missing dynamic evolution subgraphs. The informative subgraph identified by DyGIS will serve as the input of dynamic graph masked autoencoder (DGMAE), effectively ensuring the integrity of the evolutionary spatio-temporal information within dynamic graphs. Extensive experiments on eleven datasets demonstrate that DyGIS achieves state-of-the-art performance across multiple tasks.
Pengfei Jiao, Xinxun Zhang, Mengzhou Gao 0001, Tianpeng Li, Zhidong Zhao
ICDM5
2024 Contrastive representation learning on dynamic networks
Pengfei Jiao, Hongjiang Chen 0001, Huijun Tang, Qing Bao, Zhidong Zhao, Huaming Wu
Neural Networks6
2024 Inductive Link Prediction via Interactive Learning Across Relations in Multiplex Networks
abstract
Network embedding is an important class of link prediction methods, which can use the distance between learned low-dimensional node representations to characterize the similarity between nodes. Traditional network embedding methods focus on single-layer networks, while in reality, a large part of complex networks are not isolated, but interdependent and interrelated, forming multiplex complex networks. Also, how to effectively exploit layer correlations in multiplex networks to learn more robust and valuable representations, to improve link prediction performance, has been a hot research topic in the field of complex network analysis. However, previous studies mainly focus on inferring intralinks in each layer of complex networks or anchor links among layers. Another issue that has not been discussed is how to predict potential links or reconstruct the network in unobserved relations based on existing multiplex networks. To this issue, we define a novel inductive link prediction problem in multiplex networks, in which most existing multichannel network embedding methods fail to solve. This is either because they only emphasize the specific structure information of an individual layer or only capture the common information for all layers. To effectively address this problem, we propose a novel embedding method termed interactive learning across relations (ILAR), to capture and fully exploit the multiple relations and complex layer correlations in multiplex networks. We leverage two convolutional modules and ILAR to capture the sufficient complementary and correlations in multiplex networks. Moreover, during interactive learning, a disparity constraint is introduced, which enforces the features encoded from two convolutional modules to be different and prevents information redundancy. Finally, the extensive experiments in several real-world datasets show that our model can significantly outperform the existing state-of-the-art network embedding methods on the novel link prediction problem in multiplex networks.
Mengzhou Gao 0001, Pengfei Jiao, Ruili Lu, Huaming Wu, Yinghui Wang 0005, Zhidong Zhao
IEEE Trans. Comput. Soc. Syst.6
2024 VGGM: Variational Graph Gaussian Mixture Model for Unsupervised Change Point Detection in Dynamic Networks
abstract
Change point detection in dynamic networks aims to detect the points of sudden change or abnormal events within the network. It has garnered substantial interest from researchers due to its potential to enhance the stability and reliability of real-world networks. Most change point detection methods are based on statistical characteristics and phased training, and some methods are required to set the percent of change points. Meanwhile, existing methods for change point detection suffer from two limitations. On one hand, they struggle to extract snapshot features that are crucial for accurate change point detection, thereby limiting their overall effectiveness. On the other hand, they are typically tailored for specific network types and lack the versatility to adapt to networks of varying scales. To solve these issues, we propose a novel unified end-to-end framework called Variational Graph Gaussian Mixture model (VGGM) for change point detection in dynamic networks. Specifically, VGGM combines Variational Graph Auto-Encoder (VGAE) and Gaussian Mixture Model (GMM) through joint training, incorporating a Mixture-of-Gaussians prior to model dynamic networks. This approach yields highly effective snapshot embeddings via VGAE and a dedicated readout function, while automating change point detection through GMM. The experimental results, conducted on both real-world and synthetic datasets, clearly demonstrate the superiority of our model in comparison to the current state-of-the-art methods for change point detection.
Xinxun Zhang, Pengfei Jiao, Mengzhou Gao 0001, Tianpeng Li, Yiming Wu 0001, Huaming Wu, Zhidong Zhao
IEEE Trans. Inf. Forensics Secur.7
2023 On Multi-Modal Fusion Learning in Pathological Diagnosis of Fetal Distress
abstract
Cardiotocography (CTG) is an important medical diagnostic tool when it comes to monitoring fetal wellbeing. It records Fetal Heart Rate (FHR) and uterine contraction activity, and can be used to detect whether the fetus is receiving oxygen adequately or in distress. Unfortunately, the interpretation of CTG recordings is highly subjective which can lead to unnecessary medical intervention that represents a risk for both the mother and the fetus. In this regard, intelligent CTG (ICTG) classification is a challenging research that can assist obstetricians in making clinical decisions, thereby improving the efficiency and accuracy of pregnancy management. But, many of these models focus on one specific modality that lack generalization to unseen or test data samples. In this study, a multi-modal fusion learning approach is proposed for pathological diagnosis of fetal distress. It combines signal and image modalities for multi-modal inputs and develops a Multi-modal Encoder Network (MENet) model based on DNN for capturing the underlying distribution of multi-modal data samples. Experimental results demonstrate that under the constraints of same classifier structure, MENet performs well in terms of classification accuracy and stability, far superior to several existing ICTG algorithms.
Yefei Zhang, Zhidong Zhao, Yanjun Deng, Pengfei Jiao
HealthCom2
2022 FHRGAN: Generative adversarial networks for synthetic fetal heart rate signal generation in low-resource settings
Yefei Zhang, Zhidong Zhao, Yanjun Deng, Xiaohong Zhang 0005
Inf. Sci.2
2022 Reconstruction of Missing Samples in Antepartum and Intrapartum FHR Measurements Via Mini-Batch-Based Minimized Sparse Dictionary Learning
abstract
Fetal Heart Rate (FHR), an important recording in Cardiotocography (CTG)-based fetal health status monitoring, is the only information that clinical obstetricians can directly obtain and use. A challenge, however, is that missing samples are very common in FHR due to various causes such as fetal movements and sensor malfunctions. The aim is the development of an inpainting tool which is suitable for different missing lengths$q$and various total missing percentages$Q$, as well as for use in online mode. This study focused on two major impediments to existing inpainting methods: the longer the missing length, the more difficult it is to recover with mathematical methods; the reliance on tens of thousands of training samples, and the computational burden caused by full batch-based dictionary learning algorithms. We present a regularized minimization approach to signal recovery, which combines a${{\rm{L}}_{{0}{\rm{.6}}}}{\rm{ - norm}}$minimized sparse dictionary learning algorithm (MSDL) and a model optimization strategy for using a mini-batch version for signal recovery. Using 100 FHR recordings with 2 protocols designed to simulate missing clinical data scenarios, the combined method performed favorably in terms of 5 data analysis metrics and 3 clinical indicators. Comparing 4 inpainting methods, we were able to prove the superiority of the proposed algorithm for both large$q$and large$Q$. The experimental results showed the lowest values (2.64 (MAE), 4.68 (RMSE)) when${\rm{Q}} = {\rm{5\% }}$with short interval lengths. The developed architecture provides a reference value for the practical application of recovering missing samples online.
Yefei Zhang, Zhidong Zhao, Yanjun Deng, Xiaohong Zhang 0005, Yu Zhang 0077
IEEE J. Biomed. Health Informatics2
2021 ECGID: a human identification method based on adaptive particle swarm optimization and the bidirectional LSTM model
abstract
Physiological signal based biometric analysis has recently attracted attention as a means of meeting increasing privacy and security requirements. The real-time nature of an electrocardiogram (ECG) and the hidden nature of the information make it highly resistant to attacks. This paper focuses on three major bottlenecks of existing deep learning driven approaches: the lengthy time requirements for optimizing the hyperparameters, the slow and computationally intense identification process, and the unstable and complicated nature of ECG acquisition. We present a novel deep neural network framework for learning human identification feature representations directly from ECG time series. The proposed framework integrates deep bidirectional long short-term memory (BLSTM) and adaptive particle swarm optimization (APSO). The overall approach not only avoids the inefficient and experience-dependent search for hyperparameters, but also fully exploits the spatial information of ordinal local features and the memory characteristics of a recognition algorithm. The effectiveness of the proposed approach is thoroughly evaluated in two ECG datasets, using two protocols, simulating the influence of electrode placement and acquisition sessions in identification. Comparing four recurrent neural network structures and four classical machine learning and deep learning algorithms, we prove the superiority of the proposed algorithm in minimizing overfitting and self-learning of time series. The experimental results demonstrated an average identification rate of 97.71%, 99.41%, and 98.89% in training, validation, and test sets, respectively. Thus, this study proves that the application of APSO and LSTM techniques to biometric human identification can achieve a lower algorithm engineering effort and higher capacity for generalization.
Yefei Zhang, Zhidong Zhao, Yanjun Deng, Xiaohong Zhang 0005, Yu Zhang 0077
Frontiers Inf. Technol. Electron. Eng.2
2021 Heart biometrics based on ECG signal by sparse coding and bidirectional long short-term memory
Yefei Zhang, Zhidong Zhao, Yanjun Deng, Xiaohong Zhang 0005, Yu Zhang 0077
Multim. Tools Appl.2
2020 Simulating study on RHCRP protocol in utility tunnel WSN
Zhixin Zhou, Chenning Shao, Huimin Zhou, Xiong-wei Lou, Jian Li 0050, Guohua Hui, Zhidong Zhao
Wirel. Networks9