VLDB 2026 Research / reviewers in the wild / expert
Zhifu Zhao
dblp:207/7535
· DBLP profile ↗
18ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-4136-6362ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LGCM: Referring image segmentation with language-guided channel modulation
Chengyuan Chang, Fei Qi 0001, Zhifu Zhao |
Neurocomputing | 5 |
| 2026 | NVFusion: Lightweight infrared and low-light night vision image fusion in dark environments
Fu Li 0002, Zhifu Zhao |
Neurocomputing | 4 |
| 2026 | VGRF Signal-Based Gait Analysis for Parkinson's Disease Detection: A Multi-Scale Directed Graph Neural Network ApproachabstractParkinson's Disease (PD) is often characterized by abnormal gait patterns, which can be objectively and quantitatively diagnosed using Vertical Ground Reaction Force (VGRF) signals. Previous studies have demonstrated the effectiveness of deep learning in VGRF signal analysis. However, the inherent graph structure of VGRF signals has not been adequately considered, limiting the representation of dynamic gait characteristics. To address this, we propose a Multi-Scale Adaptive Directed Graph Neural Network (MS-ADGNN) approach to distinguish the gaits between Parkinson's patients and healthy controls. This method models the VGRF signal as a multi-scale directed graph, capturing the distribution relationships within the plantar sensors and the dynamic pressure conduction during walking. MS-ADGNN integrates an Adaptive Directed Graph Network (ADGN) unit and a Multi-Scale Temporal Convolutional Network (MSTCN) unit. ADGN extracts spatial features from three scales of the directed graph, effectively capturing local and global connectivity. MSTCN extracts multi-scale temporal features, capturing short to long-term dependencies. The proposed method outperforms existing methods on three widely used datasets. In cross-dataset experiments, the average improvements in terms of accuracy, F1-score, and geometric mean are 2.46$\%$, 1.25$\%$, and 1.11$\%$ respectively. Meanwhile, in 10-fold cross-validation experiments, the improvements are 0.78$\%$, 0.83$\%$, and 0.81$\%$ respectively. Xiaotian Wang 0001, Xuanhang Xu, Zhifu Zhao, Fu Li 0002, Fei Qi 0001, Shuo Liang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Adaptive Progressive Attention Graph Neural Network for EEG Emotion RecognitionabstractIn recent years, numerous neuroscientific studies have demonstrated that specific brain regions are associated with human emotional responses, with these regions exhibiting variability across individuals and emotions. To effectively leverage these neural patterns, we propose an Adaptive Progressive Attention Graph Neural Network (APAGNN), which dynamically models the spatial relationships among brain regions during emotional processing. APAGNN employs three specialized expert modules that progressively analyze brain topology. The first expert captures global brain connectivity patterns, the second extracts localized regional features, and the third focuses on emotion-related channel interactions. This progressive refinement strategy enables hierarchical feature extraction from coarsegrained to fine-grained neural representations. Furthermore, a weight generator integrates the outputs of all three experts, adaptively balancing their contributions for final emotion recognition. Extensive experiments conducted on SEED, SEED-IV and MPED datasets demonstrate that our method significantly improves EEG emotion recognition performance, achieving superior results compared to baseline methods. Tianzhi Feng, Chennan Wu, Fu Li 0002, Yang Li 0019, Boxun Fu, Zhifu Zhao, Xiaotian Wang 0001 |
BIBM | 7 |
| 2025 | Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action RecognitionabstractSkeleton sequences for action recognition are with complex temporal dynamics due to various factors such as speed variation and different activities. It is crucial and essential to model variation changes in the temporal dimension. In recent years, skeleton sequence is always modeled as a graph structure, and Graph Convolution Network (GCN) is employed to extract spatial and temporal features of actions. Though GCN has obtained great achievements, they typically employ fixed-size temporal kernels for temporal modeling, which ignore the complex temporal dynamic of actions, especially for long-term as well as short-term modeling. To capture this complex motion pattern effectively, we propose a Dual Multi-Scale Graph Convolutional Network (DMS-GCN), which is mainly composed of a Deformable Temporal Kernel (DTK) block and a dual multi-scale strategy. Specifically, the DTK block is proposed to flexibly capture complex temporal information of the skeleton sequence. And the dual multi-scale strategy is used to simultaneously accommodate long-term and short-term dynamic information at different scales globally as well as locally. The effectiveness of our proposed method is verified through experiments conducted on two widely used datasets, NTU-RGB+D 60 and NTU-RGB+D 120. Jianan Li 0003, Yangtao Zhou, Hua Chu, Zhifu Zhao, Fei Li 0030, Qingshan Li |
ICASSP | 5 |
| 2025 | Multi-Scale Adaptive Skeleton Transformer for action recognition
Xiaotian Wang 0001, Zhifu Zhao, Guangming Shi, Xuemei Xie, Xiang Jiang 0011 |
Comput. Vis. Image Underst. | 3 |
| 2025 | Exploring interaction: Inner-outer spatial-temporal transformer for skeleton-based mutual action recognition
Xiaotian Wang 0001, Xiang Jiang 0011, Zhifu Zhao |
Neurocomputing | 3 |
| 2025 | IMFR-Net: Interval Measurement and Full Recovery Network for video compressive sensing
Wanxin Zhang, Zhifu Zhao, Fu Li 0002, Jianan Li 0003 |
Neurocomputing | 3 |
| 2024 | DIDA: Dynamic Individual-to-integrateD Augmentation for Self-supervised Skeleton-Based Action Recognition
Haobo Huang, Jianan Li 0003, Zhifu Zhao, Yangtao Zhou |
PRCV (7) | 4 |
| 2024 | STDM-transformer: Space-time dual multi-scale transformer network for skeleton-based action recognition
Zhifu Zhao, Jianan Li 0003, Xuemei Xie, Xiaotian Wang 0001, Guangming Shi |
Neurocomputing | 1 |
| 2024 | Glimpse and Zoom: Spatio-Temporal Focused Dynamic Network for Skeleton-Based Action RecognitionabstractGCN-based methods have achieved remarkable performance in skeleton-based action recognition. However, existing methods have not explicitly attempted to remove temporal and spatial redundancy that might introduce additional computational costs. Inspired by the fact that humans always tend to glimpse at overall motion and then zoom into the most important spatio-temporal regions, we propose a Spatio Temporal Focused Dynamic Network (STFD-Net) trained with reinforcement learning for skeleton-based action recognition. Specifically, we first propose a global extractor with Skeleton Pooling Module (SPM) to enable the network to focus on overall motion information with a refined skeleton structure. Then, a local extractor, containing pair-wise part partition, tubelet proposal network, and Partition-Grouped Module (PGM), is proposed to extract local motion details as a complement to the overall motion information. Finally, the dynamic classifier utilizes a recurrent neural network to dynamically terminate the process once the network is adequately confident. Extensive experiments have demonstrated that the proposed network achieves SOTA level performance with lower computational cost on the NTU 60 and NTU 120 dataset. Zhifu Zhao, Jianan Li 0003, Xiaotian Wang 0001, Xuemei Xie, Wanxin Zhang, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Adaptive Spatio-Temporal Directed Graph Neural Network for Parkinson's Detection using Vertical Ground Reaction ForceabstractVertical Ground Reaction Force (VGRF) signal obtained from foot-worn sensors, also known as plantar data, provides a highly informative and detailed representation of an individual's gait features. Existing methods, such as CNNs, LSTMs and Transformers, have revealed the efficiency of deep learning in Parkinson's Disease (PD) diagnosis using VGRF signal. However, the intrinsic topologic graph and pressure transmission characteristics of plantar data are overlooked in those approaches, which are essential features for gait analysis. In this paper, we propose to construct a plantar directed topologic graph to fully exploit the plantar topology in gait circles. It can facilitate the expression of gait information by representing sensors as nodes and pressure transmissions as directional edges. Accordingly, an Adaptive Spatio-Temporal Directed Graph Neural Network (AST-DGNN) is proposed to extract the connection features of the plantar directed topologic graph. Each AST-DGNN Unit includes an Adaptive Directed Graph Network (ADGN) block and a Temporal Convolutional Network (TCN) block. In order to capture both local and global spatial relationships among sensor nodes and pressure transmission edges, the ADGN block performs message passing on the plantar directed topologic graph in an adaptive manner. To capture the temporal features of sensor nodes and pressure transmission edges, the TCN block defines a temporal feature extraction process for each node and edge in the graph. Moreover, the data augmentation is introduced for plantar data to improve the generalization ability of the AST-DGNN. Experimental results on Ga, Ju, and Si datasets demonstrate that the proposed method outperforms the existing methods under both cross-dataset validation and mixed-data cross-validation. Especially in cross-dataset validation, there is an average improvement of 2.13%, 7.73%, and 12.27% in accuracy, F1 score, and G-mean, respectively. Xiaotian Wang 0001, Shuo Liang, Zhifu Zhao, Xuanhang Xu |
ACM Multimedia | 3 |
| 2023 | View-Normalized and Subject-Independent Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest in computer vision. For this task, a challenging problem concerns the large intraclass variances of skeleton data, which are mainly caused by diverse viewpoints and subjects, and greatly increase the difficulty of modeling actions through a network. To address the above problem, we propose a variance reduction (VaRe) framework for skeleton-based action recognition, which consists of a view-normalization generative adversarial network (VN-GAN), a subject-independent network (SINet) and a classification network. First, the VN-GAN is responsible for reducing view-induced intraclass variances. Specifically, this network, comprising a generator and a discriminator, is aimed at learning a mapping from a diverse-view skeleton distribution to a unified-view skeleton distribution in an unsupervised manner, thereby generating a view-normalized skeleton. Second, taking the view-normalized skeleton as input, the SINet focuses on reducing the influences of the personal habits of subjects on action recognition. To generate SI skeleton data, the SINet automatically adjusts the human pose according to the human kinematic structure under a classification loss constraint. Finally, without the interference of view- and subject-induced variances, the classification network can concentrate more on learning discriminative action features to predict classes. Furthermore, by combining the joint and bone modalities, the proposed framework achieves competitive performance on three benchmarks: NTU RGB+D, NTU-120 RGB+D and Northwestern-UCLA Multiview Action 3D. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | View-normalized Skeleton Generation for Action RecognitionabstractSkeleton-based action recognition has attracted great interest due to low cost of skeleton data acquisition and high robustness to external conditions. A challenging problem of skeleton-based action recognition is the large intra-class gap caused by various viewpoints of skeleton data, which makes the action modeling difficult for network. To alleviate this problem, a feasible solution is to utilize label supervised methods to learn a view-normalization model. However, since the skeleton data in real scenes is acquired from diverse viewpoints, it is difficult to obtain the corresponding view-normalized skeleton as label. Therefore, how to learn a view-normalization model without the supervised label is the key to solving view-variance problem. To this end, we propose a view normalization-based action recognition framework, which is composed of view-normalization generative adversarial network (VN-GAN) and classification network. For VN-GAN, the model is designed to learn the mapping from diverse-view distribution to normalized-view distribution. In detail, it is implemented by graph convolution, where the generator predicts the transformation angles for view normalization and discriminator classifies the real input samples from the generated ones. For classification network, view-normalized data is processed to predict the action class. Without the interference of view variances, classification network can extract more discriminative feature of action. Furthermore, by combining the joint and bone modalities, the proposed method reaches the state-of-the-art performance on NTU RGB+D and NTU-120 RGB+D datasets. Especially in NTU-120 RGB+D, the accuracy is improved by 3.2% and 2.3% under cross-subject and cross-set criteria, respectively. Qingzhe Pan, Zhifu Zhao, Xuemei Xie, Jianan Li 0003, Guangming Shi |
ACM Multimedia | 2 |
| 2021 | Knowledge embedded GCN for skeleton-based two-person interaction recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Neurocomputing | 5 |
| 2020 | SGM-Net: Skeleton-guided multimodal network for action recognition
Jianan Li 0003, Xuemei Xie, Qingzhe Pan, Zhifu Zhao, Guangming Shi |
Pattern Recognit. | 5 |
| 2019 | Visualizing and understanding of learned compressive sensing with residual network
Zhifu Zhao, Xuemei Xie, Chenye Wang, Wan Liu 0001, Guangming Shi, Jiang Du 0011 |
Neurocomputing | 1 |
| 2019 | ROI-CSNet: Compressive sensing network for ROI-aware image recovery
Zhifu Zhao, Xuemei Xie, Chenye Wang, Siying Mao, Wan Liu 0001, Guangming Shi |
Signal Process. Image Commun. | 1 |