Aite Zhao

dblp:169/4591 · DBLP profile ↗
← Back
25ranked-venue papers
13as first author
22since 2021 · last 2026
0000-0003-3494-175XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VoxTransPD: A Reusable Framework for Noise-Resilient Speech Analysis and Early Parkinson's Detection
Tengfei Qu, Yanmiao Kong, Aite Zhao
MSR4
2025 ASA-DSNet: A Dual-Stage Framework with Ambiguous Sample Augmentation for Robust Parkinson's Disease Identification
abstract
Parkinson's disease (PD) is a neurodegenerative disorder with rising prevalence, and early diagnosis is critical to slowing disease progression. Traditional diagnostic methods rely on clinical rating scales, which are inherently subjective and prone to misdiagnosis. This paper proposes ASA-DSNet, a dualstage learning framework for PD detection using smartwatch motion data. The first stage$(M_{1})$fuses multi-scale time-frequency features and identifies ambiguous samples (i.e., boundary and misclassified samples), which are augmented using the Ambiguous Sample Augmenter (ASA). The second stage$(M_{2})$employs hierarchical attention mechanisms to refine decision boundaries. Experimental results demonstrate that ASA-DSNet achieves accuracies of 80.85% and 95.17% on the PADS and Monipar datasets, outperforming representative baselines. This model provides a viable solution for early PD screening and tremor grading.
Xinyue Xing, Ranxuan Tang, Yanmiao Kong, Aite Zhao
BIBM4
2025 An Efficient Hierarchical Network for Segmentation of Linear Structures in Biomedical Images
abstract
Accurate segmentation of line-like structures in biomedical images (e.g., vasculature, muscle fibers) is critical, for their morphology, density, and connectivity serve as vital biomarkers for analysis, yet it remains challenging. Mainstream segmentation models have massive scale and produce fragment filamentous structures, while specialized methods often require complex post-processing. Therefore, we propose Line-like Structure Segmentation Network (LSSeg), an ultra-lightweight hierarchical architecture (28 K parameters) that directly outputs predictions for this specific task. LSSeg uniquely integrates segmentation and edge detection strengths by combining multi-level feature perception (via Feature Decimation Modules) with locally aware reconstruction (via Focus Locally Block) to achieve high-quality output while maintaining extreme efficiency (2.1 GFLOPs). Experiments across diverse datasets (retinal vessels, muscle aponeurosis, microtubule) show LSSeg outperforms universal segmentation models as well as edge detection models. The code is available at https://github.com/BM-AI-Lab/LSSeg.
Xinglin Yu, Xiutao Shi, Guangdian Zhang, Aite Zhao
BIBM6
2025 A dynamic saliency enhanced fusion network for multimodal sleep recognition
Aite Zhao, Huimin Wu 0002
Multim. Tools Appl.1
2024 TF-FusNet: A Novel Framework for Parkinson's Disease Detection via Time-Frequency Domain Fusion
Weijie Yu 0002, Aite Zhao, Huiyu Zhou 0001
NLPCC (2)4
2024 A Triplet Multimodel Transfer Learning Network for Speech Disorder Screening of Parkinson's Disease
abstract
Deterioration in the quality of a person’s voice and speech is an early sign of Parkinson’s disease (PD). Although a number of computer-based methods have been invested to use patients’ speech for early diagnosis of Parkinson’s disease, they only focus on a fixed pronunciation test, such as the subjects’ monosyllabic pronunciation is analyzed to determine whether they have potential possibility of PD. Moreover, only using traditional speech analysis methods to extract single-view speech features cannot provide a comprehensive feature representation. This paper is dedicated to the study of various pronunciation tests for patients with PD, including the pronunciation of five monosyllabic vowels and a spontaneous dialogue. A triplet multimodel transfer learning network is designed and proposed for identifying subjects with PD in these two groups of tests. First, multisource data extract mel frequency cepstrum coefficient (MFCC) features of speech for preprocessing. Subsequently, a pretrained triplet model represents features from three dimensions as the upstream task of the transfer learning framework. Finally, the pretrained model is reconstructed as a novel model that integrates the triplet model, temporal model, and auxiliary layer as the downstream task, and weights are updated through fine-tuning to identify abnormal speech. Experimental results show that the highest PD detection rates in the two groups of tests are 99% and 90% , respectively, which outperform a large number of internationally popular pattern recognition algorithms and serve as a baseline for other academic researchers in this field.
Aite Zhao, Xuesen Niu, Huimin Wu 0002
Int. J. Intell. Syst.1
2024 A multi-level feature attention network for COVID-19 detection based on multi-source medical images
Aite Zhao, Huimin Wu 0002
Multim. Tools Appl.1
2023 A coordinate attention enhanced swin transformer for handwriting recognition of Parkinson's disease
abstract
Abstract Diagnosing Parkinson's disease (PD) in its early stages is a significant challenge in medicine. Hand tremors and dysgraphia, which are typical early motor symptoms of PD, can manifest for decades before a formal diagnosis is made. Therefore, handwriting analysis has become an important tool for detecting PD. While many machine learning algorithms have been applied in this area, they struggle to capture the subtle changes in handwriting and must describe features from various perspectives. To address these issues, this paper proposes a Coordinate Attention Enhanced Swin Transformer (CAS Transformer) model for PD handwriting recognition. It establishes the long‐term dependence of features on the joint coordinate attention application, which enables the model to more accurately localize the important features of handwriting data and also extract the fuzzy edge features of handwriting images.These characteristics of the CAS Transformer enable it to outperform current advanced deep learning methods in classification, with an accuracy of 92.68% in experiments conducted on two handwritten datasets.
Xuesen Niu, Yiyang Yuan, Yunze Sun, Guoliang You, Aite Zhao
IET Image Process.7
2023 DCACorrCapsNet: A deep channel-attention correlative capsule network for COVID-19 detection based on multi-source medical images
abstract
Abstract The raging trend of COVID‐19 in the world has become more and more serious since 2019, causing large‐scale human deaths and affecting production and life. Generally speaking, the methods of detecting COVID‐19 mainly include the evaluation of human disease characterization, clinical examination and medical imaging. Among them, CT and X‐ray screening is conducive to doctors and patients' families to observe and diagnose the severity and development of the COVID‐19 more intuitively. Manual diagnosis of medical images leads to low the efficiency, and long‐term tired gaze will decline the diagnosis accuracy. Therefore, a fully automated method is needed to assist processing and analysing medical images. Deep learning methods can rapidly help differentiate COVID‐19 from other pneumonia‐related diseases or healthy subjects. However, due to the limited labelled images and the monotony of models and data, the learning results are biased, resulting in inaccurate auxiliary diagnosis. To address these issues, a hybrid model: deep channel‐attention correlative capsule network, for channel‐attention based spatial feature extraction, correlative feature extraction, and fused feature classification is proposed. Experiments are validated on X‐ray and CT image datasets, and the results outperform a large number of existing state‐of‐the‐art studies.
Aite Zhao, Huimin Wu 0002
IET Image Process.1
2023 A Spatio-Temporal Siamese Neural Network for Multimodal Handwriting Abnormality Screening of Parkinson's Disease
abstract
Currently, hand motion recognition of single‐modality data has been extensively explored for the analysis of various contact and noncontact sensors, and it is recognized that all the existing technologies have both strengths and limitations. As a significant motor symptom, hand tremor is usually utilized for the diagnosis and evaluation of Parkinson’s disease; furthermore, a multimodal analysis of the handwriting pattern of the patient has made up for the one‐sided way of learning the hand movement in a single measurement dimension. Especially, considering a variety of measurement resources, it shows promising performance in recognizing handwriting patterns of Parkinson’s disease. In this work, a novel Spatio‐temporal Siamese neural network (ST‐SiamNN) is proposed to learn the handwriting differences between healthy individuals and patients with Parkinson’s disease, process data onto multiple sensors, and enhance the characteristics of handwriting in Parkinson’s disease. Uniquely, it is a discriminative model of multilabel and multinetwork constructed by a Siamese network, which consists of four modules: a preprocessor for handwritten data enhancement, a Siamese bidirectional memory neural network (SiamBiMNN) for temporal and texture feature extraction and difference enhancement, a Siamese octave convolutional neural network (SiamOctCNN) for spatial feature extraction and difference enhancement, and a decision‐making layer to rejudge the output features of the Siamese networks to obtain more accurate auxiliary diagnosis results. The framework proposed in this article is verified on two handwritten datasets of multiple modalities, i.e., images, smart pen signals, and graphics tablet signals, which are compared with several state‐of‐the‐art studies.
Aite Zhao, Huimin Wu 0002
Int. J. Intell. Syst.1
2023 A significantly enhanced neural network for handwriting assessment in Parkinson's disease detection
Aite Zhao
Multim. Tools Appl.1
2023 Multi-attribute Graph Convolution Network for Regional Traffic Flow Prediction
Yue Wang 0052, Aite Zhao, Zhiqiang Lv, Chuanhao Dong, Haoran Li 0021
Neural Process. Lett.2
2023 Transferable Self-Supervised Instance Learning for Sleep Recognition
abstract
Although the importance of sleep is increasingly recognized, the lack of general and transferable algorithms hinders scalable sleep assessment in healthy persons and those with sleep disorders. A deep understanding of the sleep posture, state, or stage is the premise of diagnosing and treating sleep diseases. At present, most existing methods draw support from supervised learning to monitor the whole sleep process. However, in the absence of sufficient labeled sleep data, it is difficult to guarantee the reliability of sleep recognition networks. To solve this problem, we propose a transferable self-supervised instance learning model for three sleep recognition tasks, i.e., sleep posture, state, and stage recognition. Firstly, a SleepGAN is designed to generate sleep data, and then, we combine enough self-supervised rotating sleep data and original data for non-parametric classification at the instance-level, finally, different sleep postures, states, or stages can be distinguished precisely. The proposed model can be applied to multimodal sleep data such as signals and images, and makeup for the inaccuracy caused by insufficient data, and can be transferred to sleep datasets of different sizes. The experimental results show that our algorithm for the physiological changes in the sleep process is superior to several state-of-the-art studies, which may be helpful to promote the intelligence of sleep assessment and monitoring.
Aite Zhao, Yue Wang 0052
IEEE Trans. Multim.1
2022 Socially Acceptable Trajectory Prediction for Scene Pedestrian Gathering Area
Rongkun Ye, Zhiqiang Lv, Aite Zhao
WASA (1)3
2022 A deep spatio-temporal meta-learning model for urban traffic revitalization index prediction in the COVID-19 pandemic
Yue Wang 0052, Zhiqiang Lv, Zhaoyu Sheng, Haokai Sun 0002, Aite Zhao
Adv. Eng. Informatics5
2022 Two-channel lstm for severity rating of parkinson's disease using 3d trajectory of hand motion
Aite Zhao
Multim. Tools Appl.1
2022 Multimodal Gait Recognition for Neurodegenerative Diseases
abstract
In recent years, single modality-based gait recognition has been extensively explored in the analysis of medical images or other sensory data, and it is recognized that each of the established approaches has different strengths and weaknesses. As an important motor symptom, gait disturbance is usually used for diagnosis and evaluation of diseases; moreover, the use of multimodality analysis of the patient's walking pattern compensates for the one-sidedness of single modality gait recognition methods that only learn gait changes in a single measurement dimension. The fusion of multiple measurement resources has demonstrated promising performance in the identification of gait patterns associated with individual diseases. In this article, as a useful tool, we propose a novel hybrid model to learn the gait differences between three neurodegenerative diseases, between patients with different severity levels of Parkinson's disease, and between healthy individuals and patients, by fusing and aggregating data from multiple sensors. A spatial feature extractor (SFE) is applied to generating representative features of images or signals. In order to capture temporal information from the two modality data, a new correlative memory neural network (CorrMNN) architecture is designed for extracting temporal features. Afterward, we embed a multiswitch discriminator to associate the observations with individual state estimations. Compared with several state-of-the-art techniques, our proposed framework shows more accurate classification results.
Aite Zhao, Junyu Dong, Lin Qi 0004, Qianni Zhang, Ning Li 0012, Xin Wang 0068, Huiyu Zhou 0001
IEEE Trans. Cybern.1
2022 Associated Spatio-Temporal Capsule Network for Gait Recognition
abstract
It is a challenging task to identify a person based on her/his gait patterns. State-of-the-art approaches rely on the analysis of temporal or spatial characteristics of gait, and gait recognition is usually performed on single modality data (such as images, skeleton joint coordinates, or force signals). Evidence has shown that using multi-modality data is more conducive to gait research. Therefore, we here establish an automated learning system, with an associated spatio-temporal capsule network (ASTCapsNet) trained on multi-sensor datasets, to analyze multimodal information for gait recognition. Specifically, we first design a low-level feature extractor and a high-level feature extractor for spatio-temporal feature extraction of gait with a novel recurrent memory unit and a relationship layer. Subsequently, a Bayesian model is employed for the decision-making of class labels. Extensive experiments on several public datasets (normal and abnormal gait) validate the effectiveness of the proposed ASTCapsNet, compared against several state-of-the-art methods.
Aite Zhao, Junyu Dong, Lin Qi 0004, Huiyu Zhou 0001
IEEE Trans. Multim.1
2021 Multimodal Traffic Travel Time Prediction
abstract
With the continuous growth of urban population, it is urgent for people to accurately plan the travel time. Therefore, travel time prediction of urban areas has become a key research direction in the field of smart cities. At present, several studies on travel time prediction are only conducted on a single mode, where the prediction process only treats a certain vehicle as an isolated traffic state on the route. However, the factors affecting traffic are extremely complex, thus making it very difficult to produce a comprehensive forecast. Based on this situation, the mixed existing model and mutual influence of multiple modes of transportation in the city are fully considered, and a multimodal deep learning model namely MC-GRU (Multimodal Convoluted Gated Recurrent Unit Network) is proposed. At the same time, to solve the problem of some objective factors, such as departure time and travel distance, we propose an attribute module to deal with these implicit factors. In addition, to explore the interaction between different modes of vehicles, a feature fusion module for obtaining the interaction effect between different modes of vehicles is proposed. Finally, we use GRU to learn the long-term dependence. MC-GRU can realize the accurate prediction of travel time in multimodal traffic state, as well as implement travel time prediction for three types of travel modes. The experimental results show that MC-GRU achieves higher prediction accuracy on a challenging real world dataset as compared with MAE, MAPE and RMSE.
Shizhen Fan, Zhiqiang Lv, Aite Zhao
IJCNN4
2021 Temporal Attention-Based Graph Convolution Network for Taxi Demand Prediction in Functional Areas
Yue Wang 0052, Aite Zhao, Zhiqiang Lv, Guangquan Lu
WASA (1)3
2021 Perceptual Underwater Image Enhancement With Deep Learning and Physical Priors
abstract
Underwater image enhancement, as a pre-processing step to support the following object detection task, has drawn considerable attention in the field of underwater navigation and ocean exploration. However, most of the existing underwater image enhancement strategies tend to consider enhancement and detection as two fully independent modules with no interaction, and the practice of separate optimisation does not always help the following object detection task. In this article, we propose two perceptual enhancement models, each of which uses a deep enhancement model with a detection perceptor. The detection perceptor provides feedback information in the form of gradients to guide the enhancement model to generate patch level visually pleasing or detection favourable images. In addition, due to the lack of training data, a hybrid underwater image synthesis model, which fuses physical priors and data-driven cues, is proposed to synthesise training data and generalise our enhancement model for real-world underwater images. Experimental results show the superiority of our proposed method over several state-of-the-art methods on both real-world and synthetic underwater datasets.
Long Chen 0019, Zheheng Jiang, Aite Zhao, Qianni Zhang, Junyu Dong, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Multi-View Mouse Social Behaviour Recognition With Deep Graphic Model
abstract
Home-cage social behaviour analysis of mice is an invaluable tool to assess therapeutic efficacy of neurodegenerative diseases. Despite tremendous efforts made within the research community, single-camera video recordings are mainly used for such analysis. Because of the potential to create rich descriptions for mouse social behaviors, the use of multi-view video recordings for rodent observations is increasingly receiving much attention. However, identifying social behaviours from various views is still challenging due to the lack of correspondence across data sources. To address this problem, we here propose a novel multi-view latent-attention and dynamic discriminative model that jointly learns view-specific and view-shared sub-structures, where the former captures unique dynamics of each view whilst the latter encodes the interaction between the views. Furthermore, a novel multi-view latent-attention variational autoencoder model is introduced in learning the acquired features, enabling us to learn discriminative features in each view. Experimental results on the standard CRMI13 and our multi-view Parkinson's Disease Mouse Behaviour (PDMB) datasets demonstrate that our proposed model outperforms the other state of the arts technologies, has lower computational cost than the other graphical models and effectively deals with the imbalanced data problem.
Zheheng Jiang, Feixiang Zhou, Aite Zhao, Xin Li 0052, Ling Li 0010, Dacheng Tao, Xuelong Li 0001, Huiyu Zhou 0001
IEEE Trans. Image Process.3
2020 SpiderNet: A spiderweb graph neural network for multi-view gait recognition
Aite Zhao, Manzoor Ahmed
Knowl. Based Syst.1
2018 A hybrid spatio-temporal model for detection and severity rating of Parkinson's disease from gait data
Aite Zhao, Lin Qi 0004, Jie Li 0001, Junyu Dong, Hui Yu 0001
Neurocomputing1
2018 Dual channel LSTM based multi-feature extraction in gait for diagnosis of Neurodegenerative diseases
Aite Zhao, Lin Qi 0004, Junyu Dong, Hui Yu 0001
Knowl. Based Syst.1