Huaiyu Zhu 0004

dblp:17/5913-4 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-6918-4088ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Inferring Smartphone Application Types via Inaudible Charging Sound: A New Acoustic Side-channel Attack
abstract
The security and privacy protection of mobile devices have been critical issues. This study explores a novel acoustic side-channel attack method that infers the category of applications (apps) being used by analyzing the inaudible sounds emitted from chargers during the smartphone charging process. Unlike traditional methods that require direct contact or modification to the device, this approach demonstrates widespread availability and enhanced concealment. Through preprocessing, feature extraction, and model training, we conducted experiments using various smartphones, apps and chargers. The results indicate that both Random Forest and Transformer models achieve nearly perfect performance in app type inference under random splits cross-validation test, while the Transformer with end-to-end feature learning shows higher stability and generalizability under strict leave-one-out cross-validation, with an average accuracy of 63.52%. Finally, we give further analysis on effectiveness and efficiency of the proposed method, highlighting the potential threat of acoustic side-channel attacks and emphasizing the necessity of enhancing privacy protection and improved charging technologies.
Junchen Meng, Huaiyu Zhu 0004
SMC4
2025 Swallow-PPG: Photoplethysmography Templates for Comprehensive Temporal Analysis of Swallowing Anatomical Actions
abstract
In clinical practice, Videofluoroscopic Swallowing Study (VFSS) is commonly used to monitor the activity of anatomical structures during swallowing. However, it is limited by ionizing radiation exposure, adverse effects of barium contrast agents, and the high cost of specialized equipment. In this study, we propose a framework for analyzing swallowing behaviors in photoplethysmography (PPG) waveforms, which includes generalizing the manifestation of swallowing in PPG (i.e., swallowing templates generation) and conducting comprehensive temporal analysis of swallowing anatomical actions (TASAA). For swallowing templates generation, we cluster and average the samples to obtain waveforms of templates, followed by conducting shape-based mapping and averaging on 28 time indicators to derive template unified time indicators (TUTIs). For comprehensive TASAA, we leverage templates waveforms and TUTIs to estimate time indicators based on the mapping relationship between samples and their respective templates. We evaluate the proposed framework on 357 swallowing PPG samples from 41 elderly subjects. The average relative error across all time indicators is 0.123, and 6 indicators notably excel with errors below 0.1. The proposed template-based swallowing analysis framework is expected to become a low-cost and non-ionizing alternative to VFSS for comprehensive TASAA.
Ying Zhang 0128, Huaiyu Zhu 0004
IEEE J. Biomed. Health Informatics4
2024 Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
abstract
Non-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT.
Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng
DAC11
2024 Towards a Deeper Insight Into Face Detection in Neonatal Wards
Yisheng Zhao, Huaiyu Zhu 0004, Qi Shu, Ruohong Huan, Shuohui Chen
MICCAI (5)2
2024 Heterogeneous Graph Network for Action Detection
abstract
Spatio-temporal action detection is a fundamental task that detects persons and recognizes their actions from videos. It requires reasoning about the spatial-temporal interactions between persons and their surroundings. Recently, more modalities have been found by researchers, which puts higher demands on the reasoning capability of the method, yet a method capable of holistic reasoning is still lacking. To this end, we propose a heterogeneous graph network, which aims to reason the spatial-temporal interactions among different types of nodes (video entities) and edges (inter-entity relations). Concretely, it includes spatial and temporal graphs, which are alternately updated. The spatial graph contains nodes of person appearance, person pose, object appearance, and hand interaction, and the temporal graph has person nodes at different moments. For information aggregation, we propose a person-centric heterogeneous graph reasoning algorithm, which introduces heterogeneity into the graphs through node-type-specific projections and modulated edge-type-specific representations. We find that the introduction of heterogeneity enriches the model’s ability to understand multi-modality, which facilitates better parsing of complex semantic relations in videos and potentially leads to further mining of spatial-temporal interactions between entities in the future. Experimental results on four public datasets demonstrate the superiority of our method. Code will be available after acceptance.
Yisheng Zhao, Huaiyu Zhu 0004, Ruohong Huan, Yaoqi Bao
IEEE Trans. Circuits Syst. Video Technol.2
2024 GMAEEG: A Self-Supervised Graph Masked Autoencoder for EEG Representation Learning
abstract
Annotated electroencephalogram (EEG) data is the prerequisite for artificial intelligence-driven EEG autoanalysis. However, the scarcity of annotated data due to its high-cost and the resulted insufficient training limits the development of EEG autoanalysis. Generative self-supervised learning, represented by masked autoencoder, offers potential but struggles with non-Euclidean structures. To alleviate these challenges, this work proposes a self-supervised graph masked autoencoder for EEG representation learning, named GMAEEG. Concretely, a pretrained model is enriched with temporal and spatial representations through a masked signal reconstruction pretext task. A learnable dynamic adjacency matrix, initialized with prior knowledge, adapts to brain characteristics. Downstream tasks are achieved by finetuning pretrained parameters, with the adjacency matrix transferred based on task functional similarity. Experimental results demonstrate that with emotion recognition as the pretext task, GMAEEG reaches superior performance on various downstream tasks, including emotion, major depressive disorder, Parkinson's disease, and pain recognition. This study is the first to tailor the masked autoencoder specifically for EEG representation learning considering its non-Euclidean characteristics. Further, graph connection analysis based on GMAEEG may provide insights for future clinical studies.
Zanhao Fu, Huaiyu Zhu 0004, Yisheng Zhao, Ruohong Huan, Yi Zhang 0126, Shuohui Chen
IEEE J. Biomed. Health Informatics2
2024 Video-Based Neonatal Pain Assessment in Uncontrolled Conditions
abstract
OBJECTIVE: Neonatal pain can have long-term adverse effects on newborns' cognitive and neurological development. Video-based Neonatal Pain Assessment (NPA) method has gained increasing attention due to its performance and practicality. However, existing methods focus on assessment under controlled environments while ignoring real-life disturbances present in uncontrolled conditions. METHODS: We propose a video-based NPA method, which is robust to four real-life disturbances and adaptively highlights keyframes. Our method involves a region-channel-attention module for extracting facial features under the disturbances of facial occlusion and pose variation; a body language analysis module robust to disturbances from body occlusion and movement interference, which utilizes skeleton sequences to represent the neonate's body; and a keyframes-aware convolution to get rid of information located at non-contributing moments. For evaluation, we built an NPA video dataset of 1091 neonates with disturbance annotations. RESULTS: The results show that our method consistently outperforms state-of-the-art methods on the full dataset and nine subsets, where it achieves an accuracy of 91.04% on the full dataset with an accuracy increment of 6.27%. Contributions: We present the problem of video-based NPA under uncontrolled conditions, propose a method robust to four disturbances, and construct a video NPA dataset, thus facilitating the practical applications of NPA.
Huaiyu Zhu 0004, Yisheng Zhao, Feixiang Luo, Lingli Mei, Shuohui Chen
IEEE J. Biomed. Health Informatics1