Shiqi Zhao 0001

dblp:68/887-1 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-7208-7075ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection
abstract
Depression represents a global mental health challenge requiring efficient and reliable automated detection methods. Current Transformer- or Graph Neural Networks (GNNs)-based multimodal depression detection methods face significant challenges in modeling individual differences and cross-modal temporal dependencies across diverse behavioral contexts. Therefore, we propose P³HF (Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network) with three key innovations: (1) personality-guided representation learning using LLMs to transform discrete individual features into contextual descriptions for personalized encoding; (2) Hypergraph-Former architecture modeling high-order cross-modal temporal relationships; (3) event-level domain disentanglement with contrastive learning for improved generalization across behavioral contexts. Experiments on MPDD-Young dataset show P³HF achieves around 10% improvement on accuracy and weighted F1 for binary and ternary depression classification task over existing methods. Extensive ablation studies validate the independent contribution of each architectural component, confirming that personality-guided representation learning and high-order hypergraph reasoning are both essential for generating robust, individual-aware depression-related representations.
Changzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian, Shiqi Zhao 0001
AAAI5
2026 BioSeek: A Design Generation Framework of Biosignal Processors with Large-Language Models for Edge Healthcare Applications
abstract
Deep neural network (DNN)-based methodologies have shown impressive performance and robustness in the detection of abnormalities and decoding of multi-modal biosignals. While the use of DNNs provides promising classification and decoding capabilities, it also introduces significant design and cost challenges for the implementation of biomedical System on Chips (SoC). To address the increasing demand for advanced and efficient DNN-based healthcare solutions at the edge, we propose BioSeek, an agile design generation framework enhanced by cutting-edge large-language models (LLM). BioSeek offers a comprehensive solution to the design challenges associated with biosignal processors. The effectiveness of BioSeek is evaluated through the design generation of both application-specific and versatile biosignal processors, demonstrating performance that is competitive with existing solutions.
Fengshi Tian, Jiakun Zheng, Hui Wu 0010, Zilu Liu, Jinbo Chen 0002, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng
ISCAS6
2026 A plug-and-play hybrid pruning framework for spike-driven transformers via adaptive spiking statistical importance scoring
Hanfei Liu, Shiqi Zhao 0001, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Neurocomputing2
2026 HAQ-ViT: A hardware-aware post-training quantization for efficient vision transformer inference
Shiqi Zhao 0001, Haozu Sun, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Knowl. Based Syst.1
2026 Memory-Efficient Intrinsic Gating Adaptation for Enhanced On-Device Epilepsy Diagnosis
abstract
Recently, advances in neuroscience and the rise of artificial intelligence have significantly enhanced the capabilities of epilepsy diagnosis. While EEG-based diagnosis offer a promising avenue for detecting and predicting seizure activity, practical implementation in real-world scenarios remains hindered by the heterogeneity of epilepsy and the variability of patient-specific biomarkers over time. Conventional deep learning models, trained on historical EEG, often fail to adapt to such biomarker variations, leading to degraded performance. Moreover, the computational and memory constraints of edge devices further exacerbate the challenge of on-device learning. To address these challenges, we introduce a novel framework, Memory-Efficient Intrinsic Gating Adaptation (MEIGA), designed to enhance real-world epilepsy diagnosis on resource-constrained edge devices. Our approach pre-trains a model using historical EEG data and employs lightweight adapter networks for efficient on-device tuning across new sessions, addressing session-to-session variability. By leveraging Direct Feedback Alignment (DFA), MEIGA reduces memory usage and computational overhead while maintaining high classification accuracy. Extensive experiments on the CHB-MIT epilepsy dataset demonstrate that MEIGA outperforms the pretrained-only Vision Transformer baseline, raising seizure prediction accuracy from 47.88% to 86.77% with only 3,908 tunable parameters (5.05% of the backbone). For seizure detection, MEIGA improves accuracy from 85.06% to 96.29% by adapting 2,008 parameters (17.40% of the base architecture). Further experiments on the AES dataset demonstrate that MEIGA consistently delivers strong performance across subjects and scales effectively to larger networks.
Shanjin Li, Di Wu 0057, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics3
2025 The First MPDD Challenge: Multimodal Personality-aware Depression Detection
abstract
Depression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detection methods primarily focus on young adults, neglecting the broader age spectrum and individual differences that influence depression manifestation. Current approaches often establish a direct mapping between multimodal data and depression indicators, failing to capture the complexity and diversity of depression across individuals. This challenge includes two tracks based on age-specific subsets: Track 1 uses the MPDD-Elderly dataset for detecting depression in older adults, and Track 2 uses the MPDD-Young dataset for detecting depression in younger participants. The Multimodal Personality-aware Depression Detection (MPDD) Challenge aims to address this gap by incorporating multimodal data alongside individual difference factors. We provide a baseline model that fuses audio and video modalities with individual difference information to detect depression manifestations in diverse populations. This challenge aims to promote the development of more personalized and accurate de pression detection methods, advancing mental health research and fostering inclusive detection systems. More details are available on the official challenge website: https://hacilab.github.io/MPDDChallenge.github.io.
Changzeng Fu, Zelin Fu, Qi Zhang 0124, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Junfeng Yao, Yuliang Zhao, Shiqi Zhao 0001, Siyang Song, Yuichiro Yoshikawa, Björn W. Schuller, Hiroshi Ishiguro
ACM Multimedia11
2025 BoostViT: Booth-Serial Skipping and Tunable Scaling for Vision Transformers
abstract
Vision Transformers (ViTs) have emerged as a dominant architecture in computer vision (CV), surpassing conventional neural network counterparts across diverse visual tasks. Despite their exceptional performance, ViTs incur substantial computational overhead characterized by high memory footprint, long inference latency, and elevated energy consumption. Current acceleration strategies for ViTs primarily focus on pruning operations or leveraging the inherent sparsity, requiring complex address control logic or position encoding. Alternatively, some software-based approaches attempt to pre-compute and separate dense and sparse matrix position encoding, while hardware solutions typically spend additional time and resources to obtain position encoding, decomposing matrix multiplications into structured forms. Through an analysis of ViTs’ parameters, we found that approximately 91.03% of the most significant bits (MSBs) are either 111s or 000s, and nearly 45% of 3 adjacent bits are identical. To leverage this characteristic of ViTs, we propose the Booth-Serial Skipping algorithm, which transforms the computation of consecutive 111 or 000 sequences into skip steps that require no additional computation time. Furthermore, the 4th to 6th bits of ViT weights can undergo aggressive scaling, enhancing the likelihood of Booth-skip operations with minimal impact on accuracy. The key innovation of this paper lies in exploiting the high proportion of naturally consecutive 0s or 1s in 8-bit weights during ViT inference and further expanding the skippable range through the Tunable Scaling strategy. At the hardware level, we develop a specialized accelerator to coordinate the proposed acceleration strategies. The processing element array in the accelerator is optimized for general matrix multiplication, it not only significantly improves the computation of multi-head self-attention but also enables resource reuse for linear transformations, ultimately optimizing end-to-end inference. Our design achieves$50.3\times $,$21.9\times $,$17.37\times $,$7.47\times $, and$1.49\times $an average end-to-end speedup on DeiT over CPU (Intel Xeon Gold 6152), EdgeGPU (NVIDIA Jetson Xavier NX), GPU (TITAN Xp), ViTCoD, and ViT-slice, respectively.
Shiqi Zhao 0001, Chaoming Fang, Fengshi Tian, Jinbo Chen 0002, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 A New Neuromorphic Computing Approach for Epileptic Seizure Prediction
abstract
Several high specificity and sensitivity seizure prediction methods with convolutional neural networks (CNNs) are reported. However, CNNs are computationally expensive and power hungry. These inconveniences make CNN-based methods hard to be implemented on wearable devices. Motivated by the energy-efficient spiking neural networks (SNNs), a neuromorphic computing approach for seizure prediction is proposed in this work. This approach uses a designed gaussian random discrete encoder to generate spike sequences from the EEG samples and make predictions in a spiking convolutional neural network (Spiking-CNN) which combines the advantages of CNNs and SNNs. The experimental results show that the sensitivity, specificity and AUC can remain 95.1%, 99.2% and 0.912 respectively while the computation complexity is reduced by 98.58% compared to CNN, indicating that the proposed Spiking-CNN is hardware friendly and of high precision.
Fengshi Tian, Jie Yang 0033, Shiqi Zhao 0001, Mohamad Sawan
ISCAS3
2020 Binary Single-Dimensional Convolutional Neural Network for Seizure Prediction
abstract
Nowadays, several deep learning methods are proposed to tackle the challenge of epileptic seizure prediction. However, these methods still cannot be implemented as part of implantable or efficient wearable devices due to their large hardware and corresponding high-power consumption. They usually require complex feature extraction process, large memory for storing high precision parameters and complex arithmetic computation, which greatly increases required hardware resources. Moreover, available yield poor prediction performance, because they adopt network architecture directly from image recognition applications fails to accurately consider the characteristics of EEG signals. We propose in this paper a hardware-friendly network called Binary Single-dimensional Convolutional Neural Network (BSDCNN) intended for epileptic seizure prediction. BSDCNN utilizes 1D convolutional kernels to improve prediction performance. All parameters are binarized to reduce the required computation and storage, except the first layer. Overall area under curve, sensitivity, and false prediction rate reaches 0.915, 89.26%, 0.117/h and 0.970, 94.69%, 0.095/h on American Epilepsy Society Seizure Prediction Challenge (AES) dataset and the CHB-MIT one respectively. The proposed architecture outperforms recent works while offering 7.2 and 25.5 times reductions on the size of parameter and computation, respectively.
Shiqi Zhao 0001, Jie Yang 0033, Yankun Xu, Mohamad Sawan
ISCAS1