Zhenge Jia

dblp:235/0320 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-0554-3608ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SysVCoder: An LLM-Driven Framework for Systematic Generation of System-Level Design
Jian Zuo, Junzhe Liu, Xianyong Wang, Navya Goli, Umamaheswara Rao Tida, Zhenge Jia, Zhaoyan Shen, Mengying Zhao
APPT7
2025 Exploiting Large Language Models for Software-Defined Solid-State Drives Design
Tianren Zhou, Zhenge Jia, Mengying Zhao, Zhaoyan Shen
APPT5
2025 A 10.60 μW 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia Detection
abstract
This paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves 50% sparsity in a quantized 1D CNN using a sparse processing element (SPE) architecture. Measurement on the prototype chip TSMC 40nm CMOS low-power (LP) process for the VA classification task demonstrates that it consumes 10.60 μW of power while achieving a performance of 150 GOPS and a diagnostic accuracy of 99.95%. The computation power density is only 0.57 μW/mm2, which is 14.23× smaller than state-of-the-art works, making it highly suitable for implantable and wearable medical devices.
Zhenge Jia, Zheyu Yan, Jay Mok, Manto Yung, Yu Liu 0007, Wujie Wen, Luhong Liang, Kwang-Ting Cheng, Xiaobo Sharon Hu, Yiyu Shi 0001
ASP-DAC2
2025 Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment Perspective
abstract
Reconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improve the task accuracy through architectural or algorithmic innovations, we tackle this task from image quality assessment (IQA) perspective, an under-explored direction in the field. Due to the limitations of conventional metrics such as £1 in capturing the nuanced differences in reconstructed images for medical anomaly detection, we propose fusion quality, a novel metric that wisely integrates the structure-level sensitivity of Structural Similarity Index Measure (SSIM) with the pixel-level precision of £1. The metric offers a more comprehensive assessment of reconstruction quality, considering intensity (subtractive property of l1and divisive property of SSIM), contrast, and structural similarity. Furthermore, the proposed metric makes subtle regional variations more impactful in the final assessment. Thus, considering the inherent divisive properties of SSIM, we design an average intensity ratio (AIR)-based data transformation that amplifies the divisive discrepancies between normal and abnormal regions, thereby enhancing anomaly detection. By fusing the aforementioned two components, we devise the IQA approach. Experimental results on two distinct brain MRI datasets show that our IQA approach significantly enhances medical anomaly detection performance when integrated with state-of-the-art baselines. Code is provided here.
Zixuan Pan, Jun Xia 0003, Zheyu Yan, Guoyue Xu, Yawen Wu, Zhenge Jia, Jianxu Chen 0001, Yiyu Shi 0001
BIBM8
2025 Enabling On-Tiny-Device Model Personalization via Gradient Condensing and Alternant Partial Update
abstract
On-device training enables the model to adapt to user-specific data by fine-tuning a pre-trained model locally. As embedded devices become ubiquitous, on-device training is increasingly essential since users can benefit from the personalized model without transmitting data and model parameters to the server. Despite significant efforts toward efficient training, ondevice training still faces a major challenge: The prohibitive cost of multi-layer backpropagation strains the limited resources of tiny devices. In this paper, we propose an algorithm-system cooptimization framework TinyMP that enables self-adaptive on-tiny-device model personalization. To mitigate backpropagation costs, we introduce Gradient Condensing to condense the gradient map structure, significantly reducing the computational complexity and memory consumption of backpropagation while preserving model performance. To further reduce computation overhead, we propose Alternant Partial Update, a mechanism that locally and alternatively selects essential parameters to update without requiring retraining or offline evolutionary search. Our framework is evaluated through extensive experiments using various CNN models (e.g., MobileNetV2, MCUNet) on embedded devices with minimal resources (e.g., OpenMV-H7 with less than 1MB SRAM and 2 MB Flash). Experimental results show that our framework achieves up to $2.4 \times$ speedup, 80.8% memory saving, and 30.3% accuracy improvement on downstream tasks, outperforming SOTA approaches.
Zhenge Jia, Yiyang Shi, Zeyu Bao, Xin Pang, Huiguo Liu, Zhaoyan Shen, Mengying Zhao
DAC1
2025 Routability-aware Packing for High-density Nonvolatile FPGAs
abstract
Nonvolatile field-programmable gate arrays (NVFPGAs) can use multi-level cell (MLC) nonvolatile memories (NVMs) to enhance their logic density. However, the highdensity design of NVFPGAs degrades the intra-routability of configurable logic blocks (CLBs), which significantly prolongs the time consumed by the packing process in the computer-aided design (CAD) flow. To relieve the efficiency degradation, in this paper, we propose a routability-aware re-pair stage to adjust the logical-physical look-up table (LUT) assignments to mitigate the congestion and improve their intra-routability, thereby reducing the packing time. In addition, exploiting the structural equivalence of MLC LUTs, we remove unnecessary intra-routing attempts from packing to further improve efficiency. Evaluation shows the proposed strategies reduce packing time by $41.48 \%$ on average. Index Terms-nonvolatile memory (NVM), multi-level cell (MLC), field-programmable gate array (FPGA), computer-aided design (CAD), packing.
Huichuan Zheng, Yuqing Xiong, Jian Zuo, Zhenge Jia, Mengying Zhao
DAC5
2025 DiffECG: Diffusion Model-Powered Label-Efficient and Personalized Arrhythmia Diagnosis
abstract
Arrhythmia diagnosis using electrocardiogram (ECG) is critical for preventing cardiovascular risks. However, existing deep learning-based methods struggle with label scarcity and contrastive learning-based methods suffer from false-negative samples, which lead to poor model generalization. Besides, due to inter-subject variability, pre-trained models cannot achieve evenly performance across individuals. Conducting model fine-tuning for each individual is computationally expensive and does not guarantee improvement. We propose DiffECG, a diffusion-based self-supervised learning framework for label-efficient and personalized arrhythmia detection. Our method utilizes a diffusion model to extract robust ECG representations, coupled with a novel feature extractor and a multi-modal feature fusion strategy to obtain a well-generalized model. Moreover, we propose an efficient model personalization mechanism based on zeroth-order optimization. It personalizes the model by tuning the noise-adding step t in the diffusion process, significantly reducing computational costs compared to model fine-tuning. Experimental results show that our proposed method outperforms the SOTA method by 37.9% and 23.9% in generalization and personalization performance, respectively. The source code is available at: https://github.com/Auguuust/DiffEC
Tianren Zhou, Zhenge Jia, Dongxiao Yu, Zhaoyan Shen
IJCAI2
2025 QC-CNN: Highly Quantized Compressive CNN for Efficient Ventricular Arrhythmia Detection in Implantable Cardioverter Defibrillators
abstract
The Implantable Cardioverter Defibrillator (ICD) is a device designed to reduce the risk of Sudden Cardiac Death (SCD) by detecting life-threatening ventricular arrhythmias (VAs) through intracardiac electrograms (IEGMs). Ensuring accurate VA detection across diverse patient populations remains a challenge, as traditional ICDs either rely on manual tuning of patient-specific parameters or suffer from low detection accuracy. While machine learning models, particularly convolutional neural networks (CNNs), have demonstrated improved accuracy and flexibility, their large model size and high sampling frequency make them hard to be deployed in ICDs due to limited memory and energy capacity on ICDs. To address these constraints, we propose a compressive sensing (CS)-inspired CNN architecture that reduces model size by 50× and reduce the sample frequency by 25×. By reconstructing signals with fewer measurements and applying low-bit quantization, our model is optimized for efficient execution on analog domain frontend ADC. Experimental results show our model could use only 3% power consumption of the classic approach on ADC with no accuracy loss.
Zhenge Jia, Yiyu Shi 0001
ISCAS4
2025 Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
abstract
The scaling laws have become the de facto guidelines for designing large language models (LLMs), but they were studied under the assumption of unlimited computing resources for both training and inference. As LLMs are increasingly used as personalized intelligent assistants, their customization (i.e., learning through fine-tuning) and deployment onto resource-constrained edge devices will become more and more prevalent. An urgent but open question is how a resource-constrained computing environment would affect the design choices for a personalized LLM. We study this problem empirically in this work. In particular, we consider the tradeoffs among a number of key design factors and their intertwined impacts on learning efficiency and accuracy. The factors include the learning methods for LLM customization, the amount of personalized data used for learning customization, the types and sizes of LLMs, the compression methods of LLMs, the amount of time afforded to learn, and the difficulty levels of the target use cases. Through extensive experimentation and benchmarking, we draw a number of surprisingly insightful guidelines for deploying LLMs onto resource-constrained devices. For example, an optimal choice between parameter learning and RAG may vary depending on the difficulty of the downstream task, the longer fine-tuning time does not necessarily help the model, and a compressed LLM may be a better choice than an uncompressed LLM to learn from limited personalized data.
Ruiyang Qin, Dancheng Liu, Chenhui Xu, Zheyu Yan, Zhaoxuan Tan, Zhenge Jia, Amir Nassereldine, Jiajie Li 0002, Meng Jiang 0001, Ahmed Abbasi, Jinjun Xiong, Yiyu Shi 0001
ACM Trans. Design Autom. Electr. Syst.6
2024 Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and Synthesis
abstract
After a large language model (LLM) is deployed on edge devices, it is desirable for these devices to learn from user-generated conversation data to generate user-specific and personalized responses in real-time. However, user-generated data usually contains sensitive and private information, and uploading such data to the cloud for annotation is not preferred if not prohibited. While it is possible to obtain annotation locally by directly asking users to provide preferred responses, such annotations have to be sparse to not affect user experience. In addition, the storage of edge devices is usually too limited to enable large-scale fine-tuning with full user-generated data. It remains an open question how to enable on-device LLM personalization, considering sparse annotation and limited on-device storage. In this paper, we propose a novel framework to select and store the most representative data online in a self-supervised way. Such data has a small memory footprint and allows infrequent requests of user annotations for further fine-tuning. To enhance fine-tuning quality, multiple semantically similar pairs of question texts and expected responses are generated using the LLM. Our experiments show that the proposed framework achieves the best user-specific content-generating capability (accuracy) and fine-tuning speed (performance) compared with vanilla baselines. To the best of our knowledge, this is the very first on-device LLM personalization framework.
Ruiyang Qin, Jun Xia 0003, Zhenge Jia, Meng Jiang 0001, Ahmed Abbasi, Peipei Zhou 0001, Jingtong Hu, Yiyu Shi 0001
DAC3
2024 Federation-Paced Learning: Towards Efficient Federated Learning with Synchronized Pace
abstract
Federated learning (FL) is a distributed machine learning approach that allows multiple devices or computing nodes to jointly train models without sharing raw data. However, in real-world application scenarios, FL usually encounters a critical challenge of data heterogeneity. Recent studies have revealed that the client’s model suffers severe bias between the local model and global model, leading to global performance degradation. Improving the generalization of local learning would inherently reduce bias. It has been proved that self-paced learning on a single device can greatly achieve a better generalization result. However, it is not well explored how it can be applied to federated learning with a number of distributed nodes working cooperatively. Specifically, self-paced learning suggests using easy data and then gradually difficult data during model training. It is not straightforward to differentiate “easy” and “difficult” data at the local since global data distribution is not available, especially with severe data heterogeneity. To address the above issues, we propose a novel federated learning framework, Federation-Paced Learning (FedPL), which enables a self-paced process in federated learning and effectively improves the model performance. First, we propose schemes to analyze the data characteristics in terms of difficulty. Then we define a stage controller to synchronize the learning process across cooperative nodes to follow the easy-to-hard rule. Finally, we propose a client selection strategy to further improve the learning efficacy. We evaluate the performance of FedPL on several generic public datasets. Experiment results show that the proposed FedPL outperforms existing methods by up to 13.50% in terms of accuracy. Code is available at https://github.com/tnghua/FedPL.
Mei Cao, Zhenge Jia, Jianbo Lu 0001, Zhaoyan Shen, Dongxiao Yu, Mengying Zhao
ECAI3
2024 Towards Uncertainty-Quantifiable Biomedical Intelligence: Mixed-signal Compute-in-Entropy for Bayesian Neural Networks
abstract
To enhance AI robustness of mission-critical biomedical applications, Bayesian Neural Networks (BNNs) are instrumental for their structured approach to AI uncertainty estimation. However, implementing BNNs on edge devices is challenging due to significant resource demands for dynamic model updates and extensive inference sampling. Addressing this, we introduce a novel mixed-signal Compute-in-Memory with Entropy (CIE) hardware architecture that segregates dynamically-generated weights into analog entropy and digital parameters within a compute-in-memory framework, greatly reducing hardware overhead. We conducted thorough evaluations of the CIE architecture, assessing its performance against varying hardware imperfections, such as digital quantization errors, analog distribution imperfections, and device process variations, with a focus on both general and specialized tasks like Ventricular Arrhythmia (VA) detection. Our contributions include (1) a generic BNN acceleration strategy suitable for various CIM techniques and emerging devices, (2) a custom circuit design that improves hardware efficiency by 19.2×-440× compared to existing BNN accelerators, (3) a CIE-based BNN for VA detection enhancing accuracy, reducing uncertainty estimation time and energy/latency to 1.29μJ/1.55ms, and (4) identification of tolerable quantization error and device variation limits for BNNs in uncertainty estimation.
Likai Pei, Zephan M. Enciso, Boyang Cheng, Steven Davis, Zhenge Jia, Michael T. Niemier, Yiyu Shi 0001, Xiaobo Sharon Hu, Ningyuan Cao
ICCAD7
2024 Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
abstract
Large Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG.
Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001
ICCAD4
2024 FairQuantize: Achieving Fairness Through Weight Quantization for Dermatological Disease Diagnosis
Zhenge Jia, Jingtong Hu, Yiyu Shi 0001
MICCAI (10)2
2024 TinyML Design Contest for Life-Threatening Ventricular Arrhythmia Detection
abstract
The first ACM/IEEE TinyML Design Contest (TDC) held at the 41st International Conference on Computer-Aided Design (ICCAD) in 2022 is a challenging, multimonth, research and development competition. TDC’22 focuses on real-world medical problems that require the innovation and implementation of artificial intelligence/machine learning (AI/ML) algorithms on implantable devices. The challenge problem of TDC’22 is to develop a novel AI/ML-based real-time detection algorithm for life-threatening ventricular arrhythmia (VA) over low-power microcontrollers utilized in implantable cardioverter-defibrillators (ICDs). The dataset contains more than 38000 5-s intracardiac electrograms (IEGMs) segments over eight different types of rhythm from 90 subjects. The dedicated hardware platform is NUCLEO-L432KC manufactured by STMicroelectronics. TDC’22, which is open to multiperson teams world-wide, attracted more than 150 teams from over 50 organizations. This article first presents the medical problem, dataset, and evaluation procedure in detail. It further demonstrates and discusses the designs developed by the leading teams as well as representative results. This article concludes with the direction of improvement for the future TinyML design for health monitoring applications.
Zhenge Jia, Dawei Li 0012, Liqi Liao, Xiaowei Xu 0004, Lichuan Ping, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Personalized Meta-Federated Learning for IoT-Enabled Health Monitoring
abstract
Federated learning (FL) has been widely adopted in IoT-enabled health monitoring on biosignals thanks to its advantages in data privacy preservation. However, the global model trained from FL generally performs unevenly across subjects since biosignal data is inherent with complex temporal dynamics. The morphological characteristics of biosignals with the same label can vary significantly among different subjects (i.e., inter-subject variability) while biosignals with varied temporal patterns can be collected on the same subject (i.e., intra-subject variability). To address the challenges, we present the Personalized Meta-Federated learning (PMFed) framework for personalized IoT-enabled health monitoring. Specifically, in the federated learning stage, a novel momentum-based model aggregating strategy is introduced to aggregate clients' models based on domain similarity in the meta-federated learning paradigm to obtain a well-generalized global model while speeding up the convergence. In the model personalizing stage, an adaptive model personalization mechanism is devised to adaptively tailor the global model based on the subject-specific biosignal features while preserving the learned cross-subject representations. We develop an IoT-enabled computing framework to evaluate the effectiveness of PMFed over three real-world health monitoring tasks. Experimental results show that the PMFed excels at detection performances in terms of F1 and accuracy by up to 9.4% and 8.7%, and reduces training overhead and throughput by up to 56.3% and 63.4% when compared with the SOTA federated learning algorithms.
Zhenge Jia, Tianren Zhou, Zheyu Yan, Jingtong Hu, Yiyu Shi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Opportunistic Communication with Latency Guarantees for Intermittently-Powered Devices
abstract
Energy-harvesting wireless sensor nodes have found widespread adoption due to their low cost and small form factor. However, uncertainty in the available power supply introduces significant challenges in engineering communications between intermittently-powered nodes. We propose a constraint-based model for energy harvests that together with a hardware model can be used to enable consistent, opportunistic communication with worst-case latency guarantees. We show that greedy approaches that attempt communication whenever energy is available lead to prolonged latencies in real-world environments. Our approach offers bounded worst-case latency while providing a performance improvement over a conservative, offline approach planned around the worst-case energy harvest.
Kacper Wardega, Wenchao Li 0001, Hyoseung Kim 0001, Yawen Wu, Zhenge Jia, Jingtong Hu
DATE5
2022 Personalized Neural Network for Patient-Specific Health Monitoring in IoT: A Metalearning Approach
abstract
The Internet of Things (IoT) has been widely applied in personal health monitoring on biosignals. Conventional detection methods in the field count on a variety of heuristic criteria by utilizing extracted features, which are carefully selected through extensive clinical trials and experts’ experiences. Recently, deep learning (DL) gains rapidly growing attention in health monitoring. The most significant advantage of DL-based methods is that DL could execute feature engineering automatically with only labeled data, which results in a great reduction in the expertise involved and manual works in the detection method’s design. However, individual differences among various patients (subjects) can lead to accuracy degradation of the pretrained deep model. Simply fine-tuning the deep model with the patient-specific data cannot alleviate the problem since the pretrained model may not generalize well to new data. To address the problem, we propose a metalearning-based personalization method to generate the personalized neural network for each patient to conduct patient-specific detection. Specifically, the proposed metalearning method leverages a novel patientwise training tasks formatting strategy to train the neural network that ends up with a well-generalized model initialization containing across-patient knowledge. The well-generalized model initialization would then be utilized to perform a quick adaptation to the specific patient’s data domain. In this way, a new patient could be immediately assigned with a personalized neural network using limited labeled data. Experimental results show that the proposed metalearning-based personalization method achieves 8.2%, 2.5%, and 6.4% higher accuracy when compared with the existing DL detection methods in VF detection, AF detection, and human activity recognition, respectively.
Zhenge Jia, Yiyu Shi 0001, Jingtong Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Cooperative Communication Between Two Transiently Powered Sensor Nodes by Reinforcement Learning
abstract
Energy harvesting (EH)-powered sensor nodes can achieve theoretically unlimited lifetime by scavenging energy from ambient power sources, such as radio-frequency (RF) and kinetic energy. The nodes can collect and transmit data wirelessly with the harvested energy. However, the transmission between two sensor nodes is successful only when both nodes have enough energy at the same time. While the receiver can be actively listening, it may deplete the energy long before the sender has accumulated enough energy. Thus, given the scarce, unpredictable, and unevenly distributed energy among sensor nodes, it is challenging to ensure efficient data transmission between them. To address this challenge, we propose a sensor node architecture with multiple radios, each with different energy consumption on the sender and receiver. A node can be put into sleep when charged up and wakes up for communication when it infers that both nodes have enough energy based on its observations. What is more, two nodes can cooperatively and dynamically select different radios according to the stored energy and historical information to maximize the data throughput. To achieve cooperative communication adaptively, the communication procedure is modeled as a cooperative Markov game with partial observability on each node, and multiagent reinforcement learning (MARL) is employed to achieve the best results. Experimental results on hardware prototype and by simulation show that the proposed approaches achieve up to 89.1% of the optimal throughput and significantly outperform other online algorithms.
Yawen Wu, Zhenge Jia, Fei Fang 0001, Jingtong Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Lightweight Run-Time Working Memory Compression for Deployment of Deep Neural Networks on Resource-Constrained MCUs
abstract
This work aims to achieve intelligence on embedded devices by deploying deep neural networks (DNNs) onto resource-constrained microcontroller units (MCUs). Apart from the low frequency (e.g., 1-16 MHz) and limited storage (e.g., 16KB to 256KB ROM), one of the largest challenges is the limited RAM (e.g., 2KB to 64KB), which is needed to save the intermediate feature maps of a DNN. Most existing neural network compression algorithms aim to reduce the model size of DNNs so that they can fit into limited storage. However, they do not reduce the size of intermediate feature maps significantly, which is referred to as working memory and might exceed the capacity of RAM. Therefore, it is possible that DNNs cannot run in MCUs even after compression. To address this problem, this work proposes a technique to dynamically prune the activation values of the intermediate output feature maps in the runtime to ensure that they can fit into limited RAM. The results of our experiments show that this method could significantly reduce the working memory of DNNs to satisfy the hard constraint of RAM size, while maintaining satisfactory accuracy with relatively low overhead on memory and run-time latency.
Zhepeng Wang 0001, Yawen Wu, Zhenge Jia, Yiyu Shi 0001, Jingtong Hu
ASP-DAC3
2021 Enabling On-Device Model Personalization for Ventricular Arrhythmias Detection by Generative Adversarial Networks
abstract
Implantable Cardioverter Defibrillator (ICD) is an ultra-low-power device which monitors heart rate and delivers in-time defibrillation on detected ventricular arrhythmias (VAs). The parameters of VAs detection mechanism on each recipient’s ICD are supposed to be fine-tuned to obtain accurate detection due to the individual’s unique rhythm features. However, the process extremely relies on clinical expertise and thus must be conducted manually and routinely by cardiologists diagnosing massive amount of rhythm data. In this paper, we introduce a novel self-supervised on-device personalization of convolutional neural network (CNNs) for VAs detection. We first propose a computing framework consisting of an edge device and an ICD to enable efficient on-device CNNs personalization and real-time inference respectively. Then, we propose a generative model that learns to synthesize patient-specific intracardiac EGMs signals, which can then be used as personalized training data to improve patient-specific VAs detection performance on ICDs. Evaluations on three detection models show that the self-supervised on-device personalization significantly improve VAs detection performance under a patient-specific setting.
Zhenge Jia, Lichuan Ping, Yiyu Shi 0001, Jingtong Hu
DAC1
2021 Learning to Learn Personalized Neural Network for Ventricular Arrhythmias Detection on Intracardiac EGMs
abstract
Life-threatening ventricular arrhythmias (VAs) detection on intracardiac electrograms (IEGMs) is essential to Implantable Cardioverter Defibrillators (ICDs). However, current VAs detection methods count on a variety of heuristic detection criteria, and require frequent manual interventions to personalize criteria parameters for each patient to achieve accurate detection. In this work, we propose a one-dimensional convolutional neural network (1D-CNN) based life-threatening VAs detection on IEGMs. The network architecture is elaborately designed to satisfy the extreme resource constraints of the ICD while maintaining high detection accuracy. We further propose a meta-learning algorithm with a novel patient-wise training tasks formatting strategy to personalize the 1D-CNN. The algorithm generates a well-generalized model initialization containing across-patient knowledge, and performs a quick adaptation of the model to the specific patient's IEGMs. In this way, a new patient could be immediately assigned with personalized 1D-CNN model parameters using limited input data. Compared with the conventional VAs detection method, the proposed method achieves 2.2% increased sensitivity for detecting VAs rhythm and 8.6% increased specificity for non-VAs rhythm.
Zhenge Jia, Zhepeng Wang 0001, Lichuan Ping, Yiyu Shi 0001, Jingtong Hu
IJCAI1
2021 On-device Prior Knowledge Incorporated Learning for Personalized Atrial Fibrillation Detection
abstract
Atrial Fibrillation (AF), one of the most prevalent arrhythmias, is an irregular heart-rate rhythm causing serious health problems such as stroke and heart failure. Deep learning based methods have been exploited to provide an end-to-end AF detection by automatically extracting features from Electrocardiogram (ECG) signal and achieve state-of-the-art results. However, the pre-trained models cannot adapt to each patient’s rhythm due to the high variability of rhythm characteristics among different patients. Furthermore, the deep models are prone to overfitting when fine-tuned on the limited ECG of the specific patient for personalization. In this work, we propose a prior knowledge incorporated learning method to effectively personalize the model for patient-specific AF detection and alleviate the overfitting problems. To be more specific, a prior-incorporated portion importance mechanism is proposed to enforce the network to learn to focus on the targeted portion of the ECG, following the cardiologists’ domain knowledge in recognizing AF. A prior-incorporated regularization mechanism is further devised to alleviate model overfitting during personalization by regularizing the fine-tuning process with feature priors on typical AF rhythms of the general population. The proposed personalization method embeds the well-defined prior knowledge in diagnosing AF rhythm into the personalization procedure, which improves the personalized deep model and eliminates the workload of manually adjusting parameters in conventional AF detection method. The prior knowledge incorporated personalization is feasibly and semi-automatically conducted on the edge, device of the cardiac monitoring system. We report an average AF detection accuracy of 95.3% of three deep models over patients, surpassing the pre-trained model by a large margin of 11.5% and the fine-tuning strategy by 8.6%.
Zhenge Jia, Yiyu Shi 0001, Samir Saba, Jingtong Hu
ACM Trans. Embed. Comput. Syst.1
2020 Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered Devices
abstract
This work aims to enable persistent, event-driven sensing and decision capabilities for energy-harvesting (EH)-powered devices by deploying lightweight DNNs onto EH-powered devices. However, harvested energy is usually weak and unpredictable and even lightweight DNNs take multiple power cycles to finish one inference. To eliminate the indefinite long wait to accumulate energy for one inference and to optimize the accuracy, we developed a power trace-aware and exit-guided network compression algorithm to compress and deploy multi-exit neural networks to EH-powered microcontrollers (MCUs) and select exits during execution according to available energy. The experimental results show superior accuracy and latency compared with state-of-the-art techniques.
Yawen Wu, Zhepeng Wang 0001, Zhenge Jia, Yiyu Shi 0001, Jingtong Hu
DAC3
2020 Design Insights of Non-volatile Processors and Accelerators in Energy Harvesting Systems
abstract
There is growing interest in deploying energy harvesting processors and accelerators in Internet of Things (IoT). Energy harvesting harnesses the energy scavenged from the environment to power a system. Although it has many advantages over battery-operated systems such as lightweight, compact size, and no necessity of recharging and maintenance, it may suffer frequently power-down and a fluctuating power supply even with power on. Non-volatile processor (NVP) is a promising architecture for effective computing in energy harvesting scenarios. Recently, non-volatile accelerators (NVA) have been proposed to perform computations of deep learning algorithms. In this paper, we overview the recent studies of NVP and NVA across the layers of hardware, architecture, software and their co-design. Especially, we present the design insights of how the state-of-the-art works adapt their specific designs to the intermittent and fluctuating power conditions with the energy harvesting technology. Finally, we discuss recent trends using NVP and NVA in energy harvesting scenarios.
Keni Qiu, Mengying Zhao, Zhenge Jia, Jingtong Hu, Chun Jason Xue, Kaisheng Ma, Xueqing Li 0002, Yongpan Liu, Narayanan Vijaykrishnan
ACM Great Lakes Symposium on VLSI3
2020 Personalized Deep Learning for Ventricular Arrhythmias Detection on Medical loT Systems
abstract
Life-threatening ventricular arrhythmias (VA) are the leading cause of sudden cardiac death (SCD), which is the most significant cause of natural death in the US [6]. The implantable cardioverter defibrillator (ICD) is a small device implanted to patients under high risk of SCD as a preventive treatment. The ICD continuously monitors the intracardiac rhythm and delivers shock when detecting the life-threatening VA. Traditional methods detect VA by setting criteria on the detected rhythm. However, those methods suffer from a high inappropriate shock rate and require a regular follow-up to optimize criteria parameters for each ICD recipient. To ameliorate the challenges, we propose the personalized computing framework for deep learning based VA detection on medical IoT systems. The system consists of intracardiac and surface rhythm monitors, and the cloud platform for data uploading, diagnosis, and CNN model personalization. We equip the system with real-time inference on both intracardiac and surface rhythm monitors. To improve the detection accuracy, we enable the monitors to detect VA collaboratively by proposing the cooperative inference. We also introduce the CNN personalization for each patient based on the computing framework to tackle the unlabeled and limited rhythm data problem. When compared with the traditional detection algorithm, the proposed method achieves comparable accuracy on VA rhythm detection and 6.6% reduction in inappropriate shock rate, while the average inference latency is kept at 71ms.
Zhenge Jia, Zhepeng Wang 0001, Lichuan Ping, Yiyu Shi 0001, Jingtong Hu
ICCAD1
2018 Prototyping Energy Harvesting Powered Systems with Nonvolatile Processor (Invited Paper)
abstract
Energy harvesting is a promising solution to power ubiquitous Internet-of-Things (IoT) devices. But the frequent and inevitable power failure incurs significant backup overhead, greatly degrading performance and energy efficiency. Nonvolatile processor (NVP), which can checkpoint processor states, is designed to tackle this problem. The conventional system-level design method involves repeated system modification and verification on hardware, in which measurement on hardware consumes the majority time. To expedite the NVP-based system design process, we propose a rapid system prototyping flow to eliminate repeated hardware measurement in the design flow. This method involves an NVP system-level simulator, which takes the harvester power trace, system characteristics extracted from hardware, and user design as the input, and analyzes system energy and time profile under this power trace. Iterative system optimization and verification are conducted on the simulator, with only the final verification on hardware. We demonstrate the advantages of this method by two design cases, in which time, energy efficiency and the impact of different capacitor size are optimized.
Yawen Wu, Zhenge Jia, Lefan Zhang, Yongpan Liu, Jingtong Hu
RSP3