EDBT 2026 Demo / reviewers in the wild / expert
Chih-Kai Kang
dblp:59/5211
· DBLP profile ↗
12ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0001-8258-1993ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Intermittent-Friendly Neural Architecture Search: Demystifying Accuracy and Overhead TradeoffsabstractThe fusion of tiny energy harvesting devices with deep neural networks (DNN) optimized for intermittent execution is vital for sustainable intelligent applications at the edge. However, current intermittent-aware neural architecture search (NAS) frameworks overlook the inherent intermittency management overhead (IMO) of DNNs, leading to under-performance upon deployment. Moreover, we observe that straightforward IMO minimization within NAS may degrade solution accuracy. This work explores the relationship between DNN architectural characteristics, IMO, and accuracy, uncovering the varying sensitivity toward IMO across different DNN characteristics. Inspired by our insights, we present two guidelines for leveraging IMO sensitivity in NAS. First, the overall architecture search space can be reduced to exclude parameters with low IMO sensitivity, and second, network blocks with high IMO sensitivity can be primarily focused during the search, facilitating the discovery of highly accurate networks with low IMO. We incorporate these guidelines into TiNAS, which integrates cutting-edge tiny NAS and intermittent-aware NAS frameworks. Evaluations are conducted across various datasets and latency requirements, as well as deployment experiments on a Texas Instruments device under different intermittent power profiles. Compared to two variants, one minimizing IMO and the other disregarding IMO, TiNAS, respectively, achieves up to 38% higher accuracy and 33% lower IMO, with greater improvements for larger datasets. Its deployed solutions also achieve up to a 1.33 times inference speedup, especially under fluctuating power conditions. Hashan R. Mendis, Chih-Hsuan Yen, Chih-Kai Kang, Pi-Cheng Hsiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Deep Reorganization: Retaining Residuals in TinyMLabstractDesigning intelligent, tiny devices with limited memory is immensely challenging, exacerbated by the additional memory requirement of residual connections in deep neural networks. In contrast to existing approaches that eliminate residuals to reduce peak memory usage at the cost of significant accuracy degradation, this paper presents DERO, which reorganizes residual connections by leveraging insights into the types and interdependencies of operations across residual connections. Evaluations were conducted across diverse model architectures designed for common computer vision applications. DERO consistently achieves peak memory usage comparable to plain-style models without residuals, while closely matching the accuracy of the original models with residuals. Hashan R. Mendis, Chih-Kai Kang, Chun-Han Lin, Ming-Syan Chen, Pi-Cheng Hsiu |
DAC | 2 |
| 2022 | More Is Less: Model Augmentation for Intermittent Deep InferenceabstractEnergy harvesting creates an emerging intermittent computing paradigm but poses new challenges for sophisticated applications such as intermittent deep neural network (DNN) inference. Although model compression has adapted DNNs to resource-constrained devices, under intermittent power, compressed models will still experience multiple power failures during a single inference. Footprint-based approaches enable hardware-accelerated intermittent DNN inference by tracking footprints, independent of model computations, to indicate accelerator progress across power cycles. However, we observe that the extra overhead required to preserve progress indicators can severely offset the computation progress accumulated by intermittent DNN inference. This work proposes the concept of model augmentation to adapt DNNs to intermittent devices. Our middleware stack, JAPARI, appends extra neural network components into a given DNN, to enable the accelerator to intrinsically integrate progress indicators into the inference process, without affecting model accuracy. Their specific positions allow progress indicator preservation to be piggybacked onto output feature preservation to amortize the extra overhead, and their assigned values ensure uniquely distinguishable progress indicators for correct inference recovery upon power resumption. Evaluations on a Texas Instruments device under various DNN models, capacitor sizes, and progress preservation granularities show that JAPARI can speed up intermittent DNN inference by 3× over the state of the art, for common convolutional neural architectures that require heavy acceleration. Chih-Kai Kang, Hashan R. Mendis, Chun-Han Lin, Ming-Syan Chen, Pi-Cheng Hsiu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | Intermittent-Aware Neural Architecture SearchabstractThe increasing paradigm shift towards i ntermittent computing has made it possible to intermittently execute d eep neural network (DNN) inference on edge devices powered by ambient energy. Recently, n eural architecture search (NAS) techniques have achieved great success in automatically finding DNNs with high accuracy and low inference latency on the deployed hardware. We make a key observation, where NAS attempts to improve inference latency by primarily maximizing data reuse, but the derived solutions when deployed on intermittently-powered systems may be inefficient, such that the inference may not satisfy an end-to-end latency requirement and, more seriously, they may be unsafe given an insufficient energy budget. This work proposes iNAS, which introduces intermittent execution behavior into NAS to find accurate network architectures with corresponding execution designs, which can safely and efficiently execute under intermittent power. An intermittent-aware execution design explorer is presented, which finds the right balance between data reuse and the costs related to intermittent inference, and incorporates a preservation design search space into NAS, while ensuring the power-cycle energy budget is not exceeded. To assess an intermittent execution design, an intermittent-aware abstract performance model is presented, which formulates the key costs related to progress preservation and recovery during intermittent inference. We implement iNAS on top of an existing NAS framework and evaluate their respective solutions found for various datasets, energy budgets and latency requirements, on a Texas Instruments device. Compared to those NAS solutions that can safely complete the inference, the iNAS solutions reduce the intermittent inference latency by 60% on average while achieving comparable accuracy, with an average 7% increase in search overhead. Hashan R. Mendis, Chih-Kai Kang, Pi-Cheng Hsiu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | Everything Leaves Footprints: Hardware Accelerated Intermittent Deep InferenceabstractCurrent peripheral execution approaches for intermittently powered systems require full access to the internal hardware state for checkpointing or rely on application-level energy estimation for task partitioning to make correct forward progress. Both requirements present significant practical challenges for energy-harvesting, intelligent edge Internet-of-Things devices, which perform hardware-accelerated deep neural network (DNN) inference. Sophisticated compute peripherals may have an inaccessible internal state, and the complexity of DNN models makes it difficult for programmers to partition the application into suitably sized tasks that fit within an estimated energy budget. This article presents the concept of inference footprinting for intermittent DNN inference, where accelerator progress is accumulatively preserved across power cycles. Our middleware stack, HAWAII, tracks and restores inference footprints efficiently and transparently to make inference forward progress, without requiring access to the accelerator internal state and application-level energy estimation. Evaluations were carried out on a Texas Instruments device, under varied energy budgets and network workloads. Compared to a variety of task-based intermittent approaches, HAWAII improves the inference throughput by 5.7%-95.7%, particularly achieving higher performance on heavily accelerated DNNs. Chih-Kai Kang, Hashan R. Mendis, Chun-Han Lin, Ming-Syan Chen, Pi-Cheng Hsiu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Quality-Enhanced OLED Power Savings on Mobile DevicesabstractIn the future, mobile systems will increasingly feature more advanced organic light-emitting diode (OLED) displays. The power consumption of these displays is highly dependent on the image content. However, existing OLED power-saving techniques either change the visual experience of users or degrade the visual quality of images in exchange for a reduction in the power consumption. Some techniques attempt to enhance the image quality by employing a compound objective function. In this article, we present a win-win scheme that always enhances the image quality while simultaneously reducing the power consumption. We define metrics to assess the benefits and cost for potential image enhancement and power reduction. We then introduce algorithms that ensure the transformation of images into their quality-enhanced power-saving versions. Next, the win-win scheme is extended to process videos at a justifiable computational cost. All the proposed algorithms are shown to possess the win-win property without assuming accurate OLED power models. Finally, the proposed scheme is realized through a practical camera application and a video camcorder on mobile devices. The results of experiments conducted on a commercial tablet with a popular image database and on a smartphone with real-world videos are very encouraging and provide valuable insights for future research and practices. Chun-Han Lin, Chih-Kai Kang, Pi-Cheng Hsiu |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | HomeRun: HW/SW Co-Design for Program Atomicity on Self-Powered Intermittent SystemsabstractSelf-powered intermittent systems featuring nonvolatile processors (NVPs) allow for accumulative execution in unstable power environments. However, frequent power failures may cause incorrect NVP execution results due to invalid data generated intermittently. This paper presents a HW/SW co-design, called HomeRun, to guarantee atomicity by ensuring that an uninterruptible program section can be run through at one execution. We design a HW module to ensure that a power pulse is sufficient for an atomic section, and develop a SW mechanism for programmers to protect atomic sections. The proposed design is validated through the development of a prototype pattern locking system. Experimental results demonstrate that the proposed design can completely guarantee atomicity and significantly improve the energy utilization of self-powered intermittent systems. Chih-Kai Kang, Chun-Han Lin, Pi-Cheng Hsiu, Ming-Syan Chen |
ISLPED | 1 |
| 2016 | CURA: A Framework for Quality-Retaining Power Saving on Mobile OLED DisplaysabstractOrganic Light-Emitting Diode (OLED) technology is regarded as a promising alternative to mobile displays. In this article, we introduce the design, algorithm, and implementation of a novel framework called CURA for quality-retaining power saving on mobile OLED displays. First, we link human visual attention to OLED power saving and model the OLED image scaling optimization problem. The objective is to minimize the power required to display an image without adversely impacting the user’s visual experience. Then, we present the algorithm used to solve the modeled problem, and prove its optimality even without an accurate power model. Finally, based on the framework, we implement two practical applications on a commercial OLED mobile tablet. The results of experiments conducted on the tablet with real images demonstrate that CURA can reduce significant OLED power consumption while retaining the visual quality of images. Chun-Han Lin, Chih-Kai Kang, Pi-Cheng Hsiu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | A win-win camera: Quality-enhanced power-saving images on mobile OLED displaysabstractMobile systems will increasingly feature emerging OLED displays, whose power consumption is highly dependent on the image content. Existing OLED power-saving techniques change users' visual experience or degrade images' visual quality in exchange for power reduction, or seek a chance to also enhance image quality by employing a compound objective function. This paper presents a win-win scheme that always enhances image quality and reduces power consumption simultaneously. We define metrics to assess the profit and the cost for potential image enhancement and power reduction. Then, we propose algorithms that ensure the transformation of images into their quality-enhanced power-saving versions. Finally, the proposed scheme is realized as a practical camera application on mobile devices. The results of experiments conducted on a commercial tablet with a popular image database are very encouraging and provide valuable insights for future research and practices. Chih-Kai Kang, Chun-Han Lin, Pi-Cheng Hsiu |
ISLPED | 1 |
| 2014 | Catch Your Attention: Quality-retaining Power Saving on Mobile OLED DisplaysabstractOrganic light-emitting diode (OLED) technology is considered as a promising alternative to mobile displays. This paper explores how to reduce the OLED power consumption by exploiting visual attention. First, we model the problem of OLED image scaling optimization, with the objective of minimizing the power required to display an image without adversely impacting the user's visual experience. Then, we propose an algorithm to solve the fundamental problem, and prove its optimality even without the accurate power model. Finally, based on the algorithm, we consider implementation issues and realize two application scenarios on a commercial OLED mobile tablet. The results of experiments conducted on the tablet with real images demonstrate that the proposed methodology can achieve significant power savings while retaining the visual quality. Chun-Han Lin, Chih-Kai Kang, Pi-Cheng Hsiu |
DAC | 2 |
| 2014 | A Hybrid Storage Access Framework for High-Performance Virtual MachinesabstractIn recent years, advances in virtualization technology have enabled multiple virtual machines to run on a physical machine, such that each virtual machine can perform independently with its own operating system. The IT industry has adopted virtualization technology because of its ability to improve hardware resource utilization, achieve low-power consumption, support concurrent applications, simplify device management, and reduce maintenance costs. However, because of the hardware limitation of storage devices, the I/O capacity could cause performance bottlenecks. To address the problem, we propose a hybrid storage access framework that exploits solid-state drives (SSDs) to improve the I/O performance in a virtualization environment. Chih-Kai Kang, Yu-Jhang Cai, Chin-Hsien Wu, Pi-Cheng Hsiu |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | A hybrid storage access framework for virtual machinesabstractIn recent years, virtualization technology enables multiple virtual machines to run on a physical machine, where each virtual machine can run independently and own its operating system. Virtualization technology has been adopted in many IT industries because of its ability to improve hardware resource utilization, achieve low-power consumption, simplify server management, and reduce maintenance cost. However, since the hardware limitation of storage devices, I/O capacity could cause performance bottleneck. In the paper, we will propose a hybrid storage access framework for virtualization environment to dynamically adjust and enhance I/O performance by using solid-state drives (SSDs). Chih-Kai Kang, Yu-Jhang Cai, Chin-Hsien Wu, Pi-Cheng Hsiu |
RTCSA | 1 |