Chun-Hsian Huang

dblp:45/4118 · DBLP profile ↗
← Back
24ranked-venue papers
16as first author
6since 2021 · last 2025
0000-0002-0508-6312ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 12 first-author · 5 since 2021Software engineering, systems software and programming languages · 4Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Edge AI and Adaptive Embedded System Design for Agricultural Robotics Applications
abstract
This work presents the AgrBot, an agricultural robot designed to intelligently estimate and predict crop pest and disease severity (PDS). The AgrBot incorporates two binarized neural network (BNN) hardware modules for recognizing target crops and estimating their PDS. In a resource-constrained FPGA-based design, these BNN hardware modules can be configured on-demand, showcasing system adaptivity. Furthermore, a multimodal model that integrates crop images, sensor data, and time features is presented for predicting PDS. Employing edge artificial intelligence (AI) through the BNN hardware modules and the multimodal model enables the AgrBot to determine if biological agents are applied to protect crops from pests and diseases, creating a comprehensive agricultural cyber-physical system (CPS). Experimental results demonstrate accuracies of 76.3% for recognizing target crops, 65.3% for estimating PDS, and 67% for predicting PDS. In comparison to existing microprocessor-based design methods, the AgrBot's BNN hardware modules improve frames per second (FPS) by a factor of 790, while the multimodal model reduces processing time by up to 50.9%.
Chun-Hsian Huang, Zhi-Rui Chen, Huai-Shu Hsu
ASP-DAC1
2025 An Edge AI-Enabled UAV System based on FPGA for Water Quality Monitoring with Hyperspectral Imaging
abstract
This paper presents an edge AI-enabled UAV system based on FPGA for water quality monitoring, integrating hyperspectral imaging with onboard intelligence. The proposed system incorporates a band selection strategy to reduce the dimensionality of hyperspectral data and employs a dual convolutional neural network (CNN) architecture for robust and accurate river pollution index (RPI) classification. To meet the real-time and energy efficiency requirements in industrial environments, FPGA-based hardware acceleration is utilized. Experimental results show that the proposed dual CNN model improves RPI classification accuracy by 13% compared to the ResNet-50 model using only RGB images. Furthermore, the edge AI-enabled UAV system achieves a 2.97x improvement in energy efficiency compared to an embedded GPU-based design (NVIDIA Jetson Nano). These findings indicate that the proposed system offers a practical and scalable solution for intelligent monitoring and inspection in industrial areas.
Chun-Hsian Huang, Shu-Ting Huang, Jo-Lin Li, Tsiai-Jung Li
IECON1
2024 ACNNE: An Adaptive Convolution Engine for CNNs Acceleration Exploiting Partial Reconfiguration on FPGAs
abstract
In this work, we propose a dynamical function exchange Convolutional Neural Networks (CNN) accelerator architecture named Adaptive CNN engine (ACNNE) that can reconfigure specific convolution layer hardware blocks according to model parameters at runtime. We mainly focus on exploiting reconfigurability for inferencing large-scale CNN on resource- constrained FPGAs. The proposed ACNNE can accelerate the process of convolution layers based on a nested-loop algorithm while a data buffering scheme is presented to reduce the iterations of memory accesses. As a study case, the VGG16 model was implemented on Xilinx ZU3EG MPSoC that can achieve 21.09 giga operations per second in the frequency of 100 MHz with a device resource utilization of 25% to 35%. Experiments show a single-image inference can be completed in 2.5 seconds with an average power consumption of 2.23 W, corresponding to a power efficiency of 9.46 GOPS/W.
Chun-Hsian Huang, Shao-Wei Tang, Pao-Ann Hsiung
ISCAS1
2024 An Edge and Trustworthy AI UAV System With Self-Adaptivity and Hyperspectral Imaging for Air Quality Monitoring
abstract
Leveraging the mobility and flexibility of unmanned aerial vehicles (UAVs), the proposed edge and trustworthy artificial intelligence (AI) UAV system (ETAUS) offers a comprehensive approach to air quality monitoring, complementing fixed monitoring stations. We propose a new convolutional neural network (CNN) model that utilizes hyperspectral imaging (HSI) data as input, enabling ETAUS to directly and accurately classify air quality index (AQI) levels without relying on a back-end AI computing platform. Additionally, ETAUS employs an FPGA-based system architecture, allowing for the integration of a neural engine, cryptographic hardware modules, and hardware protection matrices into a single FPGA device for accelerated processing, edge AI, and trustworthy AI functionalities. Based on its self-adaptivity, edge AI models, cryptographic hardware modules, and protection matrices can also be dynamically loaded into the system to support diverse functional requirements. Experiments have demonstrated that using our proposed CNN model and HSI data, the accuracy of AQI level classification can be achieved to 86.38%. ETAUS can achieve a speedup of$2.28\times $to$36.9\times $in terms of frames per second (FPS) for AQI-level classification compared to microprocessor-based and embedded GPU-based designs. ETAUS can also enhance energy efficiency by$2.7\times $compared to embedded GPU solutions, such as NVIDIA Jetson Nano. To support all the cryptographic functions and protection matrices, system adaptivity in ETAUS can significantly increase resource utilization while decreasing power consumption by up to 2.79%.
Chun-Hsian Huang, Wen-Tung Chen, Yi-Chun Chang, Kuan-Ting Wu
IEEE Internet Things J.1
2023 Area-Adaptive Air Quality Monitoring Based on FPGA with Edge AI and Hyperspectral Imaging
abstract
This work proposes an area-adaptive system design incorporating edge AI and hyperspectral imaging (HSI) to enable mobile and flexible air quality monitoring. To achieve real-time and accurate air quality monitoring in various areas such as industrial, busy traffic, and residential areas, a band selection method based on the fuzzy-rough set theory is presented. This method enables the classification of the air quality index (AQI) level by utilizing informative HSI bands closely related to air pollution. Furthermore, a band-based AQI level classification model is proposed for spectral and spatial feature learning. We further propose an area-adaptive air quality monitoring system (AAQMS) based on FPGA by integrating the band selection method and the band-based AQI level classification model. Experiments show that when all the HSI data in 11 bands are used for AQI level classification, the AAQMS design can enhance energy efficiency by 17.16x compared to the embedded GPU platform (Nvidia Jetson Nano). Through the support of area-adaptive air quality monitoring, although the classification accuracies of AQI levels are slightly reduced, compared to using all the HSI data in 11 bands for industrial, busy traffic, and residential areas, AAQMS can accelerate the AQI level classification by 1.66x, 1.54x, and 1.4x, respectively, while reducing power consumption by 30.32%, 11.61%, and 18.06%, respectively.
Wen-Tung Chen, Chun-Hsian Huang
IECON2
2023 ETAUS: An Edge and Trustworthy AI UAV System with Self-Adaptivity for Air Quality Monitoring
abstract
This work presents the ETAUS, an Edge and Trustworthy AI UAV System, as a mobile sensing platform for air quality monitoring. ETAUS employs an FPGA device as the main hardware computing architecture rather than relying solely on a microprocessor or integrating with GPUs to meet real-time processing demands and achieve adaptivity and scalability. ETAUS contains a neural engine that can execute our customized AI model for air quality index (AQI) level classification and a pre-trained model for detecting objects containing private information. ETAUS also incorporates a de-identification process, cryptographic functions, and protection matrices to safeguard information and individuals' privacy. Furthermore, cryptographic functions and protection matrices are implemented as reconfigurable modules, which can accelerate processing, protect data privacy, and be reconfigured as needed. Experiments have demonstrated ETAUS can achieve a speedup of 3.15x to 72.46x for AQI level classification compared to microprocessor-based and GPU-based designs. ETAUS can also enhance energy efficiency by 5.02x compared to embedded GPU solutions such as NVIDIA Jetson Nano. To support all the cryptographic functions and protection matrices, system adaptivity in ETAUS can significantly increase resource utilization while decreasing power consumption by up to 2.79%.
Chun-Hsian Huang, Wen-Tung Chen, Yi-Chun Chang, Kuan-Ting Wu, Ren-Hong Wang 0004
IROS1
2020 HDA: Hierarchical and dependency-aware task mapping for network-on-chip based embedded systems
Chun-Hsian Huang
J. Syst. Archit.1
2016 Introduction to the special issue on reconfigurable cyber-physical and embedded system design
Pao-Ann Hsiung, Tei-Wei Kuo, Yuan-Hao Chang 0001, Chun-Hsian Huang
J. Syst. Archit.4
2015 A Self-Adaptive System for Vehicle Information Security Applications
abstract
To provide complete vehicle information protection mechanism, this work proposes a self-adaptive system for vehicle information security applications (SAV). Different from the conventional software-based information access method, in the SAV, the access control policies are designed by the protection matrices and implemented as reconfigurable hardware modules. The information access method becomes specific and not generic, so the risks of illegal access of vehicle information can be reduced. To not only meet real-time requirements but also enhance hardware resource utilization, the cryptographic functions in the SAV are also implemented as reconfigurable hardware modules. Thus, the SAV can adapt its access control policies and cryptographic functions at runtime to different system requirements. Our experiments have also demonstrated the SAV can accelerate by up to 3.78x the processing time required by using the software-based design. Compared to the conventional embedded system design, the SAV can also reduce 27.1% of slice registers and 26.5% of slice LUTs in the Xilinx Virtex-5 XC5VLX110T FPGA.
Chun-Hsian Huang, Huang-Yi Chen, Tsung-Fu Huang, Yao-Ying Tzeng, Pei-Shan Wu
EUC1
2014 A reconfigurable point target detection system based on morphological clutter elimination
Chun-Hsian Huang
J. Syst. Archit.1
2013 An FPGA-based point target detection system using morphological clutter elimination
abstract
In this work, we propose a point target detection system (PTDS) based on the FPGA technology. Instead of adopting the traditional filter-based methods, in the PTDS, we design a pipelined morphological clutter elimination (PMCE) hardware design with the ability of pipeline and parallel computing. Using the PMCE hardware design, the infrared images can be processed in real-time, which thus enhances system performance significantly. To provide seamless data transfers between the PMCE hardware design and the microprocessor, a hardware/software interface component is also designed in the PTDS. Experiments with an application of point target detection for processing realtime 320 × 240 pixel infrared images demonstrate that the PTDS can speed up 18 times the execution time required by using the software method.
Chun-Hsian Huang
ISCAS1
2013 Virtualizable hardware/software design infrastructure for dynamically partially reconfigurable systems
abstract
In most existing works, reconfigurable hardware modules are still managed as conventional hardware devices. Further, the software reconfiguration overhead incurred by loading corresponding device drivers into the kernel of an operating system has been overlooked until now. As a result, the enhancement of system performance and the utilization of reconfigurable hardware modules are still quite limited. This work proposes a virtualizable hardware/software design infrastructure (VDI) for dynamically partially reconfigurable systems. Besides the gate-level hardware virtualization provided by the partial reconfiguration technology, VDI supports the device-level hardware virtualization. In VDI, a reconfigurable hardware module can be virtualized such that it can be accessed efficiently by multiple applications in an interleaving way. A Hot-Plugin Connector (HPC) replaces the conventional device driver, such that it not only assists the device-level hardware virtualization but can also be reused across different hardware modules. To facilitate hardware/software communication and to enhance system scalability, the proposed VDI is realized as a hierarchical design framework. User-designed reconfigurable hardware modules can be easily integrated into VDI, and are then executed as hardware tasks in an operating system for reconfigurable systems (OS4RS). A dynamically partially reconfigurable network security system was designed using VDI, which demonstrated a higher utilization of reconfigurable hardware modules and a reduction by up to 12.83% of the processing time required by using the conventional method in a dynamically partially reconfigurable system.
Chun-Hsian Huang, Pao-Ann Hsiung
ACM Trans. Reconfigurable Technol. Syst.1
2011 Model-Based Verification and Estimation Framework for Dynamically Partially Reconfigurable Systems
abstract
Unified Modeling Language (UML), an industry de-facto standard, has been used to analyze dynamically partially reconfigurable systems (DPRS) that can reconfigure their hardware functionalities on-demand at runtime. To make model-driven architecture (MDA) more realistic and applicable to the DPRS design in an industrial setting, a model-based verification and estimation (MOVE) framework is proposed in this work. By taking advantage of the inherent features of DPRS and considering real-time system requirements, a semiautomatic model translator converts the UML models of DPRS into timed automata models with transition urgency semantics for model checking. Furthermore, a UML-based hardware/software co-design platform (UCoP) is proposed to support the direct interaction between the UML models and the real hardware architecture. The two-phase verification process, including exhaustive functional verification and physical-aware performance estimation, is completely model-based, thus reducing system verification efforts. We used a dynamically partially reconfigurable network security system (DPRNSS) as a case study. The related experiments have demonstrated that the model checker in MOVE can alleviate the impact of the state-space-explosion problem. Compared to the synthesis-based estimation method having inaccuracies ranging from -43.4% to 18.4%, UCoP can provide accurate and efficient platform-specific verification and estimation through actual time measurements.
Chun-Hsian Huang, Pao-Ann Hsiung
IEEE Trans. Ind. Informatics1
2010 Learning-based adaptation to applications and environments in a reconfigurable Network-on-Chip
abstract
The set of applications communicating via a Network-on-Chip (NoC) and the NoC itself both have varying run-time requirements on reliability and power-efficiency. To meet these requirements, we propose a novel Power-aware and Reliable Encoding Schemes Supported reconfigurable Network-on-Chip (PRESSNoC) architecture which allows processing elements, routers, and data encoding methods to be reconfigured at runtime. Further, an intelligent selection of encoding methods is achieved through a REasoning And Learning (REAL) framework at run-time. An instance of PRESSNoC was implemented on a Xilinx Virtex 4 FPGA device, which required 25.5% lesser number of slices compared to a conventional NoC with a full-fledged encoding method. The average benefit to overhead ratio of the proposed architecture is greater than that of a conventional NoC by 71%, 32%, and 277% when we consider the individual effects of interference rate per instruction, application domains, and system characteristics, respectively. Experiments have thus shown that PRESSNoC induces a higher probability toward the reduction of crosstalk interferences and dynamic power consumption, at the same amount of overheads in performance and hardware usage.
Jih-Sheng Shen, Chun-Hsian Huang, Pao-Ann Hsiung
DATE2
2010 A Self-Adaptive Hardware/Software System Architecture for Ubiquitous Computing Applications
Chun-Hsian Huang, Jih-Sheng Shen, Pao-Ann Hsiung
UIC1
2010 UML-based hardware/software co-design platform for dynamically partially reconfigurable network security systems
Chun-Hsian Huang, Pao-Ann Hsiung, Jih-Sheng Shen
J. Syst. Archit.1
2010 Model-based platform-specific co-design methodology for dynamically partially reconfigurable systems with hardware virtualization and preemption
Chun-Hsian Huang, Pao-Ann Hsiung, Jih-Sheng Shen
J. Syst. Archit.1
2010 Scheduling and Placement of Hardware/Software Real-Time Relocatable Tasks in Dynamically Partially Reconfigurable Systems
abstract
With the gradually fading distinction between hardware and software, it is now possible to relocate tasks from a microprocessor to reconfigurable logic and vice versa. However, existing hardware-software scheduling can rarely cope with such runtime task relocation. In this work, we propose a new Relocatable Hardware-Software Scheduling (RHSS) method that not only can be applied to dynamically relocatable hardware-software tasks, but also increases the reconfigurable hardware resource utilization, reduces the reconfigurable hardware resource fragmentation with realistic placement methods, and makes best efforts at meeting the real-time constraints of tasks. The feasibility of the proposed relocatable hardware-software scheduling algorithm was proved by applying it to some randomly generated examples and a real dynamically reconfigurable network security system example. Compared to the quadratic time complexity of the state-of-the-art Adaptive Hardware-Software Allocation (AHSA) method, RHSS is linear in time complexity, and improves the reconfigurable hardware utilization by as much as 117.8%. The scheduling and placement time and the memory usage are also drastically reduced by as much as 89.5% and 96.4%, respectively.
Pao-Ann Hsiung, Chun-Hsian Huang, Jih-Sheng Shen, Cheng-Chi Chiang
ACM Trans. Reconfigurable Technol. Syst.2
2009 On the Use of a UML-Based HW/SW Co-Design Platform for Reconfigurable Cryptographic Systems
abstract
In this work, we use our proposed UML-based HW/SW co-design platform (UCoP) to implement a reconfigurable cryptographic system for network multimedia applications. The UCoP is categorized into three reusable models, including software application, hardware configuration and system management models. Through the use of the reusable models, the proposed dynamically partially reconfigurable system can be easily implemented in UCoP. Furthermore, the direct interaction with the real system architecture in UCoP is very helpful to designers in validating and analyzing system correctness and performance at a high-level, which can significantly reduce system development efforts. We compare the use of UCoP with a lower-bound estimation method by implementing a network multimedia application with the data encryption/decryption. According to our experiment results, we can clearly see how UCoP plays a key role in helping designers to develop dynamically partially reconfigurable systems with hard real-time constraints.
Chun-Hsian Huang, Pao-Ann Hsiung
ISCAS1
2009 Modeling and verification of real-time embedded systems with urgency
Pao-Ann Hsiung, Shangwei Lin 0001, Yean-Ru Chen, Chun-Hsian Huang, Chihhsiong Shih, William C. Chu
J. Syst. Softw.4
2007 Dynamically Swappable Hardware Design in Partially Reconfigurable Systems
abstract
In this work, we propose two wrapper designs for arbitrary digital hardware circuit designs such that they can be enhanced with the capability for dynamic swapping controlled by software. A hardware design with either of the proposed wrappers can thus be swapped out of the partially reconfigurable logic at runtime in some intermediate state of computation and then swapped in when required to continue from that state. The context data is saved to a buffer in the wrapper at interruptible states, and then the wrapper takes care of saving the hardware context to communication memory through a peripheral bus, and later restoring the hardware context after the design is swapped in. The overheads of the hardware standardization and the wrapper in terms of additional reconfigurable logic resources and the time for context switching are small and generally acceptable. With the capability for dynamic swapping, high priority hardware tasks can interrupt low priority tasks in real-time embedded systems so that the utilization of hardware space per unit time is increased.
Chun-Hsian Huang, Kai-Jung Shih, Chao-Sheng Lin, Shih-Shiue Chang, Pao-Ann Hsiung
ISCAS1
2006 Model Checking Timed Systems with Urgencies
Pao-Ann Hsiung, Shangwei Lin 0001, Yean-Ru Chen, Chun-Hsian Huang, Jia-Jen Yeh, Chao-Sheng Lin, Hsiao-Win Liao
ATVA4
2006 Perfecto: A Systemc-Based Performance Evaluation Framework for Dynamically Partially Reconfigurable Systems
abstract
To cope with increasing demands for higher computational power and flexibility, dynamically and partially reconfigurable logic has started to play an important role in embedded systems and systems-on-chip. However, when using traditional design methods and tools, it is difficult to estimate or analyze the performance impact of including such reconfigurable logic devices into a system design. In this work, we present an easy-to-use system-level framework, called Perfecto, which is able to perform rapid explorations of different reconfiguration alternatives and to detect system performance bottlenecks. This framework is based on the popular system-level design language SystemC, which is supported by most EDA and ESL tools. Different hardware-software co-partitioning, co-scheduling, and placement algorithms can all be embedded into the framework for analysis. Perfecto can also be used to design the algorithms to be used in an operating system for reconfigurable systems. Applications to some examples have shown advantages of having an evaluation framework such as Perfecto
Pao-Ann Hsiung, Chun-Hsian Huang, Chih-Feng Liao
FPL2
2005 Model Checking Prioritized Timed Automata
Shangwei Lin 0001, Pao-Ann Hsiung, Chun-Hsian Huang, Yean-Ru Chen
ATVA3