Chia-Heng Tu

dblp:79/3912 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0001-8967-1385ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SiliconMind-V1: Multi-Agent Distillation and Debug-Reasoning Workflows for Verilog Code Generation
abstract
Large language models (LLMs) have recently emerged as a promising approach for automating Verilog code generation; however, existing methods primarily emphasize syntactic correctness and often rely on commercial models or external verification tools, which introduces concerns regarding cost, data privacy, and limited guarantees of functional correctness. This work proposes a unified multi-agent framework for reasoning-oriented training data generation with integrated testbench-driven verification, enabling locally fine-tuned LLMs, SiliconMind-V1, to iteratively generate, test, and debug Register-Transfer Level (RTL) designs through test-time scaling. Experimental results on representative benchmarks (VerilogEval-v2, RTLLM-v2, and CVDP) demonstrate that the proposed approach outperforms the state-of-the-art QiMeng-CodeV-R1 in functional correctness while using fewer training resources.
Mu-Chi Chen, Yu-Hung Kao, Po-Hsuan Huang, Shao-Chun Ho, Hsiang-Yu Tsou, I-Ting Wu, En-Ming Huang, Yu-Kai Hung, Wei-Po Hsin, Chia-Heng Tu, Shih-Hao Hung, H. T. Kung 0001
COMPSAC11
2026 Alleviating Congestion Attacks on Traffic Signal Systems
abstract
This work addresses congestion attacks on prioritization and preemption signal applications (PPSA), an important variant of cooperative intelligent transport systems (C-ITS). In PPSA, priority vehicles, such as buses and ambulances, can request traffic signal adjustments in their favor to reduce travel time. These systems rely on vehicle-to-everything (V2X) communications, which makes them vulnerable to congestion attacks. When under attack, these systems experience disruptions in information exchange between vehicles and infrastructure, leading to traffic congestion and safety hazards. Nevertheless, existing solutions do not address the attacks at a system level and within the PPSA context. Our proposal presents a software framework designed to enhance both sustainability and safety during congestion attacks. Specifically, it introduces a whitelist-based traffic-filtering mechanism that preserves system sustainability by ensuring only legitimate priority requests are processed. Additionally, a monitoring mechanism evaluates the PPSA service rate at runtime to identify and address potential safety risks arising from insufficient service rates. Moreover, we analyze the limitations of a machine learning-based defense method and highlight risks of potential service unavailability due to false positives. We believe that this research contributes to the development of a secure and safe traffic management system for smart cities.
Tsung-Lin Tsai, Shao-Hua Wang, Chun-Ting Wu, Chia-Heng Tu, Meng-Hsun Tsai, Wei-Hsun Lee, Da-Wei Chang
ACM Trans. Cyber Phys. Syst.5
2026 Timing-Constrained Composable Inference for Intermittent Systems Using Reinforcement Learning
abstract
The increasing maturity of energy harvesting technologies has brought intermittent systems to the forefront as viable solutions for a range of applications. One critical area is environmental monitoring, where timely and accurate reporting of environmental conditions is essential. Existing approaches on systems powered by unstable ambient energy focus on maintaining the freshness of collected data but fall short when applied to neural network workloads, as they often neglect model accuracy. Prior studies have explored deploying neural networks on intermittent systems using branchy architectures, which prioritize energy-accuracy tradeoffs by terminating inference early. However, these approaches fail to address time constraints, often resulting in system failures due to processing expired results. This article introduces iTRAIN, a novel timing-aware framework for deploying neural network models on intermittent systems. iTRAIN holistically accounts for energy availability, timing constraints, and model accuracy. Unlike previous studies that depend on branchy architectures, iTRAIN leverages a composable neural network framework to broaden the solution space, enabling diverse energy-time-accuracy tradeoffs. This is achieved through runtime selection among various layer implementations, such as pruning and quantization, guided by a reinforcement learning algorithm. Experimental results demonstrate that iTRAIN outperforms state-of-the-art approaches, achieving a 65% improvement in delivered model accuracy with minimal memory and runtime overhead. iTRAIN sets a foundation for enabling complex applications on intermittent systems.
Wen Sheng Lim, Shu-Ting Cheng, Ya-Tung Tsai, Chia-Heng Tu, Yuan-Hao Chang 0001
ACM Trans. Embed. Comput. Syst.4
2025 Analyzing vertical handover of energy efficient sleep mode schemes in heterogeneous networks
Hao-Zhong Zheng, Chun-Hao Yang, Fang-Yi Lee, Chia-Heng Tu, Meng-Hsun Tsai
Comput. Commun.5
2025 iSAFE: Enabling Evenness of Data Freshness in Multipriority Networked Intermittent Systems
abstract
Environmental monitoring applications use energy harvesting to cover wide-range deployment, where devices are powered by ambient energy and operate intermittently when energy is sufficient. In such an intermittent networked system (NIS), a sink node is used to forward the environmental data collected by sensors to a central controller to reflect the physical environment status. Nevertheless, existing data forwarding algorithms for NISs cannot fulfill modern application requirements, where multiple types of data with different timeliness requirements (i.e., multipriorities) are desired to report real-time environmental data for monitoring critical situations. Without considering the multipriorities, we show in this article that it introduces a new problem: unevenness of data freshness. We then propose the sink node-based evenness-aware update forwarding (iSAFE) algorithm to provide evenness among different priorities of data sources in NISs. iSAFE consists of three important components: 1) a theoretical analysis to derive the optimal data forwarding interval between two adjacent status updates from the sensor; 2) an evenness-aware forwarding algorithm to adaptively adjust the forwarding interval based on the runtime status; and 3) a fresh-aware energy preservation algorithm to maintain the freshness of collected data. The experimental results show that iSAFE can achieve up to 682% evenness (94.47% close to the ideal) and 53.3% data freshness compared to the state of the art while being energy-efficient and scalable, suitable for modern applications.
Wen Sheng Lim, Yu-Hsuan Chu, Chia-Heng Tu, Yuan-Hao Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 PARD: A Dataflow Aware Profiler for ROS-Based Autonomous Driving Software
abstract
Robot Operating System (ROS) is an open source software platform that is well-suited for building complex robotic systems. Autonomous driving systems are one of the emerging robotic systems that are built on top of ROS. However, complicated communications among ROS processes pose a great challenge for system designers to design efficient computing hardware in autonomous vehicles. This complexity stems from the implicit data dependencies among these processes. When a process is bound by computation or communication, it can lead to increased data generation time for subsequent processes in the dependency chain. Unfortunately, existing tools for ROS are either built for tracking the communication latencies or revealing the communication performance with abstracted performance graphs. For analyzing computational resource utilization, general-purpose profiling tools like perf should be further adopted. Substantial human effort is required to consolidate and analyze all the performance data generated by the above tools so as to pinpoint potential performance issues. To address this challenge, a ROS-based dataflow aware profiler PARD is proposed. PARD employs dataflow aware analysis methods to automatically analyze collected ROS performance data, and highlights potential computation and communication performance issues on our profile graph, helping users promptly identify system bottlenecks. Thanks to the proposed data analysis methods, the system designers are able to identify and solve the performance issues rapidly through the graphical data representation that accentuates the potential performance issues. The experimental results on the commercial-grade ROS-based autonomous driving software Autoware demonstrate PARD is useful for quickly identifying the potential performance issues for autonomous valet parking. Moreover, its overall impact on the system is around 2%, indicating that PARD is indeed useful for accelerating hardware design exploration in autonomous driving vehicles.
Shao-Hua Wang, Chen-Xuan Lin, Yu-Tsung Wu, Chia-Heng Tu
ACM Trans. Cyber Phys. Syst.4
2025 QOPS: a compiler framework for quantum circuit simulation acceleration with profile-guided optimizations
Yu-Tsung Wu, Po-Hsuan Huang, Kai-Chieh Chang, Chia-Heng Tu, Shih-Hao Hung
J. Supercomput.4
2024 Toward cost-effective quantum circuit simulation with performance tuning techniques
abstract
Quantum circuit simulation is a popular approach to evaluating novel quantum algorithms before a physical quantum computer is available. Unfortunately, the simulation is often done with the full-state quantum circuit simulation scheme, and a huge memory space required by a full-state simulator is a limiting factor for the simulation of a larger qubit system. In this work, in order to support the simulation of a broadened qubit system, storage devices are introduced into the full-state quantum circuit simulation. A vertical qubit simulation design is proposed for the storage-based, full-state quantum circuit simulation, including qubit representation, threading model, and parallel state manipulations over storage devices. An empirical method of simulator parameter tuning is developed to achieve higher simulation performance. Our experimental results show that compared with the state-of-the-art memory-only simulator (QuEST), the storage-based simulation can achieve a 61x higher cost-delay ratio and can simulate a 39-qubit system on a commodity computer. The encouraging results indicate that our proposed simulator can help scale the full-state simulation for larger quantum circuits and achieve higher performance via the performance tuning method.
Nai-Wei Hsu, Chuan-Chi Wang, Chia-Hsin Hsu, Chia-Heng Tu, Shih-Hao Hung
Connect. Sci.4
2024 Markov Clustering-Based Content Placement in Roadside-Unit Caching With Deadline Constraint
abstract
With the explosive growth of mobile data traffic, roadside-unit (RSU) caching is considered an effective way to offload download traffic in vehicular ad hoc networks (VANETs). Many existing works investigate the content placement of RSU caching. However, few of them consider the download deadline constraint when caching the content in the RSUs. In this paper, the main objective is to maximize the hit rate of downloading the requested content from the RSUs before the deadline expires. We propose a Markov-based mobility model and a Markov clustering-based content placement algorithm to group the RSUs into clusters and allocate the content to the cache of the RSUs in the cluster. We also investigate the impact on the cache hit rate under different simulation parameters, such as the total number of RSUs, the cache size, and the number of RSUs visited by vehicles during the download period. According to the simulations conducted, when the region of interest (RoI) is small, the MVP method increases the cache hit rate by at least 21.40% compared to the existing methods. When the RoI is large, our approach outperforms other existing methods by at least 26.16% and at most 337.77%, which significantly increases the efficiency of the download session in VANET.
Yu-Ting Wang 0002, Sok-Ian Sou, Lo-An Chen, Meng-Hsun Tsai, Yean-Ru Chen, Chia-Heng Tu
IEEE Trans. Intell. Transp. Syst.7
2023 Data Freshness Optimization on Networked Intermittent Systems
abstract
A networked intermittent system (NIS) is often deployed in the field for environmental monitoring, where sink nodes are responsible for relaying the data captured by sensors to a central system. To evaluate the quality of the captured monitoring data, Age of Information (AoI) is adopted to quantify the freshness of the data received by the central server. As the sink nodes are powered by ambient energy sources (e.g., solar and wind), the energy-efficient design of the sink nodes is crucial in order to improve the system-wide AoI. This work proposes the energy-efficient sink node design to save energy and extend system uptime. We devise an AoI-aware data forwarding algorithm based on the branch-and-bound (B&B) paradigm for deriving the optimal solution offline. In addition, an AoI-aware data forwarding algorithm is developed to approximate the optimal solution during runtime. The experimental results show that our solution can greatly improve the average data freshness for 148% against existing well-known strategies and achieves 91 % performance of the optimal solution. Compared with the state-of-the-art algorithm, our energy-efficient design can deliver better$A^{3}oI$results by up to 9.6%.
Hao-Jan Huang, Wen Sheng Lim, Chia-Heng Tu, Chun-Feng Wu, Yuan-Hao Chang 0001
DATE3
2023 TRAIN: A Reinforcement Learning Based Timing-Aware Neural Inference on Intermittent Systems
abstract
Intermittent systems become popular to be considered as the solutions of various application domains, thanks to the maturation of energy harvesting technology. Environmental monitoring is such an example and it is a time-sensitive application domain. In order to report the perceived environmental status in a timely manner, methods have been proposed to consider the freshness of the collected information on such systems with unstable power sources. Nevertheless, these methods cannot be applied to neural network workloads since these methods do not consider the delivered model accuracy. On the other hand, while there have been studies for deploying neural network applications on intermittent systems, they depend on branchy network architectures, each branch representing an energy-accuracy tradeoff, and do not take into account a time constraint, which tends to cause system failures because of the frequent generation of expired data. In this work, the first timing-aware framework TRAIN is proposed to deploy the neural network models on the intermittent systems by considering energy, time constraint, and delivered model accuracy. Compared with the prior studies that depend on branchy network architectures, TRAIN offers a broadened solution space representing various energy/time/accuracy tradeoffs. It is achieved by allowing to choose among different implementations of each model layer during the model inference at runtime, and the smart choices are made by the proposed reinforcement learning algorithm. Our results demonstrate TRAIN outperforms the prior study by 65%, regarding the delivered model accuracy. We believe that TRAIN paves the way for building complex applications on intermittent systems.
Shu-Ting Cheng, Wen Sheng Lim, Chia-Heng Tu, Yuan-Hao Chang 0001
ICCAD3
2023 SecureTVM: A TVM-based Compiler Framework for Selective Privacy-preserving Neural Inference
abstract
Privacy-preserving neural inference helps protect both the user input data and the model weights from being leaked to others during the inference of a deep learning model. To achieve data protection, the inference is often performed within a secure domain, and the final result is revealed in plaintext. Nevertheless, performing the computations in the secure domain incurs about a thousandfold overhead compared with the insecure version, especially when the involved operations of the entire model are mapped to the secure domain, which is the computation scheme adopted by the existing works. This work is inspired by the transfer learning technique, where the weights of some parts of the model layers are transferred from a publicly available, pre-built deep learning model, and it opens a door to further boost the execution efficiency by allowing us to do the secure computations selectively on parts of the transferred model. We have built a compiler framework, SecureTVM, to automatically translate a trained model into the secure version, where the model layers to be protected can be selectively configured by its model provider. As a result, SecureTVM outperforms the state of the art, CrypTFlow2, by a factor of 55 for the transfer learning model. We believe that this work takes a step forward toward the practical uses of privacy-preserving neural inference for real-world applications.
Po-Hsuan Huang, Chia-Heng Tu, Shen-Ming Chung, Pei Yuan Wu, Tung-Lin Tsai, Yi-An Lin, Chun-Yi Dai, Tzu-Yi Liao
ACM Trans. Design Autom. Electr. Syst.2
2022 ESCA: Effective System Call Aggregation for Event-Driven Servers
abstract
The switches between a non-privileged application and the OS kernel running in the CPU’s supervisor mode have been inducing performance costs despite manufacturer efforts to provide special instructions for such transition. Software that heavily interacts with the underlying OS (e.g., I/O intensive and event-driven applications) suffers from system call overhead. To deteriorate this situation, security vulnerabilities in modern processors have prompted kernel mitigations that further increase the transition overhead. Particularly system-call-heavy applications have been reported to be slowed down by up to 30% with kernel page-table isolation (KPTI), the widely deployed mitigation for the Meltdown vulnerability. To decouple system calls from mode transitions, we revisit an old idea known as system-call batching or multi-calls: the bundling of system calls into a combined call, which only incurs the mode-transition costs of a single one. And then, we have implemented ESCA scheme to adapt system-call batching to Linux-based servers in the light of Meltdown and Spectre, effectively eliminating the slowdown of KPTI-affected applications. Our evaluation shows that the throughputs of real-world applications, benefiting from ESCA, can be improved with only 2 lines of code changed respectively: Nginx by up to 12%, lighttpd by up to 23%, and Redis by 4%. Meanwhile, using aggregated transitions, our approach allows faster system calls interleaved with full compatibility but without requiring Linux kernel patches.
Yu-Cheng Cheng, Ching-Chun Jim Huang, Chia-Heng Tu
PDP3
2022 Performance Acceleration of Secure Machine Learning Computations for Edge Applications
abstract
Edge appliances built with machine learning applications have been gradually adopted in a wide variety of application fields, such as intelligent transportation, the banking industry, and medical diagnosis. Privacy-preserving computation approaches can be used on smart appliances in order to secure the privacy of sensitive data, including application data and the parameters of machine learning models. Nevertheless, the data privacy is achieved at the cost of execution time. That is, the execution speed of a secure machine learning application is several orders of magnitude slower than that of the application in plaintext. Especially, the performance gap is enlarged for edge appliances. In this work, in order to improve the execution efficiency of secure applications, an open-source software framework CrypTen is targeted, which is widely used for building secure machine learning applications using the Secure Multi-Party Computation (SMPC) based privacy-preserving computation approach. We analyze the performance characteristics of the secure machine learning applications built with CrypTen, and the analysis reveals that the communication overhead hinders the execution of the secure applications. To tackle the issue, a communication library, OpenMPI, is added to the CrypTen framework as a new communication backend to boost the application performance by up to 50%. We further develop a hybrid communication scheme by combining the OpenMPI backend with the original communication backend with the CrypTen framework. The experimental results show that the enhanced CrypTen framework is able to provide better performance for the small-size data (LeNet5 on MNIST dataset by up to 50% of speedup) and maintain similar performance for large-size data (AlexNet on CIFAR-10), compared to the original CrypTen framework.
Zi-Jie Lin, Chuan-Chi Wang, Chia-Heng Tu, Shih-Hao Hung
RTCSA3
2022 Automatic traffic modelling for creating digital twins to facilitate autonomous vehicle development
abstract
A digital twin is often adopted in computer simulations to expedite autonomous vehicle developments by using the simulated 3D environment that reflects a physical environment. In particular, traffic simulations are a crucial part of training the driving logic before the field test of an autonomous vehicle is performed on specific regions to adapt to the region-specific, dynamic traffic conditions. Currently, the traffic conditions are either synthesised by tools (e.g. using mathematical models) or created manually (using domain knowledge), which cannot reflect the realistic, region-specific conditions or will require extensive labour works. In this article, we propose an automatic methodology to model the real-world traffic conditions captured by the sensor data and to reproduce the modeled traffic in the digital twin. We have built the tools based on the methodology and use the KITTI dataset to validate the effectiveness of the tools. To recreate the region-specific traffic, we present the results of capturing, modelling, and recreating the two-wheeler traffic condition on the Southeast Asia road. Our experimental results show that the proposed method facilitates the simulation of real-world, Southeast Asia-specific traffic conditions by removing the needs of the synthesised traffic and the labour hours.
Shao-Hua Wang, Chia-Heng Tu, Jyh-Ching Juang
Connect. Sci.2
2022 POPS: an off-peak precomputing scheme for privacy-preserving computing
Po-Hsuan Huang, Ting-Wei Chang, Chia-Heng Tu, Shen-Ming Chung
J. Supercomput.3
2021 iCheck: Progressive Checkpointing for Intermittent Systems
abstract
Energy harvesting devices powered by ambient energies, instead of batteries, have been drawn lots of attention due to their advantages of energy saving, easy deployment without relying on stable power sources, and smaller sizes, facilitating promising applications, such as environmental and health monitoring. These devices perform the computations intermittently, where the code executions are halted and resumed depending on the availability of the harvested energy. On such devices, the capacitors are present and served as the energy buffers for preserving the program states when sudden power outages occur. Nevertheless, the capacitors have relatively shorter lifetimes, compared with the rest of hardware components on the devices, and larger capacitors, which are desired by the systems requiring complex computations, hamper the achievement of device miniaturization, e.g., for medical implants or smart dust. In this article, we propose a new intermittent checkpointing strategy,iCheck, to tackle the issues raised for the program-state retaining when the capacitors are not functioning correctly (or when the capacitor-less devices are adopted). The proposediCheckis designed to perform the checkpointing-based program-state preserving progressively with being aware of the power-failure characteristics of the harvested energy source to maximize the progress forwarding and to ensure data consistency while encountering incomplete checkpoints caused by sudden power losses. The proposed design is evaluated with a series of experiments with encouraging results.
Wen Sheng Lim, Chia-Heng Tu, Chun-Feng Wu, Yuan-Hao Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 RAP: A Software Framework of Developing Convolutional Neural Networks for Resource-constrained Devices Using Environmental Monitoring as a Case Study
abstract
Monitoring environmental conditions is an important application of cyber-physical systems. Typically, the monitoring is to perceive surrounding environments with battery-powered, tiny devices deployed in the field. While deep learning-based methods, especially the convolutional neural networks (CNNs), are promising approaches to enriching the functionalities offered by the tiny devices, they demand more computation and memory resources, which makes these methods difficult to be adopted on such devices. In this article, we develop a software framework, RAP , that permits the construction of the CNN designs by aggregating the existing, lightweight CNN layers, which are able to fit in the limited memory (e.g., several KBs of SRAM) on the resource-constrained devices satisfying application-specific timing constrains. RAP leverages the Python-based neural network framework Chainer to build the CNNs by mounting the C/C++ implementations of the lightweight layers, trains the built CNN models as the ordinary model-training procedure in Chainer, and generates the C version codes of the trained models. The generated programs are compiled into target machine executables for the on-device inferences. With the vigorous development of lightweight CNNs, such as binarized neural networks with binary weights and activations, RAP facilitates the model building process for the resource-constrained devices by allowing them to alter, debug, and evaluate the CNN designs over the C/C++ implementation of the lightweight CNN layers. We have prototyped the RAP framework and built two environmental monitoring applications for protecting endangered species using image- and acoustic-based monitoring methods. Our results show that the built model consumes less than 0.5 KB of SRAM for buffering the runtime data required by the model inference while achieving up to 93% of accuracy for the acoustic monitoring with less than one second of inference time on the TI 16-bit microcontroller platform.
Chia-Heng Tu, Qihui Sun, Hsiao-Hsuan Chang
ACM Trans. Cyber Phys. Syst.1
2021 On designing the adaptive computation framework of distributed deep learning models for Internet-of-Things applications
Chia-Heng Tu, Qihui Sun, Mu-Hsuan Cheng
J. Supercomput.1
2019 Augmenting Operating Systems with OpenCL Accelerators
abstract
Heterogeneous computing leverages more than one kind of processors to boost the performance of user-space applications with the heterogeneous programming languages, e.g., OpenCL. While some works have been done to accelerate the computations required by Linux kernel software, they are either application-specific solutions or tightly coupled with the certain computing platforms and are not able to support the general-purpose in-kernel accelerations using different types of processors. In this article, the general-purpose software framework called Kernel acceleration with OpenCL (KOCL), is proposed to tackle the problem. KOCL exposes a set of the high-level programming interfaces for the Linux kernel module developers to offload compute-intensive tasks on different hardware accelerators without managing and coordinating the platform-specific computing and memory resources. The simplified programming efforts are achieved by the developed platform management and memory models, which provide a systematic means of managing the heterogeneous hardware resources. In addition, the one- and zero-copy data-buffering schemes are offered by KOCL, so that the offloaded tasks deliver high performance on the platforms with different memory architectures. We have developed the prototype system to accelerate the Network-Attached Storage server applications. Significant performance improvements are achieved with the three different types of accelerators, i.e., the multicore processor, the integrated GPU, and the discrete GPU, respectively. We believe that KOCL is useful for the design of embedded appliances to evaluate the performance of design alternatives.
Chia-Heng Tu, Te-Sheng Lin
ACM Trans. Design Autom. Electr. Syst.1
2018 An Adaptive Computation Framework of Distributed Deep Learning Models for Internet-of-Things Applications
abstract
We propose the computation framework that facilitates the inference of the distributed deep learning model to be performed collaboratively by the devices in a distributed computing hierarchy. For example, in Internet-of-Things (IoT) applications, the three-tier computing hierarchy consists of end devices, gateways, and server(s), and the model inference could be done adaptively by one or more computing tiers from the bottom to the top of the hierarchy. By allowing the trained models to run on the actually distributed systems, which has not done by the previous work, the proposed framework enables the co-design of the distributed deep learning models and systems. In particular, in addition to the model accuracy, which is the major concern for the model designers, we found that as various types of computing platforms are present in IoT applications fields, measuring the delivered performance of the developed models on the actual systems is also critical to making sure that the model inference does not cost too much time on the end devices. Furthermore, the measured performance of the model (and the system) would be a good input to the model/system design in the next design cycle, e.g., to determine a better mapping of the network layers onto the hierarchy tiers. On top of the framework, we have built the surveillance system for detecting objects as a case study. In our experiments, we evaluate the delivered performance of model designs on the two-tier computing hierarchy, show the advantages of the adaptive inference computation, analyze the system capacity under the given workloads, and discuss the impact of the model parameter setting on the system capacity. We believe that the enablement of the performance evaluation expedites the design process of the distributed deep learning models/systems.
Mu-Hsuan Cheng, Qihui Sun, Chia-Heng Tu
RTCSA3
2018 Phase-Based Profiling and Performance Prediction with Timing Approximate Simulators
abstract
Designing a system usually acquires lots of instincts and knowledge to harmonize the computing resources under the paradigm of application specific heterogeneous systems. This paper presents a phase-based profiling mechanism to speed up the process of learning how application behaviors perform on the hardware and vice versa. By analyzing program phases, performance information can be gathered in a way that highlights the performance of high-level tasks in an application running on different hardware settings. We evaluated our phase-based profiling framework using QEMU, employing approximate timing models and mechanisms to track functions/events in programs and operating systems of the guest system. Furthermore, by using timing simulations, it is possible to escape the confined boundaries of real-world machine based systems, and to rapidly explore the impact of hardware parameters on the system performance. In our experimental results, phase-based profiling yields useful information of the runtime behaviors and performance of a program, allowing developers to discover program bottlenecks, and predicts the performance of optimization ideas on the software and/or underlying hardware. Our results suggest that incorporating phase profiling with the timing approximate simulator helps to facilitate hardware and software co-design.
Chih Wei Yeh, Chia-Heng Tu, Yi-Chuan Liang, Shih-Hao Hung
RTCSA2
2017 GPU acceleration for Kernel Samepage Merging
abstract
Kernel Samepage Merging (KSM) is a Linux kernel module for improving memory utilization by searching and merging the redundant memory pages. When working with the hypervisors, such as Kernel-based Virtual Machine, KSM helps share identical memory pages of the hosted virtual servers so as to increase the server density. Nevertheless, while KSM improves the efficiency of the host system, it hurts the performance of the virtualized systems since part of the CPU cycles is spent on exploring the page sharing opportunities. In this paper, two optimization schemes, selective page comparison and checksum computation acceleration, are proposed to improve the KSM efficiency, as well as to reduce the CPU burden at the same time. Selective page comparison skips the expensive content comparison operations, which are performed by default in the original KSM for examining redundant pages, for the frequently-changing pages according to their checksums. Checksum computation acceleration shifts the CPU load of the page checksum calculation to the GPU by leveraging the in-kernel computation acceleration framework, KGPU. We implemented the two optimizations, and evaluated their performance with the web server workloads on the virtualized servers. Our results show that the speed of the page-merging process is improved by a factor of up to 1.64 with the both optimization schemes. To the best of our knowledge, we believe this is the first attempt to accelerate the KSM with the GPU, and our work paves a road toward a more sophisticated, GPU-enabled memory deduplication algorithm.
Chia-Heng Tu, Chih Wei Yeh, Shih-Hao Hung
RTCSA2
2017 Fast profiling framework and race detection for heterogeneous system
Cheng-Kung Lai, Chih Wei Yeh, Chia-Heng Tu, Shih-Hao Hung
J. Syst. Archit.3
2016 Relay-based key management to support secure deletion for resource-constrained flash-memory storage devices
abstract
The support of secure deletion on formatting a file system is to make sure that when a file system is formatted, there is no way to get any file content back again. Due to the fast-growing storage capacity, the performance of secure deletion to file systems on resource-constrained flash storage devices has become a critical issue. In contrast to the existing works that take a long time on overwriting/resetting all the file contents of a file system, we propose an efficient secure deletion scheme to securely delete all the contents of a file system without rewriting file contents. Thus, secure deletion to file systems can be efficiently achieved and can be independent of the device capacity and file systems. A series of experiments was conducted with realistic workloads to evaluate the capability of the proposed scheme. The results show that the proposed scheme achieves secure deletion with limited performance overheads in most cases.
Wei-Lin Wang, Yuan-Hao Chang 0001, Po-Chun Huang, Chia-Heng Tu, Hsin-Wen Wei, Wei-Kuan Shih
ASP-DAC4
2016 A Platform-Oblivious Approach for Heterogeneous Computing: A Case Study with Monte Carlo-based Simulation for Medical Applications
abstract
Light is important and helpful in many medical applications, such as cancer treatment. Computer modeling and simulation of light transport are often adopted to improve the quality of medical treatments. In particular, Monte Carlo-based simulations are considered to deliver accurate results, but require intensive computational resources. While several attempts to accelerate the Monte Carlo-based methods for the simulation of photon transport with platform-specific programming schemes, such as CUDA on GPU and HDL on FPGA, have been proposed, the approach has limited portability and prolongs software updates. In this paper, we parallelize the Monte Carlo modeling of light transport in multi-layered tissues (MCML) program with OpenCL, an open standard supported by a wide range of platforms. We characterize the performance of the parallelized MCML kernel program runs on CPU, GPU and FPGA. Compared to platform-specific programming schemes, our platform-oblivious approach provides a unified, highly portable code and delivers competitive performance and power efficiency.
Shih-Hao Hung, Min-Yu Tsai, Bo-Yi Huang, Chia-Heng Tu
FPGA4
2014 MobileFBP: Designing portable reconfigurable applications for heterogeneous systems
Shih-Hao Hung, Tien-Tzong Tzeng, Jyun-De Wu, Min-Yu Tsai, Yi-Chih Lu, Jeng-Peng Shieh, Chia-Heng Tu, Wen-Jen Ho
J. Syst. Archit.7
2014 Performance and power profiling for emulated Android systems
abstract
Simulation is a common approach for assisting system design and optimization. For system-wide optimization, energy and computational resources are often the two most critical issues. Monitoring the energy state of each hardware component and measuring the time spent in each state is needed for accurate energy and performance prediction. For software optimization, it is important to profile the energy and the time consumed by each software construct in a realistic operating environment with a proper workload. However, the conventional approaches of simulation often fail to produce satisfying data. First, building a cycle-accurate simulation environment for a complex system, such as an Android smartphone, is difficult and can take a long time. Second, a slow simulation can significantly alter the behavior of multithreaded, I/O-intensive applications and can affect the accuracy of profiles. Third, existing software-based profilers generally do not work on simulators, which makes it difficult for performance analysis of complicated software, for example, Java applications executed by the Dalvik VM in an Android system. To address these aforementioned problems, we proposed and prototyped a framework, called virtual performance analyzer (VPA). VPA takes advantage of an existing emulator or virtual machine monitor to reduce the complexity of building a simulator. VPA allows the user to selectively and incrementally integrate timing models and power models into the emulator with our carefully designed performance/power monitors, tracing facility, and profiling tools to evaluate and analyze the emulated system. The emulated system can perform at different levels of speed to help verify if the profile data are impacted by the emulation speed. Finally, VPA supports existing software-based profiles and enables non-intrusive tracing/profiling by minimizing the probe effect. Our experimental results show that the VPA framework allows users to quickly establish a performance/power evaluation environment and gather useful information to support system design and software optimization for Android smartphones.
Chia-Heng Tu, Hui-Hsin Hsu, Jen-Hao Chen, Chun-Han Chen, Shih-Hao Hung
ACM Trans. Design Autom. Electr. Syst.1
2012 System-wide profiling and optimization with virtual machines
abstract
Simulation is a common approach for assisting system design and optimization. For system-wide optimization, energy and computational resources are often the two most critical limitations. Modeling energy-states of each hardware component and time spent in each state is needed for accurate energy and performance prediction. Tracking software execution in a realistic operating environment with properly modeled input/output is key to accurate prediction. However, the conventional approaches can have difficulties in practice. First, for a complex system such as an Android smartphone, building a cycle-accurate simulation environment is no easy task. Secondly, for I/O-intensive applications, a slow simulation would significantly alter the application behavior and change its performance profile. Thirdly, conventional software profiling tools generally do not work on simulators, which makes it difficult for performance analysis of complicated software, e.g., Java applications executed by the Dalvik virtual machine. Recently, virtual machine technologies are widely used to emulate a variety of computer systems. While virtual machines do not model the hardware components in the emulated system, we can ease the effort of building a simulation environment by leveraging the infrastructure of virtual machines and adding performance and power models. Moreover, multiple sets of the performance and energy models can be selectively used to verify if the speed of the simulated system impacts the software behavior. Finally, performance monitoring facilities can be integrated to work with profiling tools. We believe this approach should help overcome the aforementioned difficulties. We have prototyped a framework and our case studies showed that the information provided by our tools are useful for software optimization and system design for Android smartphones.
Shih-Hao Hung, Tei-Wei Kuo, Chi-Sheng Shih 0001, Chia-Heng Tu
ASP-DAC4
2012 MCEmu: A Framework for Software Development and Performance Analysis of Multicore Systems
abstract
Developing software for heterogeneous multicore systems is particularly challenging even for experienced developers. While emulators have proven useful to application development, very few heterogeneous multicore emulators have been made available by vendors so far, as building an emulator for a heterogeneous multicore system has been a time-consuming and difficult task. Thus, we proposed a framework, called MCEmu, to speed up the process of building a heterogeneous multicore emulator by integrating existing and/or new processor emulators. MCEmu is designed to help system and application development, with a basic multicore board support package, an interprocessor communication library, and tools for debugging, tracing, and performance monitoring. In addition, MCEmu can run on a multicore host system to accelerate the emulation of data parallel applications. We show that MCEmu can be very useful for developing system software before the system becomes available, as it has helped us catch numerous functional and performance bugs which could have been hard to find. In this article, we present the design of MCEmu and demonstrate its capabilities with our case studies.
Chia-Heng Tu, Shih-Hao Hung, Tung-Chieh Tsai
ACM Trans. Design Autom. Electr. Syst.1
2011 A portable, efficient inter-core communication scheme for embedded multicore platforms
Shih-Hao Hung, Chia-Heng Tu, Wen-Long Yang
J. Syst. Archit.2
2010 Trace-based performance analysis framework for heterogeneous multicore systems
abstract
Performance evaluation is key to the optimization of computer applications on multicore systems. While many techniques and profiling tools are available for measuring performance on homogeneous multicore platforms, most of them depend on the hardware support from the vendors. For developing applications on heterogeneous multicore systems, very few analysis tools exist to help the developers. This paper describes a software-based trace collection and performance analysis framework that can be ported to a variety of platforms via code instrumentation at the source level. A pure software profiling toolkit, called ParallelTracer, were implemented based on ANTLR, an open source parser generator, to support this framework. In this paper, we present our framework and toolkit. We use the IBM Cell processor as a case study to demonstrate the capability of ParallelTrace. Our results show that ParallelTracer provided useful information for programmers to understand program behaviors and identify potential performance bottlenecks via graphical visualization. We also discuss the runtime overhead of ParallelTracer. With proper usage, the performance and code size overhead introduced by our toolkit are limited around 19% to 5% and 9%, respectively, for the benchmark program in the case study.
Shih-Hao Hung, Chia-Heng Tu, Thean-Siew Soon
ASP-DAC2
2010 V2X: An Automated Tool for Building SystemC-Based Simulation Environments in Designing Multicore Systems-on-Chips
abstract
Hardware/software (HW/SW) co-design has become an important issue for system design, and simulation environments have been utilized widely to shorten the development cycle. However, traditional hardware description languages (HDL), e.g., Verilog and VHDL, which are used by hardware designers to describe the hardware and model the hardware in a detailed simulated environment, are not appropriate for the purpose of HW/SW co-design. Instead, SystemC provides a higher-level simulation environment to the developers and is more suitable for HW/SW co-design. Furthermore, HDL-based simulation environments are far too slow to execute parallel programs as the number of processor cores increases. Thus, one would have liked an automated tool for converting existing HDL-based chip designs to SystemC or even higher-level functional descriptions so that the simulation speed would be acceptable for multicore systems. However, since existing tools failed to accomplish that, we developed an automated tool, called V2X, to convert Verilog chip designs to SystemC. In this paper, we show that complicated Verilog-based multicore chip descriptions were translated into SystemC descriptions automatically and resulted in better performance and programmability. In our case study, V2X successfully translated the 8-core OpenSPARC T1 system-on-chip into SystemC. Without further abstraction, the simulation speed was improved by ~40 times. The two-stage translation scheme makes V2X flexible and extensible, which paves the way for further abstraction to speed up the simulation environment.
Yun-Hung Liaw, Shih-Hao Hung, Chia-Heng Tu
ISPA3
2010 Designing and Implementing a Portable, Efficient Inter-core Communication Scheme for Embedded Multicore Platforms
abstract
In the recent years, multicore processor designs have become increasingly popular for embedded applications, but diversified inter-core communication mechanisms have led to the difficulties in software development, integration and migration. A unified, portable, and efficient inter-core communication mechanism would have helped reduce these difficulties significantly, but such a solution did not exist today. We proposed a scheme called MSG, which provides users with a set of essential message-passing programming interfaces adopted from MPI and MCAPI, including blocking and non-blocking point-to-point communications, one-sided communications, and collective operations. We experimented and evaluated our design methodology with the case study on the IBM CELL, a popular heterogeneous multicore platform. On the CELL platform, our MSG library fitted in the 256KB local memory on each individual processor core and outperformed two existing communication libraries, DaCS and CML. With a systematic approach, we showed how optimization could be done on the CELL platform to improve the performance of the MSG library. Hopefully, our experiences help the design and development of communication libraries for existing and future multicore platforms and embedded applications.
Shih-Hao Hung, Wen-Long Yang, Chia-Heng Tu
RTCSA3
2009 Prefetch optimizations on large-scale applications via parameter value prediction
abstract
A typical data center application requires the processor cycles of thousands of machines. Even a single-digit performance improvement can significantly reduce the cost and power consumption of a data center. Unfortunately, achieving sustained improvement, even if modest, is difficult. Data centers are dynamic environments where applications are frequently released and servers are continually upgraded. For maintainability and fault tolerance, the physical capabilities and configuration of the servers are abstracted from the application programmer.
Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Hucheng Zhou, Chinyen Chou, Chia-Heng Tu
ICS6
2009 Zero-Buffer Inter-core Process Communication Protocol for Heterogeneous Multi-core Platforms
abstract
Executing functional components in pipeline on heterogeneous multi-core platforms can greatly improve the parallelism but require great amount of data communication among processes and threads. Our studies showed that existing inter-process/thread communication protocols consist of many unnecessary memory copies and prolong the execution of the applications on heterogeneous multi-core platforms. NTU ICPC uses polling-base mail notification to unnecessary context switches, and designs a memory subsystem to manage the input and output data between the senders and receivers. The protocol was implemented and evaluated on heterogeneous multi-core platform for several use scenario including H.264 encoding process. The evaluation results show that the communication overhead on sender side is independent of the data size and that on receiver side is greatly shortened, compared to several inter-process Communication (IPC) protocols including mailbox, message queue, and shared memory. When encoding H.264 video clips, the encoding frame rates increase for more than 30%. Our experiments also showed that the communication overhead accounts 40% to 50% of total execution time in average for H.264 video decoding applications. In this paper, we present the design and implementation of zero-buffer inter-core process communication protocol, named NTU ICPC, to shorten communication overhead for pipeline executed applications on heterogeneous multi-core platforms.
Yu-Hsien Lin, Chia-Heng Tu, Chi-Sheng Shih 0001, Shih-Hao Hung
RTCSA2
2009 Machine learning-based prefetch optimization for data center applications
abstract
Performance tuning for data centers is essential and complicated. It is important since a data center comprises thousands of machines and thus a single-digit performance improvement can significantly reduce cost and power consumption. Unfortunately, it is extremely difficult as data centers are dynamic environments where applications are frequently released and servers are continually upgraded.
Shih-Wei Liao, Tzu-Han Hung, Donald Nguyen, Chinyen Chou, Chia-Heng Tu, Hucheng Zhou
SC5
2008 Optimizing the Embedded Caching and Prefetching Software on a Network-Attached Storage System
abstract
As the speed gap between memory and disk is so large today, caching and prefetch are critical to enterprise class storage applications, which demands high performance. In this paper, we present our study on performance of a mid-range storage server produced by the Quanta Computer Incorporation. We first analyzed the existing caching mechanism in the server and then developed a fast caching methodology to reduce the cache access latency and processing overhead of the storage controller. In addition, we proposed a new adaptive prefetch scheme reduces the average disk access time seen by the host. Via trace-driven simulation, we evaluated the performance of our new caching and adaptive prefetch schemes. Our results showed the performance improvement for the TPC-C on-line transaction benchmark.
Shih-Hao Hung, Chien-Cheng Wu, Chia-Heng Tu
EUC (1)3
2008 New Tracing and Performance Analysis Techniques for Embedded Applications
abstract
Performance evaluation is key to many computer applications. Many techniques and profiling tools are available for measuring performance, but most of them depend on the hardware and the software on which they run. For a new platform, or a platform which is not popular, programmers usually suffer from few analysis tools, which has been a serious problem for application development on many embedded systems. Thus, a performance analysis tool with the software mechanism is quite important for developing embedded applications. This paper describes a software mechanism for analyzing program performance on a wide range of platforms via code instrumentation at the source level. We implement this mechanism in a pure software profiling toolkit, called Module tracer, which works with a public-domain tool, CIL, to carry out code instrumentation for C programs. The toolkit aids programmers in understanding the behavior of applications by generating and analyzing traces and identify potential performance problems.
Shih-Hao Hung, Shu-Jheng Huang, Chia-Heng Tu
RTCSA3
2007 Scalable Lossless High Definition Image Coding on Multicore Platforms
Shih-Wei Liao, Shih-Hao Hung, Chia-Heng Tu, Jen-Hao Chen
EUC3