Tsung Tai Yeh

dblp:02/8471 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-2401-9916ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Omni-LUT: Energy-Efficient LUT-Based Accelerator with Hardware-Aware KV Cache Quantization
Cheng-Han Tsai, Kuan-Chen Chou, Yu-Hsin Wang, Chieh-Dun Wen, Tsung Tai Yeh
ISCA5
2025 EDA: Energy-Efficient Inter-Layer Model Compilation for Edge DNN Inference Acceleration
abstract
Modern handheld devices often employ neural processing units (NPUs) to accelerate deep neural network (DNN) inference applications. Unlike the AI accelerator of a data center, the NPU of an edge device has strict price and energy budgets. In an NPU, DRAM memory consumes much energy due to frequent data movement across on-chip and off-chip memory. The inter-layer operator scheduling has been shown to reduce off-chip memory transactions by reusing DNN operator outputs on the on-chip memory space. However, it often deeply traverses operators and substantially increases off-chip memory data traffic when the on-chip memory of an NPU decreases in size. Consequently, this work creates a DNN model compilation framework called EDA that transparently improves the energy efficiency of an NPU by adjusting the operator traversal depth of the inter-layer operator scheduling in a stacked DNN model. First, EDA transforms a DNN model into a tensor-splitting model. Second, EDA breaks the tensor-splitting model graph into multiple subgraphs. Third, the EDA inter-layer cost model quickly determines the depth of each subgraph. Fourth, EDA properly manages the on-chip shared memory space of an NPU to avoid overusing memory space. Finally, EDA devises the operator grouping method to improve the MAC unit and on-chip memory space utilization. EDA improves the geometric means of $2.08 \times$ and $2.39 \times$, respectively, in energy efficiency and performance over the NPU with designated memory buffers.
Bo Ren Pao, I-Chia Chen, En-Hao Chang, Tsung Tai Yeh
HPCA4
2025 AQB8: Energy-Efficient Ray Tracing Accelerator through Multi-Level Quantization
abstract
Ray tracing (RT) is a rendering technique that produces high-fidelity images by simulating how light physically interacts with objects in a scene.This realism comes at a high computational and memory cost, largely driven by the need to find numerous ray-object intersections.To accelerate this process, scenes are typically structured using bounding volume hierarchy (BVH) trees, where objects are grouped within bounding boxes to minimize the number of required intersection tests.Specialized hardware, known as RT accelerators, accelerates BVH processing to boost computational speed, yet memory traffic persists as a major bottleneck, largely due to the bandwidth consumed by standard 32-bit floating-point (FP32) bounding boxes.Consequently, previous work has aimed to reduce memory traffic by compressing these boxes into low-bit (e.g., 8-bit) representations.However, existing compression techniques typically require decompressing bounding boxes back to FP32 for intersection tests, thus failing to eliminate the computational burden and energy cost associated with complex FP32 arithmetic during BVH traversal.This work introduces AQB8, an RT accelerator designed to operate on a quantized BVH tree constructed using a novel multi-level quantization technique.This approach enables RT to operate directly on low-bit integers using simpler, area-efficient hardware units, thereby drastically reducing the need for FP32 arithmetic during BVH traversal while mitigating overheads associated with reduced precision.As a result, across various scenes, AQB8 achieves a 70% reduction in DRAM accesses, a 49% reduction in energy consumption, a 27% hardware area reduction, and a 1.82x performance speedup over modern GPU RT accelerators.Our code is publicly available at https://github.com/nycu-caslab/AQB8.
Yen-Chieh Huang, Chen-Pin Yang, Tsung Tai Yeh
ISCA3
2025 StreamNet++: Memory-Efficient Streaming TinyML Model Compilation on Microcontrollers
abstract
The rapid growth of on-device artificial intelligence increases the importance of TinyML inference applications. However, the stringent tiny memory space on the microcontroller unit (MCU) raises the grand challenge when deploying deep neural network (DNN) models on such a resource-constrained embedded system device. Traditionally, the machine learning system platform executes operators in a layer-wise manner. The layer-wise inference continues to the next operator before completing an operator. Thus, the DNN model compiler needs to allocate the SRAM memory space to store an operator’s entire input and output tensor when using the layer-wise inference on an MCU. However, the layer-wise inference will run out of memory quickly when an operator’s input and output tensor size in a DNN model is large. Consequently, the patch-based inference work divides a tensor into multiple small patches and only stores a small one to reduce the peak SRAM memory usage on an MCU. However, the computation of the overlapping patches tremendously increases the computational overhead of the patch-based inference and makes the patch-based inference undesirable on an MCU. Thus, this work presents StreamNet, a TinyML model compilation framework. StreamNet employs the stream buffer to eliminate redundant computation of patch-based inference while using small SRAM memory space on an MCU. StreamNet typically uses one type of patch configuration in a DNN model and does not completely eliminate the memory bottleneck of TinyML models. Unlike StreamNet, this article designs StreamNet++ patch-based variant inference that uses several types of patch configurations to completely remove the additional memory bottleneck even using StreamNet. Furthermore, StreamNet++ designs a parameter selection algorithm that quickly yields the best patch parameter candidates to meet the memory constraint of different MCUs. As a result, in 10 TinyML models, StreamNet++2D stream processing achieves a geometric mean of 5.7X speedup and removes 78% of redundant MACs over the latest patch-based inference.
Chen-Fong Hsu, Hong-Sheng Zheng, Yu-Yuan Liu, Tsung Tai Yeh
ACM Trans. Embed. Comput. Syst.4
2024 WER: Maximizing Parallelism of Irregular Graph Applications Through GPU Warp EqualizeR
abstract
Irregular graphs are becoming increasingly prevalent across a broad spectrum of data analysis applications. Despite their versatility, the inherent complexity and irregularity of these graphs often result in the underutilization of Single Instruction, Multiple Data (SIMD) resources when processed on Graphics Processing Units (GPUs). This underutilization originates from two primary issues: the occurrence of inactive threads and intra-warp load imbalances. These issues can produce idle threads, lead to inefficient usage of SIMD resources, consequently hamper throughput, and increase program execution time. To address these challenges, we introduce Warp EqualizeR (WER), a framework designed to optimize the utilization of SIMD resources on a GPU for processing irregular graphs. WER employs both software API and a specifically-tailored hardware microarchitecture. Such a synergistic approach enables workload redistribution in irregular graphs, which allows WER to enhance SIMD lane utilization and further harness the SIMD resources within a GPU. Our experimental results over seven different graph applications indicate that WER yields a geometric mean speedup of $2.52 \times$ and $1.47 \times$ over the baseline GPU and existing state-of-the-art methodologies, respectively.
En-Ming Huang, Bo-Wun Cheng, Meng-Hsien Lin, Chun-Yi Lee, Tsung Tai Yeh
ASPDAC5
2024 OC-DLRM: Minimizing the I/O Traffic of DLRM Between Main Memory and OCSSD
abstract
Due to the exponential growth of data in computing, DRAM-based main memory is now insufficient for data-intensive applications like machine learning and recommendation systems. This has led to a performance issue involving data transfer between main memory and storage devices. Conventional NAND-based SSDs are unable to efficiently handle this problem as they can't distinguish between data types from the host system. In contrast, open-channel SSDs (OCSSD) offer a solution by optimizing data placement from the host-side system. This research focuses on developing a new data access model for deep learning recommendation systems (DLRM) using OCSSD storage drives, called OC-DLRM. OC-DLRM reduces I/O traffic to flash memory by aggregating frequently-accessed data using the I/O unit of a flash memory drive. Our experiments show that OC-DLRM has significant performance improvement compared with traditional swapping space management techniques.
Shang-Hung Ti, Tseng-Yi Chen, Tsung Tai Yeh, Shuo-Han Chen, Yu-Pei Liang
DATE3
2024 TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers
abstract
Deploying deep neural network (DNN) models on Microcontroller Units (MCUs) is typically limited by the tightness of the SRAM memory budget. Previously, machine learning system frameworks often allocated tensor memory layer-wise, but this will result in out-of-memory exceptions when a DNN model includes a large tensor. Patch-based inference, another past solution, reduces peak SRAM memory usage by dividing a tensor into small patches and storing one small patch at a time. However, executing these overlapping small patches requires significantly more time to complete the inference and is undesirable for MCUs. We resolve these problems by developing a novel DNN model compiler: TinyTS. In the TinyTS, our tensor partition method creates a tensor-splitting model that eliminates the redundant computation observed in the patch-based inference. Furthermore, the TinyTS memory planner significantly reduces peak SRAM memory usage by releasing the memory space of unused split tensors for other ready split tensors early before the completion of the entire tensor. Finally, TinyTS presents different optimization techniques to eliminate the metadata storage and runtime overhead when executing multiple fine-grained split tensors. Using the TensorFlow Lite for Microcontroller (TFLM) framework as a baseline, we tested the effectiveness of TinyTS. We found that TinyTS reduces the peak SRAM memory usage of 9 TinyML models up to 5.92X over the baseline. TinyTS also achieves a geometric mean of 8.83X speedup over the patch-based inference. In resolving the two key issues when deploying DNN models on MCUs, TinyTS substantially boosts memory usage efficiency for TinyML applications. The source code of TinyTS can be obtained from https://github.com/nycu-caslab/TinyTS
Yu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu, Tsung Tai Yeh
HPCA5
2024 ReSA: Reconfigurable Systolic Array for Multiple Tiny DNN Tensors
abstract
Systolic array architecture has significantly accelerated deep neural networks (DNNs). A systolic array comprises multiple processing elements (PEs) that can perform multiply-accumulate (MAC). Traditionally, the systolic array can execute a certain amount of tensor data that matches the size of the systolic array simultaneously at each cycle. However, hyper-parameters of DNN models differ across each layer and result in various tensor sizes in each layer. Mapping these irregular tensors to the systolic array while fully utilizing the entire PEs in a systolic array is challenging. Furthermore, modern DNN systolic accelerators typically employ a single dataflow. However, such a dataflow is not optimal for every DNN model. This work proposes ReSA, a reconfigurable dataflow architecture that aims to minimize the execution time of a DNN model by mapping tiny tensors on the spatially partitioned systolic array. Unlike conventional systolic array architectures, the ReSA data path controller enables the execution of the input, weight, and output-stationary dataflow on PEs. ReSA also decomposes the coarse-grain systolic array into multiple small ones to reduce the fragmentation issue on the tensor mapping. Each small systolic sub-array unit relies on our data arbiter to dispatch tensors to each other through the simple interconnected network. Furthermore, ReSA reorders the memory access to overlap the memory load and execution stages to hide the memory latency when tackling tiny tensors. Finally, ReSA splits tensors of each layer into multiple small ones and searches for the best dataflow for each tensor on the host side. Then, ReSA encodes the predefined dataflow in our proposed instruction to notify the systolic array to switch the dataflow correctly. As a result, our optimization on the systolic array architecture achieves a geometric mean speedup of 1.87× over the weight-stationary systolic array architecture across nine different DNN models.
Ching-Jui Lee, Tsung Tai Yeh
ACM Trans. Archit. Code Optim.2
2023 COLAB: Collaborative and Efficient Processing of Replicated Cache Requests in GPU
abstract
In this work, we aim to capture replicated cache requests between Stream Multiprocessors (SMs) within an SM cluster to alleviate the Network-on-Chip (NoC) congestion problem of modern GPUs. To achieve this objective, we incorporate a per-cluster Cache line Ownership Lookup tABle (COLAB) that keeps track of which SM within a cluster holds a copy of a specific cache line. With the assistance of COLAB, SMs can collaboratively and efficiently process replicated cache requests within SM clusters by redirecting them according to the ownership information stored in COLAB. By servicing replicated cache requests within SM clusters that would otherwise consume precious NoC bandwidth, the heavy pressure on the NoC interconnection can be eased. Our experimental results demonstrate that the adoption of COLAB can indeed alleviate the excessive NoC pressure caused by replicated cache requests, and improve the overall system throughput of the baseline GPU while incurring minimal overhead. On average, COLAB can reduce 38% of the NoC traffic and improve instructions per cycle (IPC) by 43%.
Bo-Wun Cheng, En-Ming Huang, Chen-Hao Chao, Wei-Fang Sun, Tsung Tai Yeh, Chun-Yi Lee
ASP-DAC5
2023 StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller
abstract
With the emerging Tiny Machine Learning (TinyML) inference applications, there is a growing interest when deploying TinyML models on the low-power Microcontroller Unit (MCU). However, deploying TinyML models on MCUs reveals several challenges due to the MCU’s resource constraints, such as small flash memory, tight SRAM memory budget, and slow CPU performance. Unlike typical layer-wise inference, patch-based inference reduces the peak usage of SRAM memory on MCUs by saving small patches rather than the entire tensor in the SRAM memory. However, the processing of patch-based inference tremendously increases the amount of MACs against the layer-wise method. Thus, this notoriously computational overhead makes patch-based inference undesirable on MCUs. This work designs StreamNet that employs the stream buffer to eliminate the redundant computation of patch-based inference. StreamNet uses 1D and 2D streaming processing and provides an parameter selection algorithm that automatically improve the performance of patch-based inference with minimal requirements on the MCU’s SRAM memory space. In 10 TinyML models, StreamNet-2D achieves a geometric mean of 7.3X speedup and saves 81\% of MACs over the state-of-the-art patch-based inference.
Hong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai Yeh
NeurIPS4
2021 Deadline-Aware Offloading for High-Throughput Accelerators
abstract
Contemporary GPUs are widely used for throughput-oriented data-parallel workloads and increasingly are being considered for latency-sensitive applications in datacenters. Examples include recurrent neural network (RNN) inference, network packet processing, and intelligent personal assistants. These data parallel applications have both high throughput demands and real-time deadlines (40μs-7ms). Moreover, the kernels in these applications have relatively few threads that do not fully utilize the device unless a large batch size is used. However, batching forces jobs to wait, which increases their latency, especially when realistic job arrival times are considered.Previously, programmers have managed the tradeoffs associated with concurrent, latency-sensitive jobs by using a combination of GPU streams and advanced scheduling algorithms running on the CPU host. Although GPU streams allow the accelerator to execute multiple jobs concurrently, prior state-of-the-art solutions use the relatively distant CPU host to prioritize the latency-sensitive GPU tasks. Thus, these approaches are forced to operate at a coarse granularity and cannot quickly adapt to rapidly changing program behavior.We observe that fine-grain, device-integrated kernel schedulers efficiently meet the deadlines of concurrent, latency-sensitive GPU jobs. To overcome the limitations of software-only, CPU-side approaches, we extend the GPU queue scheduler to manage real-time deadlines. We propose a novel laxity-aware scheduler (LAX) that uses information collected within the GPU to dynamically vary job priority based on how much laxity jobs have before their deadline. Compared to contemporary GPUs, 3 state-of-the-art CPU-side schedulers and 6 other advanced GPU-side schedulers, LAX meets the deadlines of 1.7X - 5.0X more jobs and provides better energy-efficiency, throughput, and 99-percentile tail latency.
Tsung Tai Yeh, Matthew D. Sinclair, Bradford M. Beckmann, Timothy G. Rogers
HPCA1
2020 Dimensionality-Aware Redundant SIMT Instruction Elimination
abstract
In massively multithreaded architectures, redundantly executing the same instruction with the same operands in different threads is a significant source of inefficiency. This paper introduces Dimensionality-Aware Redundant SIMT Instruction Elimination (DARSIE), a non-speculative instruction skipping mechanism to reduce redundant operations in GPUs. DARSIE uses static markings from the compiler and information obtained at kernel launch time to skip redundant instructions before they are fetched, keeping them out of the pipeline. DARSIE exploits a new observation that there is significant redundancy across warp instructions in multi-dimensional threadblocks.
Tsung Tai Yeh, Roland N. Green, Timothy G. Rogers
ASPLOS1
2017 Pagoda: Fine-Grained GPU Resource Virtualization for Narrow Tasks
abstract
Massively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU.
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers
PPoPP1
2016 POSTER: Pagoda: A Runtime System to Maximize GPU Utilization in Data Parallel Tasks with Limited Parallelism
abstract
Massively multithreaded GPUs achieve high throughput by running thousands of threads in parallel. To fully utilize the hardware, contemporary workloads spawn work to the GPU in bulk by launching large tasks, where each task is a kernel that contains thousands of threads that occupy the entire GPU.
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, Timothy G. Rogers
PACT1
2015 An Energy-Efficient and Reliable Storage Mechanism for Data-Intensive Academic Archive Systems
abstract
Previous studies proposed energy-efficient solutions, such as multispeed disks and disk spin-down methods, to conserve power in their respective storage systems. However, in most cases, the authors did not analyze the reliability of their solutions. According to research conducted by Google and the IDEMA standard, frequently setting the disk status to standby mode will increase the disk’s Annual Failure Rate and reduce its lifespan. To resolve the issue, we propose an evaluation function called E 3 SaRC (Economic Evaluation of Energy Saving with Reliability Constraint), which considers the cost of hardware failure when applying energy-saving schemes. We also present an adaptive write cache mechanism called CacheRAID. The mechanism tries to mitigate the random access problems that implicitly exist in RAID techniques and thereby reduce the energy consumption of RAID disks. CacheRAID also addresses the issue of system reliability by applying a control mechanism to the spin-down algorithm. Our experimental results show that the CacheRAID storage system can reduce the power consumption of the conventional software RAID 5 system by 65% to 80%. Moreover, according to the E 3 SaRC measurement, the overall saved cost of CacheRAID is the largest among the systems that we compared.
Tseng-Yi Chen, Hsin-Wen Wei, Tsung Tai Yeh, Tsan-sheng Hsu, Wei-Kuan Shih
ACM Trans. Storage3