VLDB 2026 Research / reviewers in the wild / expert
Yu-Yuan Liu
dblp:369/7127
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
0009-0003-6796-8761ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Embedded and real-time systems · 50% Hardware accelerators and domain-specific architectures · 43% Memory systems · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
compiler optimization |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Embedded and real-time systems
embedded machine learning |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Embedded and real-time systems › embedded machine learning
TinyML deployment |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Hardware accelerators and domain-specific architectures › efficient inference
memory-efficient inference |
0.7 | 1 | 2023 | StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller · NeurIPS 2023 |
Hardware accelerators and domain-specific architectures › edge accelerator
microcontroller inference |
0.7 | 1 | 2023 | StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller · NeurIPS 2023 |
Memory systems
memory management |
0.2 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Machine learning › Efficient and distributed learning
memory optimization |
0.2 | 1 | 2023 | StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller · NeurIPS 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2023 | StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the Microcontroller · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
tensor partition · 1.5patch-based inference · 1.5memory planning · 1.5stream buffer · 1.3parameter selection algorithm · 1.31d and 2d streaming processing · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | StreamNet++: Memory-Efficient Streaming TinyML Model Compilation on MicrocontrollersabstractThe rapid growth of on-device artificial intelligence increases the importance of TinyML inference applications. However, the stringent tiny memory space on the microcontroller unit (MCU) raises the grand challenge when deploying deep neural network (DNN) models on such a resource-constrained embedded system device. Traditionally, the machine learning system platform executes operators in a layer-wise manner. The layer-wise inference continues to the next operator before completing an operator. Thus, the DNN model compiler needs to allocate the SRAM memory space to store an operator’s entire input and output tensor when using the layer-wise inference on an MCU. However, the layer-wise inference will run out of memory quickly when an operator’s input and output tensor size in a DNN model is large. Consequently, the patch-based inference work divides a tensor into multiple small patches and only stores a small one to reduce the peak SRAM memory usage on an MCU. However, the computation of the overlapping patches tremendously increases the computational overhead of the patch-based inference and makes the patch-based inference undesirable on an MCU. Thus, this work presents StreamNet, a TinyML model compilation framework. StreamNet employs the stream buffer to eliminate redundant computation of patch-based inference while using small SRAM memory space on an MCU. StreamNet typically uses one type of patch configuration in a DNN model and does not completely eliminate the memory bottleneck of TinyML models. Unlike StreamNet, this article designs StreamNet++ patch-based variant inference that uses several types of patch configurations to completely remove the additional memory bottleneck even using StreamNet. Furthermore, StreamNet++ designs a parameter selection algorithm that quickly yields the best patch parameter candidates to meet the memory constraint of different MCUs. As a result, in 10 TinyML models, StreamNet++2D stream processing achieves a geometric mean of 5.7X speedup and removes 78% of redundant MACs over the latest patch-based inference. Chen-Fong Hsu, Hong-Sheng Zheng, Yu-Yuan Liu, Tsung Tai Yeh |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersabstractDeploying deep neural network (DNN) models on Microcontroller Units (MCUs) is typically limited by the tightness of the SRAM memory budget. Previously, machine learning system frameworks often allocated tensor memory layer-wise, but this will result in out-of-memory exceptions when a DNN model includes a large tensor. Patch-based inference, another past solution, reduces peak SRAM memory usage by dividing a tensor into small patches and storing one small patch at a time. However, executing these overlapping small patches requires significantly more time to complete the inference and is undesirable for MCUs. We resolve these problems by developing a novel DNN model compiler: TinyTS. In the TinyTS, our tensor partition method creates a tensor-splitting model that eliminates the redundant computation observed in the patch-based inference. Furthermore, the TinyTS memory planner significantly reduces peak SRAM memory usage by releasing the memory space of unused split tensors for other ready split tensors early before the completion of the entire tensor. Finally, TinyTS presents different optimization techniques to eliminate the metadata storage and runtime overhead when executing multiple fine-grained split tensors. Using the TensorFlow Lite for Microcontroller (TFLM) framework as a baseline, we tested the effectiveness of TinyTS. We found that TinyTS reduces the peak SRAM memory usage of 9 TinyML models up to 5.92X over the baseline. TinyTS also achieves a geometric mean of 8.83X speedup over the patch-based inference. In resolving the two key issues when deploying DNN models on MCUs, TinyTS substantially boosts memory usage efficiency for TinyML applications. The source code of TinyTS can be obtained from https://github.com/nycu-caslab/TinyTS Yu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu, Tsung Tai Yeh |
HPCA | 1 |
| 2023 | StreamNet: Memory-Efficient Streaming Tiny Deep Learning Inference on the MicrocontrollerabstractWith the emerging Tiny Machine Learning (TinyML) inference applications, there is a growing interest when deploying TinyML models on the low-power Microcontroller Unit (MCU). However, deploying TinyML models on MCUs reveals several challenges due to the MCU’s resource constraints, such as small flash memory, tight SRAM memory budget, and slow CPU performance. Unlike typical layer-wise inference, patch-based inference reduces the peak usage of SRAM memory on MCUs by saving small patches rather than the entire tensor in the SRAM memory. However, the processing of patch-based inference tremendously increases the amount of MACs against the layer-wise method. Thus, this notoriously computational overhead makes patch-based inference undesirable on MCUs. This work designs StreamNet that employs the stream buffer to eliminate the redundant computation of patch-based inference. StreamNet uses 1D and 2D streaming processing and provides an parameter selection algorithm that automatically improve the performance of patch-based inference with minimal requirements on the MCU’s SRAM memory space. In 10 TinyML models, StreamNet-2D achieves a geometric mean of 7.3X speedup and saves 81\% of MACs over the state-of-the-art patch-based inference. Hong-Sheng Zheng, Yu-Yuan Liu, Chen-Fong Hsu, Tsung Tai Yeh |
NeurIPS | 2 |