EDBT 2026 Demo / reviewers in the wild / expert
Yu Fang Hu
dblp:373/2497
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Embedded and real-time systems · 87% Memory systems · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
compiler optimization |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Embedded and real-time systems
embedded machine learning |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Embedded and real-time systems › embedded machine learning
TinyML deployment |
0.8 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Memory systems
memory management |
0.2 | 1 | 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on Microcontrollers · HPCA 2024 |
Methods — techniques the papers use, named apart from their topics
tensor partition · 1.5patch-based inference · 1.5memory planning · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TinyTS: Memory-Efficient TinyML Model Compiler Framework on MicrocontrollersabstractDeploying deep neural network (DNN) models on Microcontroller Units (MCUs) is typically limited by the tightness of the SRAM memory budget. Previously, machine learning system frameworks often allocated tensor memory layer-wise, but this will result in out-of-memory exceptions when a DNN model includes a large tensor. Patch-based inference, another past solution, reduces peak SRAM memory usage by dividing a tensor into small patches and storing one small patch at a time. However, executing these overlapping small patches requires significantly more time to complete the inference and is undesirable for MCUs. We resolve these problems by developing a novel DNN model compiler: TinyTS. In the TinyTS, our tensor partition method creates a tensor-splitting model that eliminates the redundant computation observed in the patch-based inference. Furthermore, the TinyTS memory planner significantly reduces peak SRAM memory usage by releasing the memory space of unused split tensors for other ready split tensors early before the completion of the entire tensor. Finally, TinyTS presents different optimization techniques to eliminate the metadata storage and runtime overhead when executing multiple fine-grained split tensors. Using the TensorFlow Lite for Microcontroller (TFLM) framework as a baseline, we tested the effectiveness of TinyTS. We found that TinyTS reduces the peak SRAM memory usage of 9 TinyML models up to 5.92X over the baseline. TinyTS also achieves a geometric mean of 8.83X speedup over the patch-based inference. In resolving the two key issues when deploying DNN models on MCUs, TinyTS substantially boosts memory usage efficiency for TinyML applications. The source code of TinyTS can be obtained from https://github.com/nycu-caslab/TinyTS Yu-Yuan Liu, Hong-Sheng Zheng, Yu Fang Hu, Chen-Fong Hsu, Tsung Tai Yeh |
HPCA | 3 |