Haolin Chu

dblp:373/8418 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0000-7184-3150ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Embedded and real-time systems · 36% Memory systems · 31% Hardware accelerators and domain-specific architectures · 16%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Embedded and real-time systems
embedded machine learning
1.012026
RAMS: Runtime Adaptive Memory Scaling for Tiny Deep Learning on IoT Devices · IEEE Trans. Mob. Comput. 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
Act Before It's Too Late: Power-Efficient LLM Inference on Mobile Device · MobiSys 2026
Memory systems › memory management
memory footprint reduction
1.012026
RAMS: Runtime Adaptive Memory Scaling for Tiny Deep Learning on IoT Devices · IEEE Trans. Mob. Comput. 2026
Memory systems
memory management
1.012026
RAMS: Runtime Adaptive Memory Scaling for Tiny Deep Learning on IoT Devices · IEEE Trans. Mob. Comput. 2026
Embedded and real-time systems › embedded machine learning
TinyML
1.012026
RAMS: Runtime Adaptive Memory Scaling for Tiny Deep Learning on IoT Devices · IEEE Trans. Mob. Comput. 2026
Performance modeling and evaluation
profiling
0.712023
nnPerf: Demystifying DNN Runtime Inference Latency on Mobile Platforms · SenSys 2023
Energy-efficient computing
power management
0.312026
Act Before It's Too Late: Power-Efficient LLM Inference on Mobile Device · MobiSys 2026
Machine learning › Efficient and distributed learning › on-device inference
mobile inference
0.212023
nnPerf: Demystifying DNN Runtime Inference Latency on Mobile Platforms · SenSys 2023
Machine learning › Efficient and distributed learning
on-device inference
0.212023
nnPerf: Demystifying DNN Runtime Inference Latency on Mobile Platforms · SenSys 2023
GPUs and heterogeneous computing › deep learning on GPUs
GPU inference
0.212023
nnPerf: Demystifying DNN Runtime Inference Latency on Mobile Platforms · SenSys 2023

Methods — techniques the papers use, named apart from their topics

offline planning · 1.0cache-aware memory management · 1.0GPU stall analysis · 1.0
YearPublicationVenuePosition
2026 It Takes Two: Embracing Sparsity and Speculative Decoding for Efficient LLM Inference
Haolin Chu, Changyu Chen, Jian Luan 0001, Jiabin Deng, Huadong Ma, Xiaolong Zheng 0002
IWQoS1
2026 Act Before It's Too Late: Power-Efficient LLM Inference on Mobile Device
abstract
This paper presents TurboInfer, a system that enables power-efficient LLM inference on mobile devices. The core insight behind TurboInfer is that while LLMs are power-intensive due to their heavy computational demands, the model inference experiences unavoidable GPU stalls caused by tensor preparation for subsequent kernel executions at run-time. These GPU stalls arise from the unique host-controlled execution pipeline tailored to mobile phones and the significant DRAM access contention inherent to the shared memory architecture of mobile System-on-chips (SoCs). With LLM inference requiring hundreds to thousands of kernel executions, these short but frequent GPU stalls accumulate, accounting for over 74% of the token generation latency.
Haolin Chu, Jinxiao Fan, Jiabin Deng, Bensong Yu, Liguang Xie, Liang Liu 0001, Huadong Ma, Xiaolong Zheng 0002
MobiSys1
2026 RAMS: Runtime Adaptive Memory Scaling for Tiny Deep Learning on IoT Devices
abstract
Deploying Tiny Deep Learning (TinyDL) on Internet of Things (IoT) devices is gaining popularity. To accommodate the limited memory, recent methods split tensors into fine-grained parts and plan memory offline to minimize its footprint. However, they fail to adapt to dynamic memory, missing the opportunity to utilize temporarily available memory for faster inference. Additionally, existing approaches focus solely on minimizing memory size while neglecting cache usage characteristics, resulting in frequent cache misses and increased latency. In this paper, we propose RAMS, an efficient framework supporting runtime adaptive memory scaling to fully utilize the dynamic memory. We also propose a cache-friendly memory management approach that minimizes cache miss times. RAMS includes an offline planner to minimize the memory footprint essential for inference and an online manager to determine memory sizes and generate layouts for size-controllable tensors based on available memory. RAMS significantly reduces inference latency while maintaining a compact memory footprint. Extensive experiments on commercial devices running RTOS and Android systems demonstrate that, compared to the state-of-the-art methods, RAMS can efficiently reduce latency by up to 1.57× and 1.48× compared to TFLM and TinyTS, respectively using a comparable memory footprint, while reducing power consumption by 67.74% and 15.97%.
Haolin Chu, Haiteng Xin, Xiaolong Zheng 0002, Liang Liu 0001, Huadong Ma
IEEE Trans. Mob. Comput.1
2024 LLMAir: Adaptive Reprogramming Large Language Model for Air Quality Prediction
abstract
Accurate and timely air quality prediction is crucial for cities and individuals to effectively take necessary precautions against potential air pollution. Existing studies typically rely on building prediction models based on large-scale monitoring data, often designed for specific tasks. Recently, pre-trained large language models (LLMs) have achieved significant progress in various time series analysis tasks due to their powerful representation and inference capabilities. However, their application to air quality data with spatio-temporal features remains largely unexplored. In this work, we propose LLMAir, an adaptive reprogramming approach that adapts pre-trained LLMs for air quality prediction. We first construct spatiotemporal tokens based on monitoring stations by integrating value, node, and time embeddings. Next, we design an adaptive semantic-enhanced reprogramming module to compute similarity matching scores between our spatiotemporal tokens and pre-trained word embeddings for alignment. We employ a semantic regulator to generate the optimal length of word prototypes, which serve as prompt prefixes for adaptive reprogramming and guiding the spatiotemporal token embeddings into the frozen LLM. Additionally, we jointly optimize predictive error and alignment loss to train our model. Experimental results demonstrate that LLMAir achieves state-of-the-art performance in air quality prediction and few-shot forecasting across two real-world datasets.
Jinxiao Fan, Haolin Chu, Liang Liu 0001, Huadong Ma
ICPADS2
2023 nnPerf: Demystifying DNN Runtime Inference Latency on Mobile Platforms
abstract
We present nnPerf, a real-time on-device profiler designed to collect and analyze the DNN model run-time inference latency on mobile platforms. nnPerf demystifies the hidden layers and metrics used for pursuing DNN optimizations and adaptations at the granularity of operators and kernels, ensuring every facet contributing to a DNN model's run-time efficiency is easily accessible to mobile developers via well-defined APIs. With nnPerf, the mobile developers can easily identify the bottleneck in model run-time efficiency and optimize the model architecture to meet system-level objectives (SLO). We implement nnPerf on TFLite framework and evaluate its e2e-, operator-, and kernel-latency profiling accuracy across four mobile platforms. The results show that nnPerf achieves consistently high latency profiling accuracy on both CPU (98.12%) and GPU (99.87%). Our benchmark studies demonstrate that running nnPerf on mobile devices introduces the minimum overhead to model inference, with 0.231% and 0.605% extra inference latency and power consumption. We further run a case study to show how we leverage nnPerf to migrate OFA, a SOTA NAS system, to kernel-oriented model optimization on GPUs.
Haolin Chu, Xiaolong Zheng 0002, Liang Liu 0001, Huadong Ma
SenSys1