Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ruiqi Liu 0001

dblp:167/4246-1 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0004-3188-0976ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 71% Generative modeling · 29%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 87% Embedded and real-time systems · 13%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling › diffusion model
diffusion model inference
0.912025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Fast On-device LLM Inference with NPUs · ASPLOS (1) 2025
Machine learning › Efficient and distributed learning › model compression
look-up table
0.912025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Machine learning › Efficient and distributed learning › on-device inference
mobile inference
0.912025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.912025
Fast On-device LLM Inference with NPUs · ASPLOS (1) 2025
Hardware accelerators and domain-specific architectures › efficient inference
on-device LLM inference
0.912025
Fast On-device LLM Inference with NPUs · ASPLOS (1) 2025
Machine learning › Generative modeling
diffusion model
0.312025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Machine learning › Generative modeling
image generation
0.312025
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices · IEEE Trans. Mob. Comput. 2025
Embedded and real-time systems › on-device inference
mobile inference
0.312025
Fast On-device LLM Inference with NPUs · ASPLOS (1) 2025

Methods — techniques the papers use, named apart from their topics

inference optimization · 1.7NPU acceleration · 1.7quantization · 0.9lookup table · 0.9CPU-GPU co-scheduling · 0.9
YearPublicationVenuePosition
2025 Fast On-device LLM Inference with NPUs
abstract
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even mobile-sized LLMs (e.g., Gemma-2B) encounter unacceptably high inference latency, often bottlenecked by the prefill stage in tasks like screen UI understanding.
Daliang Xu, Hao Zhang 0108, Ruiqi Liu 0001, Gang Huang 0001, Mengwei Xu 0001, Xuanzhe Liu
ASPLOS (1)4
2025 Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices
abstract
Diffusion models have revolutionized image synthesis applications. Many studies focus on using approximate computation such as model quantization to reduce inference costs on mobile devices. However, due to their extensive model parameters and autoregressive inference fashion, the overhead of diffusion models remains high, which is challenging for mobile devices to handle. To reduce the inference overhead of diffusion models on mobile devices, we proposeLUT-Diff, an algorithm-system co-design specifically tailored for mobile device diffusion model inference optimization.LUT-Diffoptimizes using lookup tables and can efficiently generate a series of lookup table candidates for diffusion models without end-to-end training. During inference,LUT-Diffadaptively selects the best inference strategy based on the application/user's latency budget. Additionally,LUT-Diffincludes a parallel inference engine that rapidly completes model inference through CPU-GPU co-scheduling. Extensive experiments demonstrate thatLUT-Diffcan generate images comparable to the original model, with an up to 0.012 MSE in generated images.LUT-Diffcan also achieve up to 9.1× inference acceleration and reduce the inference memory footprint by up to 70.9% compared to baseline methods. Moreover,LUT-Diffcan save at least 3281× the learning cost of lookup tables.
Qipeng Wang 0001, Shiqi Jiang 0002, Yifan Yang 0004, Ruiqi Liu 0001, Yuanchun Li 0003, Ting Cao 0003, Xuanzhe Liu
IEEE Trans. Mob. Comput.4