Wangsong Yin

dblp:356/3416 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0009-0000-6242-4368ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
abstract
Running LLMs on devices like smartphones has become a key catalyst towards privacy-preserving mobile AI. In state-of-the-art frameworks, the attention operator falls back from the dedicated NPU to the public CPU/GPU due to its sensitivity to quantization. Such a fallback leads to hampered user experience and extra system scheduling complexity. To realize NPU-centric LLM inference, this paper presents shadowNPU, a system-algorithm codesigned sparse attention module with minimal reliance on CPU/GPU by only sparsely computing a very small number of tokens. The key idea is to hide the overhead of estimating the important tokens by offloading it to NPU. On top of this, it further incorporates insightful techniques including NPU compute graph bucketing, head-wise NPU-CPU/GPU pipeline and per-head fine-grained sparsity ratio to achieve high accuracy and efficiency. Compared to design alternatives, shadowNPU achieves the best performance with strictly limited CPU/GPU resource; it requires much less CPU/GPU resource to achieve on-par performance of SoTA frameworks.
Wangsong Yin, Daliang Xu, Mengwei Xu 0001, Gang Huang 0001, Xuanzhe Liu
MobiSys1
2026 An Efficient Context Management System for On-Device LLMaaS
abstract
Large language models (LLMs) are renovating the mobile AI, catalyzing novel mobile applications such as UI task automation. A new paradigm of mobile AI ecosystem emerges in the LLM era: LLM as a mobile OS service (LLMaaS), where LLM runs as a system service and exposes its functionality (language understanding and generation) to third-party apps. As a giant step towards on-device LLMaaS, this work presents Libra, a system that tackles with challenge in managing persistent LLM contexts (KV cache) under tight memory constraint. Libra manages the LLM contexts based on the fine-grained, chunk-wise, globally-optimized KV cache compression and swapping. Specifically, it integrates three novel techniques: (1) Tolerance-Aware Compression applies different compression rates to each chunk based on their attention scores. (2) Swapping-Recompute Pipeline simultaneously swaps and recomputes the LLM contexts to improve the hardware resource utilization. (3) Chunk Lifecycle Management judiciously determines when and what to evict to reduce the context switching overhead. Through comprehensive experiments on various edge devices, Libra reduces the context switching latency by up to 20 × and on average 9.7 × compared to competitive baselines.
Wangsong Yin, Mengwei Xu 0001, Yuanchun Li 0003, Xuanzhe Liu
SenSys1
2025 Elastic On-Device LLM Service
abstract
On-device Large Language Models (LLMs) are transforming mobile AI, catalyzing applications like UI automation without privacy concerns. Nowadays the common practice is to deploy a single yet powerful LLM as a general task solver for multiple requests. We identify a key system challenge in this paradigm: current LLMs lack the elasticity to serve requests that have diversified Service-Level Objectives (SLOs) on inference latency. To tackle this, we present ElastiLM, an on-device LLM service that elasticizes both the model and the prompt dimension of a full LLM. It incorporates (1) a one-shot neuron-reordering method, which leverages the intrinsic permutation consistency in transformer models to generate high-quality elasticized sub-models with minimal runtime switching overhead; (2) a dual-head tiny language model, which efficiently and effectively refines the prompt and orchestrates the elastification between model and prompt. We implement such an elastic on-device LLM service on multiple COTS smartphones, and evaluate ElastiLM on both standalone NLP/mobile-agent datasets and end-to-end synthesized traces. On diverse SLOs, ElastiLM outperforms 7 strong baselines in (absolute) accuracy by up to 14.83% and 10.45% on average, with <1% TTFT switching overhead, on-par memory consumption and <100 offline GPU hours.
Wangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang 0001, Mengwei Xu 0001, Xuanzhe Liu
MobiCom1
2025 EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding
abstract
Generative tasks, such as text generation and question answering, are essential for mobile applications. Given their inherent privacy sensitivity, executing them on devices is demanded. Nowadays, the execution of these generative tasks heavily relies on the Large Language Models (LLMs). However, the scarce device memory severely hinders the scalability of these models. We presentEdgeLLM, an efficient on-device LLM inference system for models whose sizes exceed the device's memory capacity.EdgeLLMis built atop speculative decoding, which delegates most tokens to a smaller, memory-resident (draft) LLM.EdgeLLMintegrates three novel techniques: (1) Instead of generating a fixed width and depth token tree,EdgeLLMproposes compute-efficient branch navigation and verification to pace the progress of different branches according to their accepted probability to prevent the wasteful allocation of computing resources to the wrong branch and to verify them all at once efficiently. (2) It uses a self-adaptive fallback strategy that promptly initiates the verification process when the smaller LLM generates an incorrect token. (3) To not block the generation,EdgeLLMproposes speculatively generating tokens during large LLM verification with the compute-IO pipeline. Through extensive experiments,EdgeLLMexhibits impressive token generation speed which is up to 9.3× faster than existing engines.
Daliang Xu, Wangsong Yin, Hao Zhang 0108, Xin Jin 0008, Ying Zhang 0012, Shiyun Wei, Mengwei Xu 0001, Xuanzhe Liu
IEEE Trans. Mob. Comput.2
2024 PieBridge: Fast and Parameter-Efficient On-Device Training via Proxy Networks
abstract
On-device training Neural Networks (NNs) has been a crucial catalyst towards privacy-preserving and personalized mobile intelligence. Recently, a novel training paradigm, namely Parameter-Efficient Training (PET), is attracting attention in both the machine learning and system community. In our preliminary measurements, we find PET well-suited for on-device scenarios; yet, its parameter efficiency does not translate coequal to time efficiency on resource-constrained devices, as the training time is dominated by the frozen layers.
Wangsong Yin, Daliang Xu, Gang Huang 0001, Ying Zhang 0012, Shiyun Wei, Mengwei Xu 0001, Xuanzhe Liu
SenSys1