EDBT 2026 Demo / reviewers in the wild / expert
Daliang Xu
dblp:253/9200
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-6775-0688ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM InferenceabstractRunning LLMs on devices like smartphones has become a key catalyst towards privacy-preserving mobile AI. In state-of-the-art frameworks, the attention operator falls back from the dedicated NPU to the public CPU/GPU due to its sensitivity to quantization. Such a fallback leads to hampered user experience and extra system scheduling complexity. To realize NPU-centric LLM inference, this paper presents shadowNPU, a system-algorithm codesigned sparse attention module with minimal reliance on CPU/GPU by only sparsely computing a very small number of tokens. The key idea is to hide the overhead of estimating the important tokens by offloading it to NPU. On top of this, it further incorporates insightful techniques including NPU compute graph bucketing, head-wise NPU-CPU/GPU pipeline and per-head fine-grained sparsity ratio to achieve high accuracy and efficiency. Compared to design alternatives, shadowNPU achieves the best performance with strictly limited CPU/GPU resource; it requires much less CPU/GPU resource to achieve on-par performance of SoTA frameworks. Wangsong Yin, Daliang Xu, Mengwei Xu 0001, Gang Huang 0001, Xuanzhe Liu |
MobiSys | 2 |
| 2026 | VLMCache: Efficient On-Device Vision-Language Model InferenceabstractVision Language Models (VLMs) are foundational for low-latency, privacy-preserving on-device AI in real-time applications like UI agents and VQA. The VLM prefilling phase, which processes the entire visual-textual input, faces the critical challenge of a long Time-to-First-Token (TTFT). One promising approach to reduce TTFT is to exploit the temporal locality by reusing block-level computations across consecutive frames. Unfortunately, current Transformer-based VLMs break the spatial invariance of CNNs and invalidate the strict-prefix KV-cache mechanism of decoder-only LLMs; in practice, even a single-pixel mismatch can prevent reuse. Yinyuan Zhang, Daliang Xu, Chenghua Wang, Ying Zhang 0012, Mengwei Xu 0001, Gang Huang 0001 |
MobiSys | 2 |
| 2026 | Training With Integer-Only Arithmetic: Energy-Efficient Federated Learning With Mobile DSP OffloadingabstractAI is making mobile applications increasingly cooler, but also introduces serious privacy risks due to the extensive user data collection. Federated learning (FL), as a privacy-preserving machine learning paradigm, enables mobile devices to collaboratively learn a shared prediction model while keeping all training data on devices. However, a key obstacle towards practical cross-device FL training is the huge energy consumption, especially for lightweight mobile devices. Prior literature mostly optimizes the convergence speed and network communication cost. In this work, we first perform the experimental analysis of improving FL performance through low-precision training with energy-friendly Digital Signal Processor (DSP) on mobile devices. Then, we demonstrate that directly integrating the state-of-the-art INT8 (8-bit integer) training algorithm and classic FL protocols will significantly degrade the model accuracy. Finally, we propose a novel FL protocol, namelyQ-FedUpdate, incorporates two critical techniques: error-compensated aggregation and pipelined batch quantization. The former can ensure the tiny model updates be accumulated and take effects, and the latter can improve the DSP cache hit rate to reduce the context switching. Extensive experiments show that,Q-FedUpdatecan effectively reduce the on-device energy consumption by 21×, and accelerate the FL convergence by 6.1× with only 2% accuracy loss. Jinliang Yuan, Daliang Xu, Mengwei Xu 0001, Yuanchun Li 0003, Xuanzhe Liu, Yunhao Liu 0001, Shangguang Wang |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Fast On-device LLM Inference with NPUsabstractOn-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even mobile-sized LLMs (e.g., Gemma-2B) encounter unacceptably high inference latency, often bottlenecked by the prefill stage in tasks like screen UI understanding. Daliang Xu, Hao Zhang 0108, Ruiqi Liu 0001, Gang Huang 0001, Mengwei Xu 0001, Xuanzhe Liu |
ASPLOS (1) | 1 |
| 2025 | Elastic On-Device LLM ServiceabstractOn-device Large Language Models (LLMs) are transforming mobile AI, catalyzing applications like UI automation without privacy concerns. Nowadays the common practice is to deploy a single yet powerful LLM as a general task solver for multiple requests. We identify a key system challenge in this paradigm: current LLMs lack the elasticity to serve requests that have diversified Service-Level Objectives (SLOs) on inference latency. To tackle this, we present ElastiLM, an on-device LLM service that elasticizes both the model and the prompt dimension of a full LLM. It incorporates (1) a one-shot neuron-reordering method, which leverages the intrinsic permutation consistency in transformer models to generate high-quality elasticized sub-models with minimal runtime switching overhead; (2) a dual-head tiny language model, which efficiently and effectively refines the prompt and orchestrates the elastification between model and prompt. We implement such an elastic on-device LLM service on multiple COTS smartphones, and evaluate ElastiLM on both standalone NLP/mobile-agent datasets and end-to-end synthesized traces. On diverse SLOs, ElastiLM outperforms 7 strong baselines in (absolute) accuracy by up to 14.83% and 10.45% on average, with <1% TTFT switching overhead, on-par memory consumption and <100 offline GPU hours. Wangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang 0001, Mengwei Xu 0001, Xuanzhe Liu |
MobiCom | 3 |
| 2025 | EdgeLLM: Fast On-Device LLM Inference With Speculative DecodingabstractGenerative tasks, such as text generation and question answering, are essential for mobile applications. Given their inherent privacy sensitivity, executing them on devices is demanded. Nowadays, the execution of these generative tasks heavily relies on the Large Language Models (LLMs). However, the scarce device memory severely hinders the scalability of these models. We presentEdgeLLM, an efficient on-device LLM inference system for models whose sizes exceed the device's memory capacity.EdgeLLMis built atop speculative decoding, which delegates most tokens to a smaller, memory-resident (draft) LLM.EdgeLLMintegrates three novel techniques: (1) Instead of generating a fixed width and depth token tree,EdgeLLMproposes compute-efficient branch navigation and verification to pace the progress of different branches according to their accepted probability to prevent the wasteful allocation of computing resources to the wrong branch and to verify them all at once efficiently. (2) It uses a self-adaptive fallback strategy that promptly initiates the verification process when the smaller LLM generates an incorrect token. (3) To not block the generation,EdgeLLMproposes speculatively generating tokens during large LLM verification with the compute-IO pipeline. Through extensive experiments,EdgeLLMexhibits impressive token generation speed which is up to 9.3× faster than existing engines. Daliang Xu, Wangsong Yin, Hao Zhang 0108, Xin Jin 0008, Ying Zhang 0012, Shiyun Wei, Mengwei Xu 0001, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Niagara+: Scheduling Live ML Analytics Across Heterogeneous Device Processors and Edge ServersabstractIntelligent applications rely significantly on the live machine learning pipeline, a couple of deep neural network (DNN) inference services, executed on mobile devices to meet functional requirements while ensuring user data privacy. However, executing these DNN services on resource-constrained mobile devices presents a considerable challenge: low throughput and high energy consumption of inference tasks. To address this issue, we proposeNiagara+, a novel system designed to enhance throughput by jointly scheduling DNN inference services across heterogeneous processors on mobile devices and offloading services to powerful edge servers. To achieve this,Niagara+encounters two critical challenges: unpredictable workload dynamics and high scheduling complexity. To effectively tackle these challenges,Niagara+employs a predictive model to forecast incoming workload patterns and orchestrates service allocation across device heterogeneous processors and edge servers through a combination of two-step offline scheduling optimization and online service dispatching strategies. We implementedNiagara+and conducted comprehensive experiments, demonstrating its superiority over state-of-the-art approaches, reducing DNN service latency by up to 2.6× under high-bandwidth networks and 9.1× under low-bandwidth networks, while consistently meeting stringent inference latency requirements. Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Qun Wei, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | SoCFlow: Efficient and Scalable DNN Training on SoC-Clustered Edge ServersabstractSoC-Cluster, a novel server architecture composed of massive mobile system-on-chips (SoCs), is gaining popularity in industrial edge computing due to its energy efficiency and compatibility with existing mobile applications. However, we observe that the deployed SoC-Cluster servers are not fully utilized, because the hosted workloads are mostly user-triggered and have significant tidal phenomena. To harvest the free cycles, we propose to co-locate deep learning tasks on them. Daliang Xu, Mengwei Xu 0001, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (1) | 1 |
| 2024 | PieBridge: Fast and Parameter-Efficient On-Device Training via Proxy NetworksabstractOn-device training Neural Networks (NNs) has been a crucial catalyst towards privacy-preserving and personalized mobile intelligence. Recently, a novel training paradigm, namely Parameter-Efficient Training (PET), is attracting attention in both the machine learning and system community. In our preliminary measurements, we find PET well-suited for on-device scenarios; yet, its parameter efficiency does not translate coequal to time efficiency on resource-constrained devices, as the training time is dominated by the frozen layers. Wangsong Yin, Daliang Xu, Gang Huang 0001, Ying Zhang 0012, Shiyun Wei, Mengwei Xu 0001, Xuanzhe Liu |
SenSys | 2 |
| 2024 | Towards Energy-efficient Federated Learning via INT8-based Training on Mobile DSPsabstractAI is making the Web an even cooler place, but also introduces serious privacy risks due to the extensive user data collection. Federated learning (FL), as a privacy-preserving machine learning paradigm, enables mobile devices to collaboratively learn a shared prediction model while keeping all training data on devices. However, a key obstacle towards practical cross-device FL training is huge energy consumption, especially for lightweight mobile devices. In this work, we perform the first-of-its-kind analysis of improving FL performance through low-precision training with an energy-friendly Digital Signal Processor (DSP) on mobile devices. We first demonstrate that directly integrating the state-of-the-art INT8 (8-bit integer) training algorithm and classic FL protocols will significantly degrade the model accuracy. Moreover, we observe that there are still unavoidable frequent quantization operations on devices that cause extreme load stress on DSP-enabled INT8 training. To address the above challenges, we present Q-FedUpdate, an FL framework that efficiently preserves model accuracy with ultra-low energy consumption. It maintains a global full-precision model and allows the tiny model updates to be continuously accumulated, instead of being erased by the quantization. Furthermore, it introduces pipelining technology to parallel CPU-based quantization and DSP-enabled training, which reduces the floating-point computation overhead of frequent data quantization. Extensive experiments show that Q-FedUpdate can effectively reduce the on-device energy consumption by 21×, and accelerate the FL convergence by 6.1× with only 2% accuracy loss. Jinliang Yuan, Shangguang Wang, Daliang Xu, Yuanchun Li 0003, Mengwei Xu 0001, Xuanzhe Liu |
WWW | 4 |
| 2024 | Efficient, Scalable, and Sustainable DNN Training on SoC-Clustered Edge ServersabstractIn the realm of industrial edge computing, a novel server architecture known as SoC-Cluster, characterized by its aggregation of numerous mobile systems-on-chips (SoCs), has emerged as a promising solution owing to its enhanced energy efficiency and seamless integration with prevalent mobile applications. Despite its advantages, the utilization of SoC-Cluster servers remains unsatisfactory, primarily attributed to the tidal patterns of user-initiated workloads. To address such inefficiency, we introduceSoCFlow+, a pioneering framework designed to facilitate the co-location of deep learning training tasks on SoC-Cluster servers, thereby optimizing resource utilization.SoCFlow+incorporates three novel techniques tailored to mitigate the inherent limitations of commercial SoC-Cluster servers. First, it employs group-wise parallelism complemented by delayed aggregation, a strategy engineered to enhance the training efficiency and scalability of deep learning models, effectively circumventing network bottlenecks. Second, it integrates a data-parallel mixed-precision training algorithm, optimized to exploit the heterogeneous processing capabilities inherent to mobile SoCs fully. Third,SoCFlow+employs an underclocking-aware workload re-balanacing mechanism to tackle the training performance degradation caused by the thermal control of mobile SoCs. Through rigorous experimental validation,SoCFlow+achieves a convergence speedup ranging from 1.6× to 740× across 32 SoCs, compared to conventional benchmarks. Furthermore, when juxtaposed with commodity GPU servers (e.g., NVIDIA V100) under identical power constraints,SoCFlow+not only exhibits comparable training speed but also achieves a remarkable reduction in energy consumption by a factor of 2.31× to 10.23×, all while preserving convergence accuracy. Mengwei Xu 0001, Daliang Xu, Chiheng Lou, Li Zhang 0133, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Niagara: Scheduling DNN Inference Services on Heterogeneous Edge Processors
Daliang Xu, Qing Li 0028, Mengwei Xu 0001, Gang Huang 0001, Shangguang Wang, Xin Jin 0008, Yun Ma 0002, Xuanzhe Liu |
ICSOC (1) | 1 |
| 2022 | Mandheling: mixed-precision on-device DNN training with DSP offloadingabstractThis paper proposes Mandheling, the first system that enables highly resource-efficient on-device training by orchestrating mixed-precision training with on-chip Digital Signal Processor (DSP) offloading. Mandheling fully explores the advantages of DSP in integer-based numerical calculations using four novel techniques: (1) a CPU-DSP co-scheduling scheme to situationally mitigate the overhead from DSP-unfriendly operators; (2) a self-adaptive rescaling algorithm to reduce the overhead of dynamic rescaling in backward propagation; (3) a batch-splitting algorithm to improve DSP cache efficiency; (4) a DSP compute subgraph-reusing mechanism to eliminate the preparation overhead on DSP. We have fully implemented Mandheling and demonstrated its effectiveness through extensive experiments. The results show that, compared to the state-of-the-art DNN engines from TFLite and MNN, Mandheling reduces per-batch training time by 5.5X and energy consumption by 8.9X on average. In end-to-end training tasks, Mandheling reduces convergence time by up to 10.7X and energy consumption by 13.1X, with only 1.9%--2.7% accuracy loss compared to the FP32 precision setting. Daliang Xu, Mengwei Xu 0001, Qipeng Wang 0001, Shangguang Wang, Yun Ma 0002, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
MobiCom | 1 |
| 2019 | An adaptive template matching-based single object tracking algorithm with parallel acceleration
Baicheng Yan, Daliang Xu, Zhaokai Wang |
J. Vis. Commun. Image Represent. | 4 |