Yuzhan Wang

dblp:331/8823 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-1797-9131ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AdaScale: Dynamic Context-Aware DNN Scaling via Automated Adaptation Loop on Mobile Devices
abstract
Deep learning is reshaping mobile applications, with a growing trend of deploying deep neural networks (DNNs) directly to mobile and embedded devices to address real-time performance and privacy. To accommodate local resource limitations, techniques like weight compression, convolution decomposition, and specialized layer architectures have been developed. However, the dynamic and diverse deployment contexts of mobile devices pose significant challenges. Adapting deep models to meet varied device-specific requirements for latency, accuracy, memory, and energy is labor-intensive. Additionally, changing processor states, fluctuating memory availability, and competing processes frequently necessitate model recompression to preserve user experience. To address these issues, we introduce AdaScale, an elastic inference framework that automates the adaptation of deep models to dynamic contexts. AdaScale leverages a self-evolutionary model to streamline network creation, employs diverse compression operator combinations to reduce the search space and improve outcomes, and integrates a resource availability awareness block and performance profilers to establish an automated adaptation loop. Our experiments demonstrate that AdaScale significantly enhances accuracy by 5.09%, reduces training overhead by 66.89%, speeds up inference latency by 1.51 to$6.2\times $, and lowers energy costs by$4.69\times $.
Yuzhan Wang, Sicong Liu 0005, Bin Guo 0001, Boqi Zhang, Yasan Ding, Hao Luo 0022, Zhiwen Yu 0001
IEEE Internet Things J.1
2025 CrowdHMTware: A Cross-Level Co-Adaptation Middleware for Context-Aware Mobile DL Deployment
abstract
There are many deep learning (DL) powered mobile and wearable applications today continuously and unobtrusively sensing the ambient surroundings to enhance all aspects of human lives. To enable robust and private mobile sensing, DL models are often deployed locally on resource-constrained mobile devices using techniques such as model compression or offloading. However, existing methods, either front-end algorithm level (i.e. DL model compression/partitioning) or back-end scheduling level (i.e. operator/resource scheduling), cannot be locally online because they require offline retraining to ensure accuracy or rely on manually pre-defined strategies, struggle withdynamic adaptability. The primary challenge lies in feeding back runtime performance from theback-endlevel to thefront-endlevel optimization decision. Moreover, the adaptive mobile DL model porting middleware withcross-level co-adaptationis less explored, particularly in mobile environments withdiversityanddynamics. In response, we introduce CrowdHMTware, a dynamic context-adaptive DL model deployment middleware for heterogeneous mobile devices. It establishes anautomated adaptation loopbetween cross-level functional components, i.e. elastic inference, scalable offloading, and model-adaptive engine, enhancing scalability and adaptability. Experiments with four typical tasks across 15 platforms and a real-world case study demonstrate that${\sf CrowdHMTware}$can effectively scale DL model, offloading, and engine actions across diverse platforms and tasks. It hides run-time system issues from developers, reducing the required developer expertise.
Sicong Liu 0005, Bin Guo 0001, Shiyan Luo, Yuzhan Wang, Hao Luo 0022, Yuan Xu 0018, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.4
2025 AdaKnife: Flexible DNN Offloading for Inference Acceleration on Heterogeneous Mobile Devices
abstract
The integration of deep neural network (DNN) intelligence into embedded mobile devices is expanding rapidly, supporting a wide range of applications. DNN compression techniques, which adapt models to resource-constrained mobile environments, often force a trade-off between efficiency and accuracy. Distributed DNN inference, leveraging multiple mobile devices, emerges as a promising alternative to enhance inference efficiency without compromising accuracy. However, effectively decoupling DNN models into fine-grained components for optimal parallel acceleration presents significant challenges. Current partitioning methods, including layer-level and operator or channel-level partitioning, provide only partial solutions and struggle with the heterogeneous nature of DNN compilation frameworks, complicating direct model offloading. In response, we introduce AdaKnife, an adaptive framework for accelerated inference across heterogeneous mobile devices. AdaKnife enables on-demand mixed-granularity DNN partitioning via computational graph analysis, facilitates efficient cross-framework model transitions with operator optimization for offloading, and improves the feasibility of parallel partitioning using a greedy operator parallelism algorithm. Our empirical studies show that AdaKnife achieves a 66.5% reduction in latency compared to baselines.
Sicong Liu 0005, Hao Luo 0022, Bin Guo 0001, Zhiwen Yu 0001, Yuzhan Wang, Yasan Ding, Yuan Yao 0004
IEEE Trans. Mob. Comput.7
2025 Cross-F$^{2}$SCIL: A Federated Few-Shot Class Incremental Learning Method for Cross Mobile Edge Network Environments
abstract
Edge Federated Learning (EFL) has demonstrated significant potential in the field of Artificial Intelligence of Things (AIoT) by protecting data privacy and reducing communication costs. However, in real-world scenarios, multiple independent edge networks seldom collaborate due to factors such as data heterogeneity and the absence of a central server. Mobile devices, acting as bridges across different environments, offer an opportunity to enable dynamic collaboration among multiple edge networks. Nevertheless, as mobile devices transition between edge networks, they may encounter new classes with only a few samples, leading to catastrophic forgetting of previous knowledge and overfitting in new environments. To address this challenge, we propose Cross-F$^{2}$SCIL, a Federated Few-Shot Class Incremental Learning method that enables on-demand dynamic collaboration in mobile edge network environments. Cross-F$^{2}$SCIL allows mobile devices to efficiently learn new class knowledge from few-shot samples upon entering new edge networks while consolidating prior knowledge to prevent forgetting. Specifically, to mitigate the forgetting caused by new class overwriting on devices and parameter dilution at the server, we design a two-phase training framework. In the first phase, we learn a local model using Prototype Augmentation to enhance the retention of prior knowledge. In the second phase, we obtain the global model via Hierarchical Personalized Parameter Aggregation to effectively integrate learned knowledge across devices. To effectively learn new class information while reducing overfitting, we incorporate Self-Supervised Knowledge Aggregation and Prototype Knowledge Fusion to enhance model generalization and seamlessly integrate new classes into the existing model. Compared to the best-performing baseline on each dataset, Cross-F$^{2}$SCIL achieves an average improvement of 5.52% in Average Accuracy across five datasets, with the maximum improvement reaching 7.97%.
Yan Liu 0045, Bin Guo 0001, Dongzhi Wang, Yuzhan Wang, Hao Luo 0022, Zhiwen Yu 0001
IEEE Trans. Serv. Comput.7
2024 AdaMEC: Towards a Context-adaptive and Dynamically Combinable DNN Deployment Framework for Mobile Edge Computing
abstract
With the rapid development of deep learning, recent research on intelligent and interactive mobile applications (e.g., health monitoring, speech recognition) has attracted extensive attention. And these applications necessitate the mobile edge computing scheme, i.e., offloading partial computation from mobile devices to edge devices for inference acceleration and transmission load reduction. The current practices have relied on collaborative DNN partition and offloading to satisfy the predefined latency requirements, which is intractable to adapt to the dynamic deployment context at runtime. AdaMEC, a context-adaptive and dynamically combinable DNN deployment framework, is proposed to meet these requirements for mobile edge computing, which consists of three novel techniques. First, once-for-all DNN pre-partition divides DNN at the primitive operator level and stores partitioned modules into executable files, defined as pre-partitioned DNN atoms. Second, context-adaptive DNN atom combination and offloading introduces a graph-based decision algorithm to quickly search the suitable combination of atoms and adaptively make the offloading plan under dynamic deployment contexts. Third, runtime latency predictor provides timely latency feedback for DNN deployment considering both DNN configurations and dynamic contexts. Extensive experiments demonstrate that AdaMEC outperforms state-of-the-art baselines in terms of latency reduction by up to 62.14% and average memory saving by 55.21%.
Sicong Liu 0005, Bin Guo 0001, Yuzhan Wang, Hao Wang 0182, Zhenli Sheng, Zhiwen Yu 0001
ACM Trans. Sens. Networks5
2022 CAQ: Toward Context-Aware and Self-Adaptive Deep Model Computation for AIoT Applications
abstract
Artificial Intelligence of Things (AIoT) has recently accepted significant interests. Remarkably, embedded artificial intelligence (e.g., deep learning) on-device transforms IoT devices into intelligent systems that robustly and privately process data. Quantization technique is widely used to compress deep models for narrowing the resource gap between computation demands and platform supply. However, existing quantization schemes induce unsatisfaction for IoT scenarios since they are oblivious to dynamic changes of application context (e.g., battery and hierarchical memory availability) during the long-term operation. Subsequently, they will mismatch the user-desired resource efficiency and application lifetime. Also, to adapt to the dynamic context, we can neither accept the latency for model retraining with existing hand-crafted quantization nor the overhead for quantization bit width researching with prior on-demand quantization. This article presents a context-aware and self-adaptive deep model quantization (CAQ) system for IoT application scenarios. CAQ integrates a novel switchable multigate quantization framework, optimizing the quantized model accuracy and energy efficiency in diverse contexts. Based on the learned model, CAQ can switch among different gating networks in a context-aware manner and then adopt it to automatically capture the representation importance of various layers for optimal quantization bit-width selection. The experimental results show that CAQ achieves up to 50% storage savings with even 2.61% higher accuracy than the state-of-the-art baselines.
Sicong Liu 0005, Yungang Wu, Bin Guo 0001, Yuzhan Wang, Liyao Xiang, Zhetao Li, Zhiwen Yu 0001
IEEE Internet Things J.4