Ziyan Fu 0001

dblp:234/4932-1 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-1039-0386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Malva: A Jitter-Aware Online Pruning Framework for DNN Inference Tasks
abstract
In fields like autonomous driving, strict constraints are imposed on the computing latency of deep neural network (DNN) inference tasks on edge servers. However, it is typical for edge servers to execute multiple tasks in parallel to serve multiple users, causing severe latency jitter due to resource competition, which seriously affects timeliness. Existing works ignore the computing jitter and regard computing latency as a deterministic value, failing to meet the timeliness requirement. To address this issue, we propose Malva, a framework for finegrained online pruning for DNN tasks, allowing flexible pruning at runtime based on jitter conditions. Specifically, Malva first partitions the DNN model into blocks and applies early exiting and pruning methods to create block variants. Then, the Malva scheduler flexibly selects the variant to be executed or exits early according to urgency and jitter conditions. Moreover, we propose a novel urgency-aware prediction strategy to estimate the accuracy impact of variants with incomplete pathway information during scheduling. Stress testing shows Malva can strictly maintain a zero deadline miss rate and significantly increase the stress required to cause the first deadline miss while still outperforming state-of-the-art methods in accuracy.
Ziyan Fu 0001, Yongheng Deng, Yingjun Wu, Zhibo Wang 0001, Su Yao, Yaoxue Zhang, Ju Ren 0001
IWQoS1
2025 Squeezer: Efficient Multi-DNN Inference for Edge Video Analytics via Cross-Model Scheduling
abstract
Video analytics at the edge is becoming increasingly prevalent in many scenarios, such as smart campuses and intelligent factories. These applications often consist of multiple subtasks, which necessitates the optimization for multi-DNN (Deep Neural Network) inference. Due to limited consideration over cross-model scheduling, current practices cannot fully leverage available computing resources, leading to suboptimal performance. To address this, we propose Squeezer, a multiDNN serving framework that holistically schedules multiple DNN models on an edge server with a single GPU. Squeezer decouples the cross-model scheduling into a two-layered approach, which involves (1) balanced operator grouping which partitions operators of multiple DNN models into groups, significantly reducing the scheduling complexity and (2) kernel scheduler which orchestrates parallel execution within each group by considering the interplay among kernels running in parallel, thereby enabling cross-model optimizations in multi-DNN inference. Performance evaluation results demonstrate that Squeezer outperforms state-of-the-art baselines, achieving up to 1.91× improvement in system throughput.
Lingxiao Ma, Ziyan Fu 0001, Yuanchun Li 0003, Ju Ren 0001, Yaoxue Zhang, Yunxin Liu 0001
IEEE Trans. Mob. Comput.3
2022 Kalmia: A Heterogeneous QoS-aware Scheduling Framework for DNN Tasks on Edge Servers
abstract
Motivated by the popularity of edge intelligence, DNN services have been widely deployed at the edge, posing significant performance pressure on edge servers. How to improve the QoS of edge DNN services becomes a crucial and challenging problem. Previous works, however, did not fully consider the heterogeneous QoS requirements on urgent and non-urgent tasks, causing frequent QoS violations. Meanwhile, our empirical study shows that severe task interference exists in concurrent DNN tasks, further degrading the timeliness of urgent tasks and throughput of non-urgent tasks. To address these issues, we propose Kalmia, a heterogeneous QoS-aware framework for DNN inference task scheduling on edge servers. Specifically, Kalmia includes an offline profiling stage and an online scheduling policy. In offline profiling, we build a regression model to predict the execution time of tasks. During online scheduling, we classify the tasks into urgent and non-urgent tasks and distribute them into two CUDA contexts. By a tailored scheduling strategy, non-urgent tasks can fully utilize the computing resources for throughput improvement, while the timeliness of urgent tasks can be guaranteed via preemption. Experimental results demonstrate that Kalmia can achieve up to 2.8× improvement in throughput and significantly reduce the deadline violation rate compared with state-of-the-art methods.
Ziyan Fu 0001, Ju Ren 0001, Yue-Zhi Zhou, Yaoxue Zhang
INFOCOM1
2022 Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCL
abstract
Despite the significant development of mobile device SoCs, they are still inefficient in computing computation-intensive workloads, such as high-resolution image processing and AR/VR applications. Offloading offers a promising way to leverage cloud or edge servers for acceleration, but existing offloading is limited to specific tasks or specific hardware/software platforms, resulting in significant engineering overhead. To address this problem, we focus on the underlying layer of these applications (i.e., OpenCL) and propose Hyperion, a generic and distributed mobile offloading framework built on OpenCL. To achieve high-performance distributed execution for Hyperion, we first take a deep insight into the OpenCL data structures and design regularity-aware kernel analyzer to analyze the data dependency of work-groups and identify the essential data to offload. Then, context-aware execution time predictor is proposed to estimate the computing time of a given partitioned kernel workload that is highly impacted by many runtime factors. These techniques are integrated into pipeline-enabled and network-adaptive scheduler to make scheduling decisions, which coordinates the kernel partition and workload scheduling to form pipeline processing between data transmission and distributed execution with flexible adaptability to network dynamics. Extensive experimental results demonstrate that Hyperion achieves superior performance with an average 3.80× speedup compared with the best baseline and flexible adaptation to dynamic network conditions and available computing resources.
Ziyan Fu 0001, Ju Ren 0001, Yunxin Liu 0001, Ting Cao 0003, Yue-Zhi Zhou, Yaoxue Zhang
SenSys1
2021 Joint Optimization of Data Transfer and Co-Execution for DNN in Edge Computing
abstract
Deep learning plays an increasingly important role in human life. However, resource-constrained IoT devices are still inefficient in performing deep neural network (DNN) inference. Existing works have attempted to improve the performance by leveraging edge computing that partitions DNN and offloads a part of workloads to the edge server. However, most of them focus on scheduling workloads among different devices, ignoring the network costs. Thus, user experience easily suffers from inferior conditions such as network congestion. To address this issue, we jointly consider network conditions and computing capabilities of the IoT device and edge server, and then propose FastCoDNN, a co-execution framework that enables high-performance DNN inference. Specifically, under inferior conditions, we conduct redundant calculation instead of data synchronization to reduce network costs. Then, we orchestrate redundant calculations and data synchronizations among the local and edge, and flexibly adjust them according to new network conditions. We propose a new algorithm based on dynamic programming to achieve this adjustment. Experimental results show that FastCoDNN achieves fewer network costs and much performance improvement compared with existing methods.
Ziyan Fu 0001, Yue-Zhi Zhou, Chao Wu 0002, Yaoxue Zhang
ICC1
2019 Enabling Flexible Resource Allocation in Mobile Deep Learning Systems
abstract
Deep learning provides new opportunities for mobile applications to achieve higher performance than before. Rather, the deep learning implementation on mobile device today is largely demanding on expensive resource overheads, imposes a significant burden on the battery life and limited memory space. Existing methods either utilize cloud or edge infrastructure that require to upload user data, however, resulting in a risk of privacy leakage and large data transfers; or adopt compressed deep models, nevertheless, downgrading the algorithm accuracy. This paper provides DeepShark, a platform to enable mobile devices with the ability of flexible resource allocation in using commercial-off-the-shelf (COTS) deep learning systems. Compared to existing approaches, DeepShark seeks a balanced point between time and memory efficiency by user requirements, breaks down sophisticated deep model into code block stream and incrementally executes such blocks on system-on-chip (SoC). Thus, DeepShark requires significantly less memory space on mobile device and achieves the default accuracy. In addition, all referred user data of model processing is handled locally, thus to avoid unnecessary data transfer and network latency. DeepShark is now developed on two COTS deep learning systems, i.e., Caffe and TensorFlow. The experimental evaluations demonstrate its effectiveness in the aspects of memory space and energy cost.
Chao Wu 0002, Lan Zhang 0002, Qiushi Li 0002, Ziyan Fu 0001, Wenwu Zhu 0001, Yaoxue Zhang
IEEE Trans. Parallel Distributed Syst.4