Jing Wu 0024

dblp:88/3604-24 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-2555-0220ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Sonnet: A Workflow-Aware Serverless Platform for Time-Sensitive Edge Computing With WebAssembly
abstract
The serverless computing paradigm has emerged as a promising solution to address the resource underutilization and inflexible service scaling in edge environments by decoupling the monolithic application into a serverless workflow. However, existing serverless platforms are primarily designed for cloud centers, relying on heavyweight isolation mechanisms that are illsuited for resource-constrained edge computing. These limitations result in high latency, low deployment density, and restricted parallelism. In this paper, we proposeSonnet, a serverless platform tailored for edge computing, capable of rapidly responding to user requests and supporting efficient and elastic service scaling. Sonnet offers these features by (i) employing lightweight WebAssembly as the execution environment for functions, (ii) leveraging serverless workflow information to optimize function deployment on resource-constrained edge environments, and (iii) designing a function deployment algorithm that achieves dynamic load balancing within the cluster. An extensive evaluation ofSonnetwith real-world serverless workflows demonstrates its effectiveness and practical applicability. Compared with SOTA and commonly used edge computing serverless solutions, our experiments show that Sonnet can reduce end-to-end latency by 27% and improve throughput by 2.83×.
Quanfeng Deng, Jing Wu 0024, Qiangyu Pei, Chuangxun Lin, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers2
2025 It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
abstract
Serverless platforms typically adopt an earlybinding approach for function sizing, requiring developers to specify an immutable size for each function within a workflow beforehand. Accounting for potential runtime variability, developers must size functions for worst-case scenarios to ensure service-level objectives (SLOs), resulting in significant resource inefficiency. To address this issue, we propose Janus, a novel resource adaptation framework for serverless platforms. Janus employs a late-binding approach, allowing function sizes to be dynamically adapted based on runtime conditions. The main challenge lies in the information barrier between the developer and the provider: developers lack access to runtime information, while providers lack domain knowledge about the workflow. To bridge this gap, Janus allows developers to provide hints containing rules and options for resource adaptation. Providers then follow these hints to dynamically adjust resource allocation at runtime based on real-time function execution information, ensuring compliance with SLOs. We implement Janus and conduct extensive experiments with real-world serverless workflows. Our results demonstrate that Janus enhances resource efficiency by up to 34.7% compared to the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Quanfeng Deng, Chen Yu 0003, Bingheng Yan, Fangming Liu
IPDPS1
2024 Graft: Efficient Inference Serving for Hybrid Deep Learning With SLO Guarantees via DNN Re-Alignment
abstract
Deep neural networks (DNNs) have been widely adopted for various mobile inference tasks, yet their ever-increasing computational demands are hindering their deployment on resource-constrained mobile devices. Hybrid deep learning partitions a DNN into two parts and deploys them across the mobile device and a server, aiming to reduce inference latency or prolong battery life of mobile devices. However, such partitioning produces (non-uniform) DNN fragments which are hard to serve efficiently on the server. This article presents Graft—an efficient inference serving system for hybrid deep learning with latency service-level objective (SLO) guarantees. Our main insight is to mitigate the non-uniformity by a core concept called DNN re-alignment, allowing multiple heterogeneous DNN fragments to be restructured to share layers. To fully exploit the potential of DNN re-alignment, Graft employs fine-grained GPU resource sharing. Based on that, we propose efficient algorithms for merging, grouping, and re-aligning DNN fragments to maximize request batching opportunities, minimizing resource consumption while guaranteeing the inference latency SLO. We implement a Graft prototype and perform extensive experiments with five types of widely used DNNs and real-world network traces. Our results show that Graft improves resource efficiency by up to 70% compared with the state-of-the-art inference serving systems.
Jing Wu 0024, Lin Wang 0015, Qirui Jin, Fangming Liu
IEEE Trans. Parallel Distributed Syst.1
2022 HiTDL: High-Throughput Deep Learning Inference at the Hybrid Mobile Edge
abstract
Deep neural networks (DNNs) have become a critical component for inference in modern mobile applications, but the efficient provisioning of DNNs is non-trivial. Existing mobile- and server-based approaches compromise either the inference accuracy or latency. Instead, a hybrid approach can reap the benefits of the two by splitting the DNN at an appropriate layer and running the two parts separately on the mobile and the server respectively. Nevertheless, the DNN throughput in the hybrid approach has not been carefully examined, which is particularly important for edge servers where limited compute resources are shared among multiple DNNs. This article presents HiTDL, a runtime framework for managing multiple DNNs provisioned following the hybrid approach at the edge. HiTDL's mission is to improve edge resource efficiency by optimizing the combined throughput of all co-located DNNs, while still guaranteeing their SLAs. To this end, HiTDL first builds comprehensive performance models for DNN inference latency and throughout with respect to multiple factors including resource availability, DNN partition plan, and cross-DNN interference. HiTDL then uses these models to generate a set of candidate partition plans with SLA guarantees for each DNN. Finally, HiTDL makes global throughput-optimal resource allocation decisions by selecting partition plans from the candidate set for each DNN via solving a fairness-aware multiple-choice knapsack problem. Experimental results based on a prototype implementation show that HiTDL improves the overall throughput of the edge by$4.3\times$compared with the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Qiangyu Pei, Xingqi Cui, Fangming Liu, Tingting Yang 0001
IEEE Trans. Parallel Distributed Syst.1