Wenzhuo Qian

dblp:368/6065 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PeerSync: Accelerating Containerized Model Inference at the Network Edge
abstract
Efficient container image distribution is crucial for enabling machine learning inference at the network edge, where resource limitations and dynamic network conditions create significant challenges. In this paper, we presentPeerSync, a decentralized P2P-based system designed to optimize image distribution in edge environments.PeerSyncemploys a popularity- and network-aware download engine that dynamically adapts to content popularity and real-time network conditions.PeerSyncfurther integrates automated tracker election for rapid peer discovery and dynamic cache management for efficient storage utilization. We implementPeerSyncwith 8000+ lines of Rust code and test its performance extensively on both large-scale Docker-based emulations and physical edge devices. Experimental results show thatPeerSyncdelivers a remarkable speed increase of 2.72×, 1.79×, and 1.28× compared to the Baseline solution, Dragonfly, and Kraken, respectively, while significantly reducing cross-network traffic by 90.72% under congested and varying network conditions.
Yinuo Deng, Hailiang Zhao, Dongjing Wang, Peng Chen 0051, Wenzhuo Qian, Jianwei Yin, Schahram Dustdar, Shuiguang Deng
IEEE Trans. Serv. Comput.5
2025 Decentralized Proactive Model Offloading and Resource Allocation for Split and Federated Learning
abstract
In the resource-constrained Internet of Things (IoT)-edge computing environment, split federated (SplitFed) learning is implemented to enhance training efficiency. This method involves each terminal device dividing its full deep neural network (DNN) model at a designated layer into a device-side model and a server-side model, then offloading the latter to the edge server. However, existing research overlooks four critical issues as follows: 1) the heterogeneity of end devices’ resource capacities and the sizes of their local data samples impact training efficiency; 2) the influence of the edge server’s computation and network resource allocation on training efficiency; 3) the data leakage risk associated with the offloaded server-side submodel; and 4) the privacy drawbacks of current centralized algorithms. Consequently, proactively identifying the optimal cut layer and server resource requirements for each end device to minimize training latency while adhering to data leakage risk rate constraint remains a challenging issue. To address these problems, this article first formulates the latency and data leakage risk of training DNN models using SplitFed learning. Next, we frame the SplitFed learning problem as a mixed-integer nonlinear programming challenge. To tackle this, we propose a decentralized proactive model offloading and resource allocation (DP-MORA) scheme, empowering each end device to determine its cut layer and resource requirements based on its local multidimensional training configuration, without knowledge of other devices’ configurations. Extensive experiments on two real-world datasets demonstrate that the DP-MORA scheme effectively reduces DNN model training latency, enhances training efficiency, and complies with data leakage risk constraints compared to several baseline algorithms across various experimental settings.
Binbin Huang 0006, Hailiang Zhao, Lingbin Wang, Wenzhuo Qian, Yuyu Yin, Shuiguang Deng
IEEE Internet Things J.4
2025 Online Workload Scheduling for Social Welfare Maximization in the Computing Continuum
abstract
Computing ecosystems are shifting toward a computing continuum paradigm designed to handle the diverse and dynamic nature of computing resources spread across various locations. It demonstrates significant potential in providing high-bandwidth and low-latency services for users. However, as a large number of users request services from distributed computing continuum systems, it is critical to schedule numerous delay-sensitive, fractional workloads and maximum parallelism-bound jobs to appropriate backend resources,e.g., cloud container instances. In addition, the scheduling strategy also needs to maximize the social welfare that incorporates the utilities of jobs and the revenue of service providers. However, current workload scheduling algorithms are based on simple heuristics and lack performance guarantees. Due to the unpredictability of online requests, the distribution of requests should not be assumed. Therefore, designing an online workload scheduling strategy without assumptions on request distributions is essential for balancing the online workload. This work first establishes a spatiotemporal integrated resource pool to reflect the computational resources provided by distributed computing continuum systems. Then, several pseudo-social welfare functions and marginal cost functions are constructed, where the latter is used to estimate the marginal cost of provisioning services to each newly arrived job based on the current resource surplus. We propose an online workload scheduling strategy namedOnSocMaxto solve the above problems. It operates by following the solutions to several convex pseudo-social welfare maximization problems and is proven to be$\alpha$-competitive for some$\alpha$with a value of at least 2. The evaluation results demonstrate thatOnSocMaxoutperforms several benchmark strategies in maximizing social welfare.
Hailiang Zhao, Ziqi Wang 0011, Guanjie Cheng, Wenzhuo Qian, Peng Chen 0051, Jianwei Yin, Schahram Dustdar, Shuiguang Deng
IEEE Trans. Serv. Comput.4