Wenxiu Xu

dblp:205/7391 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2023
0009-0008-5982-2953ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 DNN Inference Task Offloading Based on Distributed Soft Actor-Critic in Mobile Edge Computing
abstract
In mobile edge computing, DNN-driven intelligent inference service is highly sensitive to latency.Recently, collaborative inference between user devices and Edge Servers (ESs) based on DNN partition has been used in service acceleration.However, due to the limited computing resources of ESs, there is resource competition between concurrent requests, resulting in the partition tasks cannot be offloaded to ESs in time.Therefore, it is necessary to design an efficient offloading scheme for partitionbased concurrent inference tasks.Existing task offloading schemes based on Deep Reinforcement Learning (DRL) can solve complex decision-making problems in high-dimensional state space, but there are problems such as insufficient sample diversity and easily falling into local optimum.Therefore, we propose a collaborative DNN inference task offloading scheme based on distributed Soft Actor-Critic(SAC).It supports SAC Agents to explore samples in parallel and share learning experiences, and improves the randomness of the policy through the maximum entropy mechanism to avoid falling into local optimum, thus achieving efficient offloading of concurrent partition tasks.Experimental results on DNN benchmarks show that compared with the baseline schemes, the average service latency of our scheme is reduced by more than 18.3%, and it has a higher convergence speed and task success rate, which can make ESs achieve load balancing.
Wenxiu Xu, Ningjiang Chen, Huan Tu
SEKE1
2023 Collaborative Inference Acceleration Integrating DNN Partitioning and Task Offloading in Mobile Edge Computing
abstract
In mobile edge computing environment, intelligent inference services driven by DNN are highly sensitive to latency. Recently, collaborative inference between User Devices and Edge Servers (ESs) based on Deep Neural Networks (DNN) partition has achieved success in service acceleration. However, most of the existing collaborative acceleration schemes are partitioned for a single DNN inference task, which cannot quickly make partition decisions for a set of concurrent inference tasks, and often sacrifice inference accuracy. In addition, due to the limited resources of ESs, there is resource competition among concurrent requests, which makes the partitioned tasks cannot be offloaded to ESs in time for processing. Therefore, designing an efficient offloading scheme becomes essential. The task offloading schemes based on deep reinforcement learning can solve complex decision-making problems in high-dimensional state space, but they have problems such as insufficient sample diversity and easily falling into local optimum. In this paper, a Collaborative Inference Acceleration Scheme integrating DNN Partitioning and Task Offloading (CIAS-PnO) is proposed. First, while ensuring inference accuracy, the Collaborative DNN Layer Partitioning (CDLP) algorithm is designed with the goal of optimal latency. CDLP can reduce the problem scale of concurrent inference tasks partition by pruning operation and determine the partition decisions in time. Then, the Distributed Soft Actor-Critic (SAC)-based Partition Task Offloading algorithm (DSACO) is designed. DSACO supports SAC Agents to explore samples in parallel and share learning experiences, and uses the automatic entropy adjustment mechanism to improve the exploration efficiency of Agents, so as to avoid falling into local optimum and achieve efficient offloading of partition tasks. Experimental results on DNN benchmarks show that compared with the baseline acceleration schemes, CIAS-PnO achieves more than 19.8% acceleration performance improvement, and has higher convergence performance and task success rate.
Wenxiu Xu, Yin Yin, Ningjiang Chen, Huan Tu
Int. J. Softw. Eng. Knowl. Eng.1