VLDB 2026 Research / reviewers in the wild / expert
Desheng Wang 0002
dblp:19/1084-2
· DBLP profile ↗
25ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-7502-7094ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 16 since 2021Computer networks · 6 · 2 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Contribution-Aware Coalition Federated Learning in Edge-Assisted Healthcare Monitoring Systems
Hualong Wu, Weizhe Zhang, Desheng Wang 0002 |
ICC | 4 |
| 2026 | Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing EnvironmentsabstractGPU sharing is commonly employed in GPU clusters to improve utilization, with spatial sharing being one of the most widely adopted techniques. However, spatial sharing can lead to resource interference, making task execution times difficult to predict. Predictable execution times for each task are crucial in GPU cluster management and task scheduling. In this article, we propose a performance predictor for multi-DNN training tasks in GPU spatial sharing environments. We first conduct experiments on spatial sharing for multiple DNN workloads on a single GPU, demonstrating that concurrent execution of multiple tasks improves overall performance and GPU resource utilization compared to serial execution. By analyzing warp stall reasons collected during task execution, we investigate the interference for computation and memory resources under MPS on GPUs. Finally, we design a performance predictor that predicts the execution time of a target DNN training task when it runs concurrently with other tasks under GPU spatial sharing via MPS. The predictor is capable of predicting the execution time of each task for previously unseen combinations of DNN training tasks. Extensive evaluations on modern GPUs show that compared to other baseline methods, our approach exhibits higher prediction accuracy, as well as improved stability and robustness. Experiments on multiple GPU architectures, as well as at higher concurrency levels, further demonstrate that our method possesses strong generalization and scalability. We also conducted a performance analysis under diverse workload pattern and a case study to validate the practical applicability of our predictor in real scheduling environments. Sichao Chen, Desheng Wang 0002, Weizhe Zhang, Meng Hao 0002, Yu-Chu Tian |
ACM Trans. Archit. Code Optim. | 2 |
| 2026 | PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power ConstraintsabstractThe proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents PctoDL , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM/memory frequency settings. To address the throughput–power tradeoff in power-constrained multi-tenant inference, PctoDL couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform, PctoDL improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak’s coarse-grained partitioning approach, PctoDL achieves average/peak gains of 79.05%/137.93% on the RTX 3080 Ti and 26.33%/70.21% on the A100. Meng Hao 0002, Zikun Wu, Xueyang Tian, Siyu Yang 0002, Guotong Guo, Yiming Wang 0010, Farui Wang, Desheng Wang 0002, Weizhe Zhang |
ACM Trans. Archit. Code Optim. | 10 |
| 2026 | GreenDLS: An Energy-Efficient and SLO-Aware Deep Learning Serving SystemabstractThe growing demand for deploying deep learning (DL) models, particularly large language models (LLMs), has made it imperative to optimize GPU energy consumption while meeting service-level objectives (SLOs). Significant energy use and carbon dioxide (CO2) emissions from GPU-based inference tasks contribute substantially to the environmental footprint of the DL deployment. Existing approaches primarily rely on batching and dynamic voltage and frequency scaling (DVFS) to optimize service performance or throughput, but often overlook memory frequency adjustments and holistic energy optimization under dynamic workloads. This study presents GREENDLS, a DL serving system that integrates deep reinforcement learning (DRL) with offline prediction models to optimize energy consumption while adhering to inference latency SLOs, achieving significant energy savings. GREENDLS dynamically adjusts batch size, GPU streaming multiprocessor (SM) frequency, and GPU memory frequency based on inference request rates. It also accounts for GPU energy consumption during idle phases, such as batch filling, enabling multi-parameter and fine-grained energy optimization. Compared to the Clipper system, GREENDLS achieves energy savings of up to 45.57% on the RTX 3080Ti and 39.44% on the Tesla V100S. When compared to the EAIS system, which only combines batching with GPU SM frequency adjustment, GREENDLS achieves energy savings of up to 36.44% on the RTX 3080Ti and 15.44% on the Tesla V100S. Against the method proposed by Yu et al., GREENDLS attains energy savings of up to 46.34% on the RTX 3080Ti and 33.74% on the Tesla V100S. In LLM inference tasks using Qwen, GREENDLS reduces average energy consumption by 40.79% compared to Clipper, 10.42% compared to EAIS, and 42.42% compared to Yu et al. These results clearly demonstrate that GREENDLS more effectively optimizes energy consumption compared to traditional methods that rely primarily on batching or a combination of batching and DVFS, while still ensuring SLO compliance. Meng Hao 0002, Xueyang Tian, Siyu Yang 0002, Yiming Wang 0010, Desheng Wang 0002, Weizhe Zhang |
IEEE Trans. Computers | 9 |
| 2026 | Accelerating Secure Machine Learning Training on GPUs With Pipeline ParallelismabstractThe proliferation of data-driven machine learning (ML) applications makes privacy and data security risks increasingly prominent. Secure multi-party computation (MPC) offers a privacy-preserving ML method by enabling joint model training without data disclosure, but its reliance on complex cryptography adds significant delay and resource demands, challenging its practicality in large-scale applications. To address the low GPU utilization within the MPC framework, we propose an innovative privacy-preserving ML training optimization framework that exploits pipeline parallelism. Through a detailed analysis of the underlying principles of MPC-based training, we identify computation and communication as the primary bottlenecks for linear and non-linear computations, respectively. Drawing from traditional ML optimization strategies, we design a sub-network partitioning and pipeline parallelism method, specifically tailored for MPC training. This method not only allows for simultaneous training computations across different network layers, but also strategically overlaps computation with communication to enhance GPU utilization and reduce latency. Additionally, we develop a distributed communication mechanism to further improve communication efficiency. We integrate our framework into two distinct, state-of-the-art secure training frameworks: CryptGPU and Piranha. Compared to their original versions, our enhancements boost training speeds by up to 51%, significantly increasing GPU utilization, with negligible impact on model convergence speed and accuracy. Meng Hao 0002, Mingdong Xie, Weizhe Zhang, Linxuan Wang, Xueyang Tian, Desheng Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2025 | ServerlessLego: An Elastic Serverless Framework Assembling Model Building Blocks to Provide SLO-Aware Inference ServicesabstractInference of large language models (LLMs) is common in cloud environments. As the elastic resource management capabilities and the flexible pay-as-you-go billing model offered by serverless, LLM inference services are increasingly migrated to serverless platforms. However, the increasing size of LLMs in recent years has introduced a new cold start issue for serverless frameworks, which in turn impacts their scalability under dynamic workloads. To address these issues, we propose ServerlessLego, an elastic serverless computing framework. ServerlessLego partitions LLMs into layers, then groups and deploys them to different instances, and loads these groups in parallel. These instances perform a subscription-based pipeline. To address dynamically request loads, ServerlessLego models the incoming request patterns and the inference time of running requests, providing an SLO-Aware instance scheduling. Experiments show that ServerlessLego reduces the cold start time of serverless frameworks by 58.15 % and improves throughput by 43.39 % compared to the baseline for dynamic workloads. Moreover, ServerlessLego can horizontally schedule instance based on request SLOs and arrival rates. Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Yuming Feng 0002 |
ICPADS | 2 |
| 2025 | Aegis Sketch: High-Throughput and Accurate Top-$k$ Elephant Flows Detection in Large-Scale Parallel Network Traffic ProcessingabstractIn large-scale parallel network traffic processing systems, detecting Top-$k$elephant flows is essential for real-time monitoring and traffic management. However, under the dual constraints of high update rates and limited memory, existing approaches struggle to balance throughput and accuracy. Many fail to exploit the heavy-tailed distribution of network traffic, resulting in frequent hash collisions and irreversible eviction errors that severely limit detection performance. Current methods fall into two categories: counter-based approaches, which maintain a candidate set (e.g., a min-heap) for high accuracy but suffer from high synchronization overhead and unrecoverable evictions, and sketch-based approaches, which are naturally parallelizable but prone to accuracy loss under hash collisions. To address these challenges, we propose Aegis Sketch, a novel framework for high-throughput and accurate Top-$k$elephant flow detection in large-scale parallel network traffic processing. Aegis Sketch incorporates an ordered storage mechanism that fundamentally reduces hash collisions without incurring additional structural overhead, thereby significantly improving throughput and accuracy. In addition, a multi-party competitive replacement policy prioritizes the preservation of true elephant flows during contention, effectively mitigating the accuracy loss caused by irreversible evictions. Experiments on real-world traffic traces show that Aegis Sketch outperforms existing methods such as Elastic Sketch and OneSketch in throughput, detection accuracy, and memory efficiency, achieving up to 1.35 times higher throughput and 22 % higher precision. These results demonstrate its effectiveness and efficiency for large-scale parallel network traffic processing. Desheng Wang 0002, Weizhe Zhang |
ICPADS | 3 |
| 2025 | HyDLR: Load-Aware Dynamic Rescheduling for Deep Learning Hybrid DeploymentabstractResource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption. Desheng Wang 0002, Shuo Si, Sichao Chen, Weizhe Zhang |
ICPADS | 1 |
| 2025 | DynGPU: A Dynamic GPU Sharing Framework for Enhanced Resource Utilization and Task Scheduling in Concurrent DNN TrainingabstractTraining deep neural networks (DNNs) is a common task in GPU clusters. However, in practical cluster environments, multiple concurrent DNN training tasks often fail to fully leverage GPU resources, resulting in suboptimal GPU utilization. Furthermore, existing GPU sharing frameworks primarily rely on static scheduling and frequently overlook task deadlines, leading to task delays and inefficient scheduling. To address these issues, we propose a dynamic GPU sharing framework (DynGPU) that intercepts GPU kernel executions to perform resource scheduling in multi-task environments. DynGPU incorporates a dynamic task priority adjustment mechanism that adapts task priorities in real time based on task progress, historical data, and remaining time to deadlines. By guaranteeing resources for high-priority tasks while maximizing resource allocation for low-priority tasks, DynGPU reduces resource contention and improves system throughput, enabling more timely task completions. Experiments show that, compared to dedicated GPU execution, DynGPU can reserve up to 97.5 % of throughput for high-priority tasks. Compared to state-of-the-art baselines, DynGPU achieves up to an 8.4 % improvement in task completion time. Zhiji Yu, Desheng Wang 0002, Weizhe Zhang, Sichao Chen, Meng Hao 0002, Yu-Chu Tian |
ICPADS | 2 |
| 2025 | An Efficient Collaborative Algorithm for Building-Wide Mobile Edge ComputingabstractThe rapid proliferation of multi-mobile devices in smart building environments has intensified the demand for efficient computation offloading strategies in mobile edge computing. To address the challenges of task offloading, communication, and computation resource allocation while also considering the mobility of mobile devices and task priorities, we design a target server query strategy based on resource matching. This strategy accommodates the varying resource requirements of different task types and avoids increasing algorithm complexity. Based on this strategy, we propose the Greedy-Based Collaborative Algorithm to minimize the average execution time of tasks. Simulation results demonstrate that the proposed algorithm outperforms baseline algorithms and that the energy consumption of mobile devices remains acceptable. Hualong Wu, Weizhe Zhang, Desheng Wang 0002 |
INDIN | 4 |
| 2025 | EVRM: Elastic Virtual Resource Management framework for cloud virtual instances
Desheng Wang 0002, Weizhe Zhang, Zhiji Yu, Yu-Chu Tian, Keqin Li 0001 |
Future Gener. Comput. Syst. | 1 |
| 2025 | HEngine: A High Performance Optimization Framework on a GPU for Homomorphic EncryptionabstractHomomorphic encryption (HE) represents an encryption technology that allows for direct computation on encrypted data without requiring decryption. However, the substantial computational complexity and significant latency associated with HE has impeded its broader adoption in practical applications. To address these challenges, we propose a GPU-based acceleration framework, namely HEngine, tailored for homomorphic encryption tasks. Specifically, we first propose a warp shuffle-based optimization method for two key phases, i.e., inverse Chinese Remainder Theorem (ICRT) and number theoretic transformation (NTT), to mitigate synchronization overhead in homomorphic encryption. Secondly, we propose to fuse the NTT kernel with the inner product kernel to address the imbalance between memory access and computation. Thirdly, considering the potential difference in the amount of tasks of users in the real-world, we design two different encoding methods for small batch and large batch inference tasks to improve computational efficiency. Finally, experiments demonstrate that our proposed framework achieves a 218× speedup on homomorphic multiplication tasks compared with the CPU-based SEAL library. In addition, for convolutional neural network inference tasks on shallow network structures, our proposed framework achieves amortized inference performance at the millisecond level and sub-millisecond level on small batch and large batch data, respectively. For convolutional neural network inference tasks on deeper network structures (i.e., ResNet-20), our proposed framework achieves second-level inference. Meng Hao 0002, Weizhe Zhang, Desheng Wang 0002 |
ACM Trans. Archit. Code Optim. | 6 |
| 2024 | Fast Memory Disaggregation with SwiftSwap
Xiangwei Zhang, Desheng Wang 0002, Weizhe Zhang, Zhiji Yu, Meng Hao 0002 |
NPC (1) | 2 |
| 2024 | Optimizing depthwise separable convolution on DCUabstractAbstract The integration of Large Language Models (LLMs) with Convolutional Neural Networks (CNNs) is significantly advancing the development of large models. However, the computational cost of large models is high, necessitating optimization for greater efficiency. One effective way to optimize the CNN is the use of depthwise separable convolution (DSC), which decouples spatial and channel convolutions to reduce the number of parameters and enhance efficiency. In this study, we focus on porting and optimizing DSC kernel functions from the GPU to the Deep Computing Unit (DCU), a computing accelerator developed in China. For depthwise convolution, we implement a row data reuse algorithm to minimize redundant data loading and memory access overhead. For pointwise convolution, we extend our dynamic tiling strategy to improve hardware utilization by balancing resource allocation among blocks and threads, and we enhance arithmetic intensity through a channel distribution algorithm. We implement depthwise and pointwise convolution kernel functions and integrate them into PyTorch as extension modules. Experiments demonstrate that our optimized kernel functions outperform the MIOpen library on the DCU, achieving up to a 3.59 $$\times$$ × speedup in depthwise convolution and up to a 3.54 $$\times$$ × speedup in pointwise convolution. These results highlight the effectiveness of our approach in leveraging the DCU’s architecture to accelerate deep learning operations. Meng Hao 0002, Weizhe Zhang, Gangzhao Lu, Xueyang Tian, Siyu Yang 0002, Mingdong Xie, Chenyu Yuan, Desheng Wang 0002 |
CCF Trans. High Perform. Comput. | 10 |
| 2024 | Adaptive asynchronous federated learning
Renhao Lu, Weizhe Zhang, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002, Zenglin Xu, Mamoun Alazab |
Future Gener. Comput. Syst. | 7 |
| 2024 | Cooperative Service Caching in Vehicular Edge Computing Networks Based on Transportation Correlation AnalysisabstractIn vehicular edge computing, the vehicular services are cached and replaced among Roadside Units (RSU) to minimize delay in delay-sensitive services. However, without prior knowledge of these service preferences, the cached services suffer a low hit ratio and a high network delay. Before the appearance of requests from vehicles, most existing vehicular service caching methods design the service-sharing mechanisms among RSUs to optimize the hit ratio, but they ignore the influence of vehicle mobility on service-sharing. Moreover, in vehicular services replacement, the differences between actual and historical vehicle trajectories negatively impact the service-sharing mechanisms constructed before, which existing methods ignored. Thus, this paper addresses the vehicular service caching problem based on the surrounding function-features of vehicles and transportation correlations. Firstly, we use the surrounding function-features of vehicles to estimate the service preference and ensure the hit ratio of cached services. Secondly, we formulate the vehicular service caching problem as a constrained optimization problem and design a cooperation mechanism between RSUs, which is based on transportation correlations of vehicle trajectories. Thirdly, we propose two vehicular service caching methods based on Gibbs sampling to optimize the network delay of vehicular services and deal with the negative influence of vehicle mobility. Simulation in real datasets from Shenzhen, China shows that our methods have a better network delay and hit ratio than existing baselines by 15.26% and 5.6%, respectively. Chen Ling 0006, Weizhe Zhang, Qingyang Fan, Zebang Feng, Desheng Wang 0002 |
IEEE Internet Things J. | 7 |
| 2024 | Two-Stage Client Selection for Federated Learning Against Free-Riding Attack: A Multiarmed Bandits and Auction-Based ApproachabstractUtilizing the federated learning (FL) technique, data owners can collaboratively train artificial intelligence models, retaining all training data on their premises to minimize the potential for personal data breaches. However, self-interested users (e.g., free riders) bring new challenges that hinder the development of FL techniques. To this end, we propose a two-stage client selection scheme comprising a multiarmed bandit (MAB)-based candidate client selection method and an auction-based training client selection method. Specifically, our client selection scheme initially formulates the FL system into an MAB system, where clients are the arms and the server is the player. Then, we quantify the similarity between a local model and the server side, which is the designed metric for model aggregation and reward computation updating based on the fuzzy mathematical strategy. Next, based on the Thompson Sampling strategy, the server can intelligently determine the reward of each client, and clients with more significant rewards have the chance for local model training. With an auction method, the server can determine the training clients to reduce the training cost while maximizing each client’s revenue. Extensive experiments on real-world data sets demonstrate that the proposed scheme outperforms representative FL schemes (i.e., FedAvg, FedProx, FedMax, and MFL) regarding the model’s convergence rate and cost in FL systems with free riders. Renhao Lu, Weizhe Zhang, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002, Lu Shi 0002, Yuelin Guo |
IEEE Internet Things J. | 7 |
| 2024 | Minimizing Service Latency Through Image-Based Microservice Caching and Randomized Request Routing in Mobile Edge ComputingabstractIn the context of mobile edge computing (MEC), the traditional method of requesting microservices from a central cloud can result in increased delay for users due to the physical distance between the user and the cloud server. To address this issue, MEC advocates for placing servers closer to the users at the edge of the network. However, this approach is constrained by the storage capacity and computing resources of edge servers (ESs). Therefore, it is crucial to devise a strategy for processing user requests that minimizes the average request delay. To address this problem, This article formulates the microservice caching problem as an image-based microservice placement and task request routing problem. We model the problem as an integer linear programming problem with multicondition constraints. Considering the limited resources of ESs, we propose a microservice placement algorithm called approximate algorithm based on randomized task request routing. The proposed algorithm is designed to provide near-optimal solutions in polynomial time, leveraging Chernoff’s theorem. Our approach is evaluated through comparisons with two existing algorithms: 1) the image-pull-based microservice cache request algorithm and 2) the greedy-based microservice cache and request routing algorithm. The results demonstrate that our algorithm exhibits superior performance compared to existing methods. Desheng Wang 0002, Weizhe Zhang, Guanqing Lou |
IEEE Internet Things J. | 2 |
| 2024 | An IoT Device Identification Method Using Extracted Fingerprint From Sequence of Traffic Grayscale ImagesabstractWith the widespread deployment and application of various types of IoT devices, preventing illegal intrusion and impersonation attacks of IoT devices has become an important security challenge. Device identification helps to limit the behavior of suspicious devices and enhances the security of the device access process. In this paper, we propose a novel deep learning-based automatic fingerprint extraction model that addresses low efficiency and complexity of traditional feature engineering process, which are often rely on expert experience. The proposed model integrates advanced modules such as Depthwise Separable Convolution (DSC) and Gated Recurrent Unit (GRU), as well as architectures of inverted residuals and linear bottlenecks to enhance the performance of fingerprint extraction. After converting the raw device traffic into the sequence of traffic grayscale images, the model can analyze spatial and temporal features from them to generate highly distinguishable device fingerprints automatically. Additionally, we also achieve fast fingerprint search based on Hierarchical Navigable Small World (HNSW) to support device identification. Our proposed method can not only indicate deviations in device behavior from expected specifications, but also identify unknown and unreliable IoT devices. The experimental results show that our method has excellent performance and more comprehensive identification capabilities in multiple dimensions. Yuming Feng 0002, Yu Zhang 0036, Weizhe Zhang, Desheng Wang 0002 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | A Multi-Objective Virtual Network Migration Algorithm Based on Reinforcement LearningabstractVirtual network migration (VNM) helps improve network performance by remapping a subset of virtual nodes or links to physical infrastructure, aligning the resource allocation to the virtual network's changing conditions. However, existing VNM methods neglect integrating multiple objectives that affect network performance, such as energy, communication, migration, and service level agreement violation (SLAV). It is challenging to make VNM decisions to optimize the overall objective in a large-scale cloud environment. This article establishes a multi-objective optimization model and proposes a multi-objective VNM algorithm called MiOvnm. The MiOvnm employs the double deep$Q$-learning approach to cope with ample state space. It also applies an action selection method called actfilter to deal with large-scale action space. The MiOvnm finds the migration action with optimal potential reward from the candidate action set. Simulation results demonstrate the superiority of our MiOvnm to the state-of-the-art methods. More specifically, MiOvnm reduces average SLAV, communication cost, and total cost by 24.32%, 4.95%, and 12.45%, respectively. Furthermore, evaluation results in a real-world OpenStack platform reveal that making full use of computation and network resources, the MiOvnm reduces the completion time of computation- and network-intensive benchmarks by 11.35% and 10.31%, respectively, with a total cost reduction of 26.02%. Desheng Wang 0002, Weizhe Zhang, Junren Lin, Yu-Chu Tian |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Auction-Based Cluster Federated Learning in Mobile Edge Computing SystemsabstractFederated Learning (FL), allowing data owners to conduct model training without sending their raw data to third-party servers, can enhance data privacy in Mobile Edge Computing (MEC) which brings data processing closer to the data sources. However, the heterogeneity of local data and constrained local resources in MEC bring new challenges hindering the development of FL. To this end, we propose an Auction-based Cluster Federated Learning scheme, called ACFL, comprising a clustered FL framework and an auction-based client selection strategy. Our clustered FL framework first introduces a mean-shift clustering algorithm to FL, which can intelligently cluster clients according to their local data distribution. Then, we select clients from each cluster using an auction mechanism to participate in FL training, which can mitigate the impact of data heterogeneity on model convergence and balance energy consumption. Moreover, we prove the proposed clustered FL framework converges at a sublinear rate. Extensive experiments conducted on real-world datasets demonstrate that the proposed FL scheme outperforms the conventional FL schemes in terms of convergence rate and energy balance. Renhao Lu, Weizhe Zhang, Yan Wang 0002, Qiong Li 0001, Xiaoxiong Zhong, Desheng Wang 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | Repeatable Multi-Dimensional Virtual Network Embedding in Cloud Service PlatformabstractVirtual network embedding (VNE) can effectively deploy virtual networks (VNs) onto shared substrate network (SN) resources. However, with the consistent changing scalability and diversity demands of VNs, traditional VNE methods prove to be a challenging task for current cloud service platforms. Thus, we model a repeatable multi-dimensional virtual network embedding (RMD-VNE) problem for implementing multi-dimensional virtual networks (MD-VNs) that involves real servers, virtual machines, containers, and network simulators. The MD-VN is preprocessed and embedded via a heuristic method denoted asReMiDvne. Following its transformation for the containers and simulation networks, the MD-VN topology undergoes a process of coarsening, partitioning, and uncoarsening.ReMiDvnethen applies a topology-aware repeatable embedding solution to complete the embedding stage. Experimental results demonstrate thatReMiDvneoutperforms seven baseline approaches through small-, 1,000- and 10,000-scale VNE simulation experiments. Remarkably,ReMiDvneimproves the average rates of acceptance ratio, revenue, and revenue-cost ratio by up to 40.45, 40.45, and 299.03 percent, respectively, and reduces the average rate of cost by up to 64.16 percent. Furthermore, real-world VNE experiments are conducted based on the OpenStack platform. The results reveal the ability ofReMiDvneto efficiently reduce communication costs by up to 45.93 and 63.43 percent for download and upload, respectively. Weizhe Zhang, Desheng Wang 0002, Shui Yu 0001, Yan Wang 0002 |
IEEE Trans. Serv. Comput. | 2 |
| 2021 | Kalman prediction-based virtual network experimental platform for smart living
Desheng Wang 0002, Weizhe Zhang, Yang Xiang 0003, Yu-Chu Tian |
Comput. Commun. | 1 |
| 2021 | Node-Fusion: Topology-aware virtual network embedding algorithm for repeatable virtual network mapping over substrate nodesabstractSummary Cloud computing has become a new Internet application model, where network virtualization is recognized as an important technology for allowing multiple heterogeneous virtual networks (VNs) to coexist on a shared substrate network (SN). As demands in cloud computing increase, the scale of VN greatly increases as well, and providing an end‐to‐end SN to embed VNs in terms of scale is difficult. To utilize SN resources fully, we devise a topology‐aware Node‐Fusion algorithm, which is different from the traditional virtual network embedding (VNE) algorithms, for repeatable VNE over substrate nodes problem. We rank the resource of nodes through a novel solution by considering the CPU and bandwidth of adjacent link capacity and the number of adjacent links of each node as resources, and rank a node on the basis of resources. Furthermore, we embed several virtual nodes into the same substrate node together in accordance with Node‐Fusion interconnection value during the node mapping process, which can greatly improve the success ratio of the subsequent link mapping phase. Evaluation results confirm that Node‐Fusion outperforms traditional classical heuristics (Link‐opt, Node‐opt, and ORSTA), which are modified to fit into our model, with regard to acceptance ratio, long‐term revenue, long‐term cost, and revenue‐cost ratio. Desheng Wang 0002, Weizhe Zhang, Chuanyi Liu |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | RLS-VNE: Repeatable Large-Scale Virtual Network Embedding over Substrate NodesabstractEmbedding multiple virtual networks (VNs) on a shared substrate network (SN), known as virtual network embedding (VNE), is a challenging problem in cloud platforms. VNE methods can provide strategies to deploy VNs onto SN resources. However, as the scale of VN greatly increases, traditional VNE methods are time-consuming and waste link resource. Meanwhile, traditional VNE methods assign each virtual node of the same VN to different substrate nodes, whereas it is hard to provide larger scale SN to provision the VN. In order to efficiently embed large-scale VNs, multiple virtual nodes from the same VN need to share the same substrate node. We therefore model a repeatable large-scale virtual network embedding (RLSVNE) problem in this study, provisioning large-scale VNs, and propose a heuristic method (Rlsvne) to handle RLS-VNE. Rlsvne pre-processes the VN topology before embedding. In the pre-processing stage, the VN topology is processed through graph coarsening, partitioning, and uncoarsening. After the pre-processing, Rlsvne accomplishes an embedding stage with a topology- aware repeatable embedding solution. 1,000 and 10,000-scale VNE experiments are conducted to demonstrate our Rlsvne. The evaluation results demonstrate that our Rlsvne outperforms three modified heuristics. Rlsvne shows improved performance in reducing substrate cost and fully utilizing substrate resources, achieving high acceptance ratio and revenue values. Desheng Wang 0002, Weizhe Zhang, Shui Yu 0001 |
GLOBECOM | 1 |