Fei Xu 0009

dblp:42/5065-9 · DBLP profile ↗
← Back
42ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0003-1590-5323ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 6 first-author · 15 since 2021Computer networks · 12 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Duba: Cost-Efficient Serverless Cloud-Edge Collaborative Machine Learning Serving with Dual-Batching
Jianxiong Liao, Zhi Zhou 0006, Fei Xu 0009
J. Comput. Sci. Technol.4
2026 Cooling as You Wish: Component-Level Cooling for Heterogeneous Edge Datacenters
abstract
As computing shifts toward the edge, edge datacenters are becoming essential for supporting diverse real-time applications. Unlike traditional cloud datacenters, edge datacenters face unique cooling challenges due to their requirements forproximity to end users, high density, and hardware heterogeneity. While warm water cooling is a promising technique for this infrastructure, current one-size-fits-all cooling strategies significantly compromise efficiency due to severe inter- and intra-component hotspots. In this work, we present CoolEdge+, a cost-effective component–level water cooling system for enhancing the cooling efficiency of edge datacenters. Specifically, CoolEdge+dynamically adjusts the inlet water temperature for each component through a carefully designed water circulation architecture to mitigate inter-component hotspots. To address intra-component hotspots, it employs vapor chamber–based cold plates that rapidly dissipate heat without manual intervention or additional energy consumption. We further design a fine-grained cooling control framework that leverages a well-managed power capping approach to decide on customized inlet water temperatures and hardware power limits. Based on a hardware prototype and a real-world trace from Alibaba PAI, evaluation results show that CoolEdge+reduces cooling energy consumption by up to 27.19% compared to existing coarse-grained systems, while maintaining performance guarantees. Compared to the state-of-the-art CoolEdge, CoolEdge+saves 35.24% more cooling costs with comparable energy consumption and no latency violations.
Fangming Liu, Qiangyu Pei, Yongjie Yuan, Qixia Zhang, Ziyang Jia, Fei Xu 0009, Bingheng Yan
IEEE Trans. Computers8
2026 Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters
abstract
Scheduling deep learning (DL) models to train on powerful clusters with accelerators like GPUs and TPUs, presently falls short, either lacking fine-grained heterogeneity awareness or leaving resources substantially under-utilized. To fill this gap, we propose a novel task-level heterogeneity-aware scheduler for DL clusters, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>, based on an optimization framework able to boost cluster resource utilization. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> leverages the performance traits of DL jobs on a heterogeneous DL cluster to make scheduling decisions across both spatial and temporal dimensions. It characterizes the task-level performance heterogeneity for optimization and involves the primal-dual framework employing a dual subroutine, to solve the optimization problem and guide the scheduling design. Our trace-driven simulation with representative DL model training workloads demonstrates that <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> accelerates the total training time duration by 1.20× when compared with its state-of-the-art heterogeneity-aware counterpart, Gavel. Further, our <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> scheduler is enhanced to <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E by forking each job into multiple copies to let a job train concurrently on heterogeneous GPUs resided on separate available cluster nodes (i.e., machines or servers) for resource utilization enhancement. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E is evaluated extensively on physical DL clusters for comparison with <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> and Gavel. With substantial enhancement in cluster resource utilization (by 1.45×), <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E exhibits considerable speed-ups in DL model training, reducing the total training time duration by 50% (or 80%) on an Amazon’s AWS (or our lab) cluster, while producing trained DL models with consistently better inference quality than those trained by <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>.
Abeda Sultana, Nabin Pakka, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
IEEE Trans. Computers3
2026 PatternInsight: An Online Approach to Complex Pattern Detection Over Mobile Data Streams
abstract
Today's mobile applications oftentimes need to detect user-defined complex patterns (e.g., the mysterious “phantom traffic jam”) over data streams to support decision making. It is achieved by continuously creating candidate instances that have partially matched a pattern, and meanwhile aggregating common instances (across patterns) for efficiency enhancement. Existing aggregation approaches are taken in a straightforward or intuitive manner, incurring an exponential solution space and thus having to be executed offline. This paper explores how to significantly accelerate aggregation so as to make pattern detection online executable, even suited to the emerging serverless runtime that involves complicated state synchronizations among distributed cloud functions. By comprehensively investigating a wide variety of mobile data streams, we note the existence of a latent hierarchical cluster structure among complex patterns (in terms of their instance similarities), which can be utilized to quickly aggregate common instances without going through the exponential solution space. To extract the latent information, we devise a content-aware structural entropy minimization algorithm to properly determine intra-cluster patterns, together with a lightweight differential compensation mechanism to maintain those inter-cluster “residual” relations among patterns. Evaluations on real-world vehicle and sensor network data streams illustrate that the resulting approach, dubbed PatternInsight, saves the aggregation time by 10× to 50× and reduces the instance size by 40%.
Yuyang Ren, Zhenhua Li 0001, Fei Xu 0009, Yunhao Liu 0001, Guihai Chen
IEEE Trans. Mob. Comput.4
2025 Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
Qiannan Zhou, Fei Xu 0009, Lingxuan Weng, Ruixing Li, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
INFOCOM2
2025 SEAFL: Enhancing Efficiency in Semi-Asynchronous Federated Learning Through Adaptive Aggregation and Selective Training
abstract
Federated Learning (FL) is a promising distributed machine learning framework that allows collaborative learning of a global model across decentralized devices without uploading their local data. However, in real-world FL scenarios, the conventional synchronous FL mechanism suffers from inefficient training caused by slow-speed devices, commonly known as stragglers, especially in heterogeneous communication environments. Though asynchronous FL effectively tackles the efficiency challenge, it induces substantial system overheads and model degradation. Striking for a balance, semi-asynchronous FL has gained increasing attention, while still suffering from the open challenge of stale models, where newly arrived updates are calculated based on outdated weights that easily hurt the convergence of the global model. In this paper, we present SEAFL, a novel FL framework designed to mitigate both the straggler and the stale model challenges in semi-asynchronous FL. SEAFL dynamically assigns weights to uploaded models during aggregation based on their staleness and importance to the current global model. We theoretically analyze the convergence rate of SEAFL and further enhance the training efficiency with an extended variant that allows partial training on slower devices, enabling them to contribute to global aggregation while reducing excessive waiting times. We evaluate the effectiveness of SEAFL through extensive experiments on three benchmark datasets. The experimental results demonstrate that SEAFL outperforms its closest counterpart by up to$\sim 22 \%$in terms of the wall-clock training time required to achieve target accuracy.
Md Sirajul Islam, Sanjeev Panta, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
IPDPS3
2025 Multi-Width Neural Network-Assisted Hierarchical Federated Learning in Heterogeneous Cloud-Edge-Device Computing
abstract
Federated learning (FL), an emerging data-secure distributed training paradigm, unites massive isolated Internet of Things (IoT) device nodes to collaboratively train a global neural network (NN) model without the exposure of their local multimedia data. However, constrained by the synchronous NN model integration nature of FL, there is a training latency inconsistency among heterogeneous devices, which significantly deteriorates FL training efficiency. Meanwhile, frequent local NN training and transmission impose high energy consumption pressure on users. To tackle these issues, this paper proposes a premium multi-width NN-assisted hierarchical FL (HFL) framework in heterogeneous cloud-edge-device computing to achieve remarkable training speedup and energy conservation. Specifically, a heterogeneity-aware NN width coefficient determination algorithm, which flexibly assigns a subnet with a suitable width to each user device based on its computing ability, is first applied to shorten the HFL training latency. Subsequently, to integrate subnets with different width topologies, we design a width-aware adaptive NN model integration approach to effectively ensure the accuracy of the integrated global NN model. Finally, a latency-aware energy saving strategy is introduced to reduce energy consumption. Experimental results demonstrate that our proposed framework outperforms state-of-the-art benchmarks, and attains up to 42.42% enhancement in accuracy, 81.5% reduction in training latency, and 40.9% optimization in energy cost.
Guobing Zou, Fei Xu 0009, Yangguang Cui, Tongquan Wei
ACM Multimedia3
2025 Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
abstract
GPUs have become thedefactohardware devices for accelerating Deep Neural Network (DNN) inference workloads. However, the conventionalsequential execution mode of DNN operatorsin mainstream deep learning frameworks cannot fully utilize GPU resources, even with the operator fusion enabled, due to the increasing complexity of model structures and a greater diversity of operators. Moreover, theinadequate operator launch orderin parallelized execution scenarios can lead to GPU resource wastage and unexpected performance interference among operators. In this paper, we proposeOpara, a resource- and interference-aware DNNOperatorparallel scheduling framework to accelerate DNN inference on GPUs. Specifically,Oparafirst employsCUDA StreamsandCUDA Graphtoparallelizethe execution of multiple operators automatically. To further expedite DNN inference,Oparaleverages the resource demands of operators to judiciously adjust the operator launch order on GPUs, overlapping the execution of compute-intensive and memory-intensive operators. We implement and open source a prototype ofOparabased on PyTorch in anon-intrusivemanner. Extensive prototype experiments with representative DNN and Transformer-based models demonstrate thatOparaoutperforms the default sequentialCUDA Graphin PyTorch and the state-of-the-art operator parallelism systems by up to$1.68\boldsymbol{\times}$and$1.29\boldsymbol{\times}$, respectively, yet with acceptable runtime overhead.
Aodong Chen, Fei Xu 0009, Li Han 0001, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers2
2024 COUPLE: Orchestrating Video Analytics on Heterogeneous Mobile Processors
abstract
Video analytics is considered the killer application of edge computing and has been successfully deployed across diverse domains. Yet, executing video analytics on mobile devices presents notable challenges owing to the considerable computational demands and frame rate requirements of DNN models. Current mobile inference frameworks often concentrate on enhancing model inference performance on the CPU or GPU, overlooking the potential of the Digital Signal Processor (DSP) – an emerging heterogeneous processor increasingly integrated into modern mobile processors. In this paper, we introduce COUPLE, an orchestration framework for video analytics on heterogeneous mobile processors, with the goal of optimizing real-time video analysis through the collaboration of CPU, GPU and DSP. To tackle the accuracy loss of DSP inference, we introduce the Anchor Frame Calibration mechanism, utilizing high-precision GPU inference results and frame similarities to mitigate accuracy loss on the DSP. Additionally, we design a lightweight progressive scheduler to distribute video frames to GPU and DSP, maximizing inference Average Precision (AP) under performance (i.e., frame rate) and power constraints. COUPLE has been implemented on the Qualcomm's Snapdragon 888 mobile SoC, extensive evaluation results demonstrate its efficacy in imnroving the inference performance and accuracy.
Hao Bao, Zhi Zhou 0006, Fei Xu 0009, Xu Chen 0004
ICDE3
2024 FedClust: Tackling Data Heterogeneity in Federated Learning through Weight-Driven Client Clustering
abstract
Federated learning (FL) is an emerging distributed machine learning paradigm that enables collaborative training of machine learning models over decentralized devices without exposing their local data. One of the major challenges in FL is the presence of uneven data distributions across client devices, violating the well-known assumption of independent-and-identically-distributed (IID) training samples in conventional machine learning. To address the performance degradation issue incurred by such data heterogeneity, clustered federated learning (CFL) shows its promise by grouping clients into separate learning clusters based on the similarity of their local data distributions. However, state-of-the-art CFL approaches require a large number of communication rounds to learn the distribution similarities during training until the formation of clusters is stabilized. Moreover, some of these algorithms heavily rely on a predefined number of clusters, thus limiting their flexibility and adaptability. In this paper, we propose FedClust, a novel approach for CFL that leverages the correlation between local model weights and the data distribution of clients. FedClust groups clients into clusters in a one-shot manner by measuring the similarity degrees among clients based on the strategically selected partial weights of locally trained models. We conduct extensive experiments on four benchmark datasets with different non-IID data settings. Experimental results demonstrate that FedClust achieves higher model accuracy up to ∼ 45% as well as faster convergence with a significantly reduced communication cost up to 2.7 × compared to its state-of-the-art counterparts.
Md Sirajul Islam, Simin Javaherian, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
ICPP3
2024 Hadar: Heterogeneity-Aware Optimization-Based Online Scheduling for Deep Learning Cluster
abstract
With the wide adoption of deep neural network (DNN) models for various applications, enterprises, and cloud providers have built deep learning clusters and increasingly deployed specialized accelerators, such as GPUs and TPUs, for DNN training jobs. To arbitrate cluster resources among multi-user jobs, existing schedulers fall short, either lacking fine-grained heterogeneity awareness or hardly generalizable to various scheduling policies. To fill this gap, we propose a novel design of a task-level heterogeneity-aware scheduler, Hadar, based on an online optimization framework that can express other scheduling algorithms. Hadar leverages the performance traits of DNN jobs on a heterogeneous cluster, characterizes the task-level performance heterogeneity in the optimization problem, and makes scheduling decisions across both spatial and temporal dimensions. The primal-dual framework is employed, with our design of a dual subroutine, to solve the optimization problem and guide the scheduling design. Extensive trace-driven simulations with representative DNN models have been conducted to demonstrate that Hadar improves the average job completion time (JCT) by 3× over an Apache YARN-based resource manager used in production. Moreover, Hadar outperforms Gavel [1], the state-of-the-art heterogeneity-aware scheduler, by 2.5× for the average JCT, shortens the queuing delay by 13%, and improves FTF (Finish-Time-Fairness) by 1.5%.
Abeda Sultana, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
IPDPS2
2024 HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
abstract
Deep Neural Network (DNN) inference on serverless functions is gaining prominence due to its potential for substantial budget savings. Existing works on serverless DNN inference solely optimize batching requests from one application with a single Service Level Objective (SLO) on CPU functions. However, production serverless DNN inference traces indicate that the request arrival rate of applications is surprisingly low, which inevitably causes a long batching time and SLO violations. Hence, there is an urgent need for batching multiple DNN inference requests with diverse SLOs (i.e., multi-SLO DNN inference) in serverless platforms. Moreover, the potential performance and cost benefits of deploying heterogeneous (i.e., CPU and GPU) functions for DNN inference have received scant attention.In this paper, we present HarmonyBatch, a cost-efficient resource provisioning framework designed to achieve predictable performance for multi-SLO DNN inference with heterogeneous serverless functions. Specifically, we construct an analytical performance and cost model of DNN inference on both CPU and GPU functions, by explicitly considering the GPU time-slicing scheduling mechanism and request arrival rate distribution. Based on such a model, we devise a two-stage merging strategy in HarmonyBatch to judiciously batch the multi-SLO DNN inference requests into application groups. It aims to minimize the budget of function provisioning for each application group while guaranteeing diverse performance SLOs of inference applications. We have implemented a prototype of HarmonyBatch on Alibaba Cloud Function Compute. Extensive prototype experiments with representative DNN inference workloads demonstrate that HarmonyBatch can provide predictable performance to serverless DNN inference workloads while reducing the monetary cost by up to 82.9% compared to the state-of-the-art methods.
Fei Xu 0009, Yikun Gu, Li Chen 0019, Fangming Liu, Zhi Zhou 0006
IWQoS2
2024 Taming Serverless Cold Start of Cloud Model Inference With Edge Computing
abstract
Serverless computing is envisioned as the de-facto standard for next-generation cloud computing. However, the cold start dilemma has impeded its adoption by delay-sensitive and burst applications. In this paper, we propose to tame serverless cold start in a cloud inference system with edge computing. Specifically, the proposed solution smooths the serverless cloud workload with user-owned edge computing, reducing the number of cold starts. Leveraging the configurability of requests and serverless functions, the proposed solution further reduces the transmission latency and serverless cost by adapting request configuration (e.g., image resolution) and function configuration (e.g., memory). To alleviate the potential inference accuracy degradation incurred by configuration adaption, we aim to strike a nice balance between inference latency, cost, and accuracy. However, achieving this goal is non-trivial since the underlying optimization is non-convex and involves future uncertain information. To simultaneously address dual challenges, the presented cold-start-aware online algorithms apply the regularization technique to decompose the problem into separate convex subproblems. Then, it applies lazy switching to smooth the number of provisioned functions and thus reduces the cold start. Through rigorous theoretical analysis, realistic prototype evaluations on AWS Lambda, and trace-driven simulations, we comprehensively validate the theoretical and empirical performance of our proposed solution.
Kongyange Zhao, Zhi Zhou 0006, Lei Jiao 0002, Shen Cai, Fei Xu 0009, Xu Chen 0004
IEEE Trans. Mob. Comput.5
2024 Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters
abstract
Long-running containerized workloads (e.g., machine learning), which typically showtime-varyingpatterns, are increasingly prevailing in shared production clusters. To improve workload performance, current schedulers mainly focus on optimizingshort-termbenefits of cluster load balancing orinitial container placementon servers. However, this would inevitably bring manyinvalid migrations(i.e., containers are migrated back and forth among servers over a short time window), leading to significant service level objective (SLO) violations. This paper introducesTetris, amodel predictive control(MPC)-based container scheduling strategy to proactively migrate long-running workloads for cluster load balancing. Specifically, we first build a discrete-time dynamic model forlong-termoptimization of container scheduling. To solve such an optimization problem,Tetristhen employs two main components: (1) a container resource predictor, which leverages time-series analysis approaches to accurately predict the container resource consumption; (2) an MPC-based container scheduler that jointly optimizes the cluster load balancing and container migration costover a certain sliding time window. We implement and open source a prototype ofTetrisbased on K8s. Extensive prototype experiments and trace-driven simulations demonstrate thatTetriscan improve the cluster load balancing degree by up to 77.8% without incurring any SLO violations, compared to the state-of-the-art container scheduling strategies.
Fei Xu 0009, Xiyue Shen, Shuohao Lin, Li Chen 0019, Zhi Zhou 0006, Fen Xiao, Fangming Liu
IEEE Trans. Serv. Comput.1
2023 DAG-Aware Optimization for Geo-Distributed Data Analytics
abstract
Geo-distributed data analytics has been proposed to analyze geographically distributed data. Existing studies have achieved significant reductions in execution time and data transfer cost ($) of data analytics jobs by optimizing task placement. Given a directed acyclic graph (DAG)-style job, however, they mainly optimize each stage independently, and they tend to distribute tasks and intermediate data across all locations, potentially inflating execution time and data transfer cost of descendent stages and the whole job.
Qingyuan Wang 0005, Bin Gao 0013, Zhi Zhou 0006, Fei Xu 0009, Chenghao Ouyang
ICPP4
2023 ACTS: Autonomous Cost-Efficient Task Orchestration for Serverless Analytics
abstract
Serverless computing has become increasingly popular for cloud applications, due to its compelling properties of high-level abstractions, lightweight runtime, high elasticity and pay-per-use billing. In this revolutionary computing paradigm shift, challenges arise when adapting data analytics applications to the serverless environment, due to the lack of support for efficient state sharing, which attract ever-growing research attention. In this paper, we aim to exploit the advantages of task-level orchestration and fine-grained resource provisioning for data analytics on serverless platforms, with the hope of fulfilling the promise of serverless deployment to the maximum extent. To this end, we present ACTS, an autonomous cost-efficient task orchestration framework for serverless analytics. ACTS judiciously schedules and coordinates function tasks to mitigate cold-start latency and state sharing overhead. In addition, ACTS explores the optimization space of fine-grained workload distribution and function resource configuration for cost efficiency. We have deployed and implemented ACTS on AWS Lambda, evaluated with various data analytics workloads. Results from extensive experiments demonstrate that ACTS achieves up to 98% monetary cost reduction while maintaining superior job completion time performance, in comparison with the state-of-the-art baselines.
Jananie Jarachanthan, Li Chen 0019, Fei Xu 0009
IWQoS3
2023 spotDNN: Provisioning Spot Instances for Predictable Distributed DNN Training in the Cloud
abstract
Distributed Deep Neural Network (DDNN) training on cloud spot instances is increasingly compelling as it can significantly save the user budget. To handle unexpected instance revocations, provisioning a heterogeneous cluster using the asynchronous parallel mechanism becomes the dominant method for DDNN training with spot instances. However, blindly provisioning a cluster of spot instances can easily result in unpre-dictable DDNN training performance, mainly because bottlenecks occur on the parameter server network bandwidth and PCIe bandwidth resources, as well as the inadequate cluster heterogeneity. To address the challenges above, we propose spotDNN, a heterogeneity-aware spot instance provisioning framework that provides predictable performance for DDNN training in the cloud. By explicitly considering the contention for bottle-neck resources, we first build an analytical performance model of DDNN training in heterogeneous clusters. It leverages the weighted average batch size and convergence coefficient to quantify the DDNN training loss in heterogeneous clusters. Through a lightweight workload profiling, we further design a cost-efficient instance provisioning strategy which incorporates the bounds calculation and sliding window techniques to effectively guarantee the training performance service level objectives (SLOs). We have implemented a prototype of spotDNN and conducted extensive experiments on Amazon EC2. Experiment results show that spotDNN can deliver predictable DDNN training performance while reducing the monetary cost by up to 68.1% compared to the existing solutions, yet with acceptable runtime overhead.
Ruitao Shang, Fei Xu 0009, Zhuoyan Bai, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IWQoS2
2023 COUPLE: Accelerating Video Analytics on Heterogeneous Mobile Processors
abstract
Deep learning has achieved tremendous success in various fields, but its significant computational demands make inference on mobile devices extremely challenging. To address this issue, we propose the COUPLE system, which enables heterogeneous processors to collaborate on mobile devices for accelerating video analytics. Additionally, we design the Co-Optimize strategy which utilizes the inference results of GPU to mitigate the accuracy loss caused by DSP. Experimental results demonstrate that COUPLE can improve the inference Average Precision by up to 5% compared to existing solutions.
Hao Bao, Zhi Zhou 0006, Qianyi Huang, Fei Xu 0009, Xu Chen 0004
MobiCom5
2023 iGniter: Interference-Aware GPU Resource Provisioning for Predictable DNN Inference in the Cloud
abstract
GPUs are essential to accelerating the latency-sensitive deep neural network (DNN) inference workloads in cloud datacenters. To fully utilize GPU resources,spatial sharingof GPUs among co-located DNN inference workloads becomes increasingly compelling. However, GPU sharing inevitably bringssevere performance interferenceamong co-located inference workloads, as motivated by an empirical measurement study of DNN inference on EC2 GPU instances. While existing works on guaranteeing inference performance service level objectives (SLOs) focus on eithertemporal sharingof GPUs orreactiveGPU resource scaling and inference migration techniques, how toproactivelymitigate such severe performance interference has received comparatively little attention. In this paper, we proposeiGniter, aninterference-awareGPU resource provisioning framework for cost-efficiently achieving predictable DNN inference in the cloud.iGniteris comprised of two key components: (1) alightweightDNN inference performance model, which leverages the system and workload metrics that are practically accessible to capture the performance interference; (2) Acost-efficientGPU resource provisioning strategy thatjointlyoptimizes the GPU resource allocation and adaptive batching based on our inference performance model, with the aim of achieving predictable performance of DNN inference workloads. We implement a prototype ofiGniterbased on the NVIDIA Triton inference server hosted on EC2 GPU instances. Extensive prototype experiments on four representative DNN models and datasets demonstrate thatiGnitercan guarantee the performance SLOs of DNN inference workloads with practically acceptable runtime overhead, while saving the monetary cost by up to$25\%$in comparison to the state-of-the-art GPU resource provisioning strategies.
Fei Xu 0009, Jianian Xu, Li Chen 0019, Ruitao Shang, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Parallel Distributed Syst.1
2022 Cost-Efficient Continuous Edge Learning for Artificial Intelligence of Things
abstract
The accelerating convergence of artificial intelligence (AI) and Internet of Things (IoT) has sparked a recent wave of interest in Artificial Intelligence of Things (AIoT). By exploiting the novel paradigm of edge intelligence, emerging computational intensive and resource demanding AIoT applications can be efficiently supported at the network edge. However, due to the limited resource capacity and/or power budget of the edge node, AIoT applications typically deploy compressed AI models to achieve the goal of low-latency and energy-efficient model inference. However, compressed models inherently suffer from the curse of data drift, i.e., the inference data at the deployment stage diverges from the training data at the training stage, leading to reduced model inference accuracy. To handle this issue, continuous learning has been proposed to periodically retrain the AI models on new data in an incremental manner. In this article, we investigate how to coordinate the edge and the cloud resources to perform cost-efficient continuous learning, with the goal of simultaneously optimizing the model performance (in terms of accuracy and robustness) and resource cost. Leveraging the Lyapunov optimization theory, we design and analyze a cost-efficient optimization framework for making online decisions upon admission control, transmission scheduling, and resource provisioning, for the dynamically arrived new data samples of various AIoT applications. We examine the effectiveness of the proposed framework on navigating the performance–cost tradeoff theoretically and empirically through trace-driven simulations.
Zhi Zhou 0006, Fei Xu 0009, Hai Jin 0001
IEEE Internet Things J.3
2022 λDNN: Achieving Predictable Distributed DNN Training With Serverless Architectures
abstract
Serverless computing is becoming a promising paradigm for Distributed Deep Neural Network (DDNN) training in the cloud, as it allows users to decompose complex model training into a number offunctionswithout managing virtual machines or servers. Though provided with a simpler resource interface (i.e., function number and memory size), inadequate function resource provisioning (either under-provisioning or over-provisioning) easily leads tounpredictableDDNN training performance in serverless platforms. Our empirical studies on AWS Lambda indicate that, suchunpredictable performanceof serverless DDNN training is mainly caused by the resource bottleneck of Parameter Servers (PS) and small local batch size. In this article, we design and implement$\lambda$λDNN, a cost-efficient function resource provisioning framework to provide predictable performance for serverless DDNN training workloads, while saving the budget of provisioned functions. Leveraging the PS network bandwidth and function CPU utilization, we build alightweightanalytical DDNN training performance model to enable our design of$\lambda$λDNNresource provisioning strategy, so as to guarantee DDNN training performance with serverless functions. Extensive prototype experiments on AWS Lambda and complementary trace-driven simulations demonstrate that,$\lambda$λDNNcan deliver predictable DDNN training performance and save the monetary cost of function resources by up to 66.7 percent, compared with the state-of-the-art resource provisioning strategies, yet with an acceptable runtime overhead.
Fei Xu 0009, Yiling Qin, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers1
2022 An Online Framework for Joint Network Selection and Service Placement in Mobile Edge Computing
abstract
With the rapid development and deployment of 5G wireless technology, mobile edge computing (MEC) has emerged as a new computing paradigm to facilitate a large variety of infrastructures at the network edge to reduce user-perceived communication delay. One of the fundamental problems in this new paradigm is to preserve satisfactory quality-of-service (QoS) for mobile users in light of densely dispersed wireless communication environment and often capacity-constrained MEC nodes. Such user-perceived QoS, typically in terms of the end-to-end delay, is highly vulnerable to both access network bottleneck and communication delay. Previous works have primarily focused on optimizing the communication delay through dynamic service placement, while ignoring the critical effect of access network selection on the access delay. In this work, we study the problem of jointly optimizing the access network selection and service placement for MEC, with the objective of improving the QoS in a cost-efficient manner by judiciously balancing the access delay, communication delay, and service switching cost. Specifically, we propose an efficient online framework to decompose a long-term time-varying optimization problem into a series of one-shot subproblems. To address the NP-hardness of the one-shot problem, we design a computationally-efficient two-phase algorithm based on matching and game theory, which achieves a near-optimal solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations are conducted to validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009, Bo Li 0001
IEEE Trans. Mob. Comput.4
2022 Astrea: Auto-Serverless Analytics Towards Cost-Efficiency and QoS-Awareness
abstract
With the ability to simplify the code deployment with one-click upload and lightweight execution, serverless computing has emerged as a promising paradigm with increasing popularity. However, there remain open challenges when adapting data-intensive analytics applications to the serverless context, in which users ofserverless analyticsencounter the difficulty in coordinating computation across different stages and provisioning resources in a large configuration space. This paper presents our design and implementation ofAstrea, which configures and orchestrates serverless analytics jobs in an autonomous manner, while taking into account flexibly-specified user requirements.Astrearelies on the modeling of performance and cost which characterizes the intricate interplay among multi-dimensional factors (e.g., function memory size, degree of parallelism at each stage). We formulate an optimization problem based on user-specific requirements towards performance enhancement or cost reduction, and develop a set of algorithms based on graph theory to obtain the optimal job execution. We deployAstreain the AWS Lambda platform and conduct real-world experiments over representative benchmarks, including Big Data analytics and machine learning workloads, at different scales. Extensive results demonstrate thatAstreacan achieve the optimal execution decision for serverless data analytics, in comparison with various provisioning and deployment baselines. For example, when compared with three provisioning baselines,Astreamanages to reduce the job completion time by 21% to 69% under a given budget constraint, while saving cost by 20% to 84% without violating performance requirements.
Jananie Jarachanthan, Li Chen 0019, Fei Xu 0009, Bo Li 0001
IEEE Trans. Parallel Distributed Syst.3
2022 Eiffel: Efficient and Fair Scheduling in Adaptive Federated Learning
abstract
Emerging machine learning (ML) technologies, in combination with the increasing computational power of mobile devices, lead to the extensive adoption of ML-based applications. Different from conventional model training that needs to collect all the user data in centralized cloud servers, federated learning (FL) has recently drawn increasing research attention as it enables privacy-preserving model training. With FL, decentralized edge devices in participation, train their model copies locally over their siloed datasets, and periodically synchronize the model parameters. However, model training is computationally extensive which easily drains the battery of mobile devices. In addition, due to the uneven distribution of siloed datasets, the shared model may become biased. To address theefficiencyandfairnessconcerns in a resource-constrained federated learning setting, in this paper, we proposeEiffelto judiciously select mobile devices to participate in the global model aggregation, and adaptively adjust the frequency of local and global model updates.Eiffelaims to make scheduling and coordination for the federated learning towards both resource efficiency and model fairness. We have conducted theoretical analysis ofEiffelfrom the perspectives of fairness and convergence. Extensive experiments with a wide variety of real-world datasets and models, both on a networked prototype system and in a larger-scale simulated environment, have demonstrated that while maintaining similar accuracy performance,Eiffeloutperforms existing baselines with respect to reducing communication overhead by up to 6× for higher efficiency and improving the fairness metric by up to 57% compared to the state-of-the-art algorithms.
Abeda Sultana, Md. Mainul Haque, Li Chen 0019, Fei Xu 0009, Xu Yuan 0001
IEEE Trans. Parallel Distributed Syst.4
2021 AMPS-Inf: Automatic Model Partitioning for Serverless Inference with Cost Efficiency
abstract
The salient pay-per-use nature of serverless computing has driven its continuous penetration as an alternative computing paradigm for various workloads. Yet, challenges arise and remain open when shifting machine learning workloads to the serverless environment. Specifically, the restriction on the deployment size over serverless platforms combining with the complexity of neural network models makes it difficult to deploy large models in a single serverless function. In this paper, we aim to fully exploit the advantages of the serverless computing paradigm for machine learning workloads targeting at mitigating management and overall cost while meeting the response-time Service Level Objective (SLO). We design and implement AMPS-Inf, an autonomous framework customized for model inferencing in serverless computing. Driven by the cost-efficiency and timely-response, our proposed AMPS-Inf automatically generates the optimal execution and resource provisioning plans for inference workloads. The core of AMPS-Inf relies on the formulation and solution of a Mixed-Integer Quadratic Programming problem for model partitioning and resource provisioning with the objective of minimizing cost without violating response time SLO. We deploy AMPS-Inf on the AWS Lambda platform, evaluate with the state-of-the-art pre-trained models in Keras including ResNet50, Inception-V3 and Xception, and compare with Amazon SageMaker and three baselines. Experimental results demonstrate that AMPS-Inf achieves up to 98% cost saving without degrading response time performance.
Jananie Jarachanthan, Li Chen 0019, Fei Xu 0009, Bo Li 0001
ICPP3
2021 Prophet: Speeding up Distributed DNN Training with Predictable Communication Scheduling
abstract
Optimizing performance for Distributed Deep Neural Network (DDNN) training has recently become increasingly compelling, as the DNN model gets complex and the training dataset grows large. While existing works on communication scheduling mostly focus on overlapping the computation and communication to improve DDNN training performance, the GPU and network resources are still under-utilized in DDNN training clusters. To tackle this issue, in this paper, we design and implement a predictable communication scheduling strategy named Prophet to schedule the gradient transfer in an adequate order, with the aim of maximizing the GPU and network resource utilization. Leveraging our observed stepwise pattern of gradient transfer start time, Prophet first uses the monitored network bandwidth and the profiled time interval among gradients to predict the appropriate number of gradients that can be grouped into blocks. Then, these gradient blocks can be transferred one by one to guarantee high utilization of GPU and network resources while ensuring the priority of gradient transfer (i.e., low-priority gradients cannot preempt high-priority gradients in the network transfer). Prophet can make the forward propagation start as early as possible so as to greedily reduce the waiting (idle) time of GPU resources during the DDNN training process. Prototype experiments with representative DNN models trained on Amazon EC2 demonstrate that Prophet can improve the DDNN training performance by up to 40% compared with the state-of-the-art priority-based communication scheduling strategies, yet with negligible runtime performance overhead.
Qiang Qi, Ruitao Shang, Li Chen 0019, Fei Xu 0009
ICPP5
2021 Astra: Autonomous Serverless Analytics with Cost-Efficiency and QoS-Awareness
abstract
With the ability to simplify the code deployment with one-click upload and lightweight execution, serverless computing has emerged as a promising paradigm with increasing popularity. However, there remain open challenges when adapting data-intensive analytics applications to the serverless context, in which users of serverless analytics encounter with the difficulty in coordinating computation across different stages and provisioning resources in a large configuration space. This paper presents our design and implementation of Astra, which configures and orchestrates serverless analytics jobs in an autonomous manner, while taking into account flexibly-specified user requirements. Astra relies on the modeling of performance and cost which characterizes the intricate interplay among multi-dimensional factors (e.g., function memory size, degree of parallelism at each stage). We formulate an optimization problem based on user-specific requirements towards performance enhancement or cost reduction, and develop a set of algorithms based on graph theory to obtain optimal job execution. We deploy Astra in the AWS Lambda platform and conduct real-world experiments over three representative benchmarks with different scales. Results demonstrate that Astra can achieve the optimal execution decision for serverless analytics, by improving the performance of 21% to 60% under a given budget constraint, and resulting in a cost reduction of 20% to 80% without violating performance requirement, when compared with three baseline configuration algorithms.
Jananie Jarachanthan, Li Chen 0019, Fei Xu 0009, Bo Li 0001
IPDPS3
2021 Rationing bandwidth resources for mitigating network resource contention in distributed DNN training clusters
Qiang Qi, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006
CCF Trans. High Perform. Comput.2
2020 E-LAS: Design and Analysis of Completion-Time Agnostic Scheduling for Distributed Deep Learning Cluster
abstract
With the prosperity of deep learning, enterprises, and large platform providers, such as Microsoft, Amazon, and Google, have built and provided GPU clusters to facilitate distributed deep learning training. As deep learning training workloads are heterogeneous, with a diverse range of characteristics and resource requirements, it becomes increasingly crucial to design an efficient and optimal scheduler for distributed deep learning jobs in the GPU cluster. This paper aims to propose a simple and yet effective scheduler, called E-LAS, with the objective of reducing the averaged training completion time of deep learning jobs. Without relying on the estimation or prior knowledge of the job running time, E-LAS leverages the real-time epoch progress rate, unique for distributed deep learning training jobs, as well as the attained services from temporal and spatial domains, to guide the scheduling decisions. The theoretical analysis for E-LAS is conducted to offer a deeper understanding on the components of scheduling criteria. Furthermore, we present a placement algorithm to achieve better resource utilization without involving much implementation overhead, complementary to the scheduling algorithm. Extensive simulations have been conducted, demonstrating that E-LAS improves the averaged job completion time (JCT) by 10 × over an Apache YARN-based resource manager used in production. Moreover, E-LAS outperforms Tiresias, the state-of-the-art scheduling algorithm customized for deep learning jobs, by almost 1.5 × for the average JCT as well as queuing time.
Abeda Sultana, Li Chen 0019, Fei Xu 0009, Xu Yuan 0001
ICPP3
2020 Towards trusted and efficient SDN topology discovery: A lightweight topology verification scheme
Xinli Huang, Fei Xu 0009
Comput. Networks4
2019 Stage Delay Scheduling: Speeding up DAG-style Data Analytics Jobs with Resource Interleaving
abstract
To increase the resource utilization of datacenters, big data analytics jobs are commonly running stages in parallel which are organized into and scheduled according to the Directed Acyclic Graph (DAG). Through an in-depth analysis of the latest Alibaba cluster trace and our motivation experiments on Amazon EC2, however, we show that the CPU and network resources are still under-utilized due to the unwise stage scheduling, thereby prolonging the completion time of a DAG-style job (e.g., Spark). While existing works on reducing the job completion time focus on either task scheduling or job scheduling, stage scheduling has received comparably little attention. In this paper, we design and implement DelayStage, a simple yet effective stage delay scheduling strategy to interleave the cluster resources across the parallel stages, so as to increase the cluster resource utilization and speed up the job performance. With the aim of minimizing the makespan of parallel stages, DelayStage judiciously arranges the execution of stages in a pipelined manner to maximize the performance benefits of resource interleaving. Extensive prototype experiments on 30 Amazon EC2 instances and complementary trace-driven simulations show that DelayStage can improve the cluster resource utilization by up to 81.8% and reduce the job completion time by up to 41.3%, in comparison to the stock Spark and the state-of-the-art stage scheduling strategies, yet with acceptable runtime overhead.
Wujie Shao, Fei Xu 0009, Li Chen 0019, Haoyue Zheng, Fangming Liu
ICPP2
2019 Cynthia: Cost-Efficient Cloud Resource Provisioning for Predictable Distributed Deep Neural Network Training
abstract
It becomes an increasingly popular trend for deep neural networks with large-scale datasets to be trained in a distributed manner in the cloud. However, widely known as resource-intensive and time-consuming, distributed deep neural network (DDNN) training suffers from unpredictable performance in the cloud, due to the intricate factors of resource bottleneck, heterogeneity and the imbalance of computation and communication which eventually cause severe resource under-utilization. In this paper, we propose Cynthia, a cost-efficient cloud resource provisioning framework to provide predictable DDNN training performance and reduce the training budget. To explicitly explore the resource bottleneck and heterogeneity, Cynthia predicts the DDNN training time by leveraging a lightweight analytical performance model based on the resource consumption of workers and parameter servers. With an accurate performance prediction, Cynthia is able to optimally provision the cost-efficient cloud instances to jointly guarantee the training performance and minimize the training budget. We implement Cynthia on top of Kubernetes by launching a 56-docker cluster to train four representative DNN models. Extensive prototype experiments on Amazon EC2 demonstrate that Cynthia can provide predictable training performance while reducing the monetary cost for DDNN workloads by up to 50.6%, in comparison to state-of-the-art resource provisioning strategies, yet with acceptable runtime overhead.
Haoyue Zheng, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
ICPP2
2019 Winning at the Starting Line: Joint Network Selection and Service Placement for Mobile Edge Computing
abstract
Mobile Edge Computing (MEC) is an emerging computing paradigm in which computational capabilities are pushed from the central cloud to the network edges. However, preserving the satisfactory quality-of-service (QoS) for user applications is non-trivial among multiple densely dispersed yet capacity constrained MEC nodes. This is mainly because both the access network and edge nodes are vulnerable to network congestion. Previous works are mostly limited to optimizing the QoS through dynamic service placement, while ignoring the critical effects of access network selection on the network congestion. In this paper, we study the problem of jointly optimizing the access network selection and service placement for MEC, towards the goal of improving the QoS by balancing the access, switching and communication delay. Specifically, we first design an efficient online framework to decompose the long-term optimization problem into a series of one-shot problems. To address the NP-hardness of the one-shot problem, we further propose an iteration-based algorithm to derive a computation efficient solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009
INFOCOM4
2019 Cost-Effective Cloud Server Provisioning for Predictable Performance of Big Data Analytics
abstract
Cloud datacenters are underutilized due to server over-provisioning. To increase datacenter utilization, cloud providers offer users an option to run workloads such as big data analytics on the underutilized resources, in the form of cheap yet revocable transient servers (e.g., EC2 spot instances, GCE preemptible instances). Though at highly reduced prices, deploying big data analytics on the unstable cloud transient servers can severely degrade the job performance due to instance revocations. To tackle this issue, this paper proposes iSpot, a cost-effective transient server provisioning framework for achieving predictable performance in the cloud, by focusing on Spark as a representative Directed Acyclic Graph (DAG)-style big data analytics workload. It first identifies the stable cloud transient servers during the job execution by devising an accurate Long Short-Term Memory (LSTM)-based price prediction method. Leveraging automatic job profiling and the acquired DAG information of stages, we further build an analytical performance model and present a lightweight critical data checkpointing mechanism for Spark, to enable our design of iSpot provisioning strategy for guaranteeing the job performance on stable transient servers. Extensive prototype experiments on both EC2 spot instances and GCE preemptible instances demonstrate that, iSpot is able to guarantee the performance of big data analytics running on cloud transient servers while reducing the job budget by up to 83.8 percent in comparison to the state-of-the-art server provisioning strategies, yet with acceptable runtime overhead.
Fei Xu 0009, Haoyue Zheng, Wujie Shao, Haikun Liu, Zhi Zhou 0006
IEEE Trans. Parallel Distributed Syst.1
2018 eBrowser: Making Human-Mobile Web Interactions Energy Efficient with Event Rate Learning
abstract
Due to the limited screen size of mobile devices, finger movements on touchscreen, such as scrolling and pinching (i.e., zooming in or out), are frequently used on mobile Web browsers and WebView-based apps, consuming considerable energy on mobile devices. While existing works on mobile Web browsers focus on reducing the power consumption or optimizing the performance of webpage loading, the power consumption of mobile Web interactions, especially after webpage loading, has received comparatively little attention. Motivated by an empirical study of the power consumption and user experience survey of human-mobile interactions, we design and implement eBrowser, an energy-efficient mobile Web interaction framework. It leverages a cloud-based machine learning model to enable personalized interaction event rate for individual users according to the interaction speed of their finger movement and the content of rendered webpages. To adapt to user behavior changes, eBrowser continuously monitors the interaction experience on each mobile device and periodically updates the personalized event rate model with incremental learning in the cloud. We implement eBrowser in Chromium and deploy the event rate model in a remote Aliyun cloud instance. Experimental results show that eBrowser reduces the energy consumption of mobile Web interactions by up to 43.8% with negligible runtime overhead, while guaranteeing user satisfaction on both mobile browsers and WebView-based apps.
Fei Xu 0009, Zhi Zhou 0006, Jia Rao
ICDCS1
2017 UFalloc: Towards Utility Max-min Fairness of Bandwidth Allocation for Applications in Datacenter Networks
Fei Xu 0009, Wangying Ye
Mob. Networks Appl.1
2016 Heterogeneity and Interference-Aware Virtual Machine Provisioning for Predictable Performance in the Cloud
abstract
Infrastructure-as-a-service (IaaS) cloud providers offer tenants elastic computing resources in the form of virtual machine (VM) instances to run their jobs. Recently, providing predictable performance (i.e., performance guarantee) for tenant applications is becoming increasingly compelling in IaaS clouds. However, the hardware heterogeneity and performance interference across the same type of cloud VM instances can bring substantial performance variation to tenant applications, which inevitably stops the tenants from moving their performance-sensitive applications to the IaaS cloud. To tackle this issue, this paper proposes Heifer, a Heterogeneity and interference-aware VM provisioning framework for tenant applications, by focusing on MapReduce as a representative cloud application. It predicts the performance of MapReduce applications by designing a lightweight performance model using the online-measured resource utilization and capturing VM interference. Based on such a performance model, Heifer provisions the VM instances of the good-performing hardware type (i.e., the hardware that achieves the best application performance) to achieve predictable performance for tenant applications, by explicitly exploring the hardware heterogeneity and capturing VM interference. With extensive prototype experiments in our local private cloud and a real-world public cloud (i.e., Microsoft Azure) as well as complementary large-scale simulations, we demonstrate that Heifer can guarantee the job performance while saving the job budget for tenants. Moreover, our evaluation results show that Heifer can improve the job throughput of cloud datacenters, such that the revenue of cloud providers can be increased, thereby achieving a win-win situation between providers and tenants.
Fei Xu 0009, Fangming Liu, Hai Jin 0001
IEEE Trans. Computers1
2015 Achieving Application-Level Utility Max-Min Fairness of Bandwidth Allocation in Datacenter Networks
Wangying Ye, Fei Xu 0009
CollaborateCom2
2015 When Software Defined Networks Meet Fault Tolerance: A Survey
Jinbang Chen, Fei Xu 0009, Min Yin
ICA3PP (3)3
2014 Managing Performance Overhead of Virtual Machines in Cloud Computing: A Survey, State of the Art, and Future Directions
abstract
Infrastructure-as-a-Service (IaaS) cloud computing offers customers (tenants) a scalable and economical way to provision virtual machines (VMs) on demand while charging them only for the leased computing resources by time. However, due to the VM contention on shared computing resources in datacenters, this new computing paradigm inevitably brings noticeable performance overhead (i.e., unpredictable performance) of VMs to tenants, which has become one of the primary issues of the IaaS cloud. Consequently, increasing efforts have recently been devoted to guaranteeing VM performance for tenants. In this survey, we review the state-of-the-art research on managing the performance overhead of VMs, and summarize them under diverse scenarios of the IaaS cloud, ranging from the single-server virtualization, a single mega datacenter, to multiple geodistributed datacenters. Specifically, we unveil the causes of VM performance overhead by illustrating representative scenarios, discuss the performance modeling methods with a particular focus on their accuracy and cost, and compare the overhead mitigation techniques by identifying their effectiveness and implementation complexity. With the obtained insights into the pros and cons of each existing solution, we further bring forth future research challenges pertinent to the modeling methods and mitigation techniques of VM performance overhead in the IaaS cloud.
Fei Xu 0009, Fangming Liu, Hai Jin 0001, Athanasios V. Vasilakos
Proc. IEEE1
2014 iAware: Making Live Migration of Virtual Machines Interference-Aware in the Cloud
abstract
Large-scale datacenters have been widely used to host cloud services, which are typically allocated to different virtual machines (VMs) through resource multiplexing across shared physical servers. Although recent studies have primarily focused on harnessing live migration of VMs to achieve load balancing and power saving among different servers, there has been little attention on the incurred performance interference and cost on both source and destination servers during and after such VM migration. To avoid potential violations of service-level-agreement (SLA) demanded by cloud applications, this paper proposes iAware, a lightweight interference-aware VM live migration strategy. It empirically captures the essential relationships between VM performance interference and key factors that are practically accessible through realistic experiments of benchmark workloads on a Xen virtualized cluster platform. iAware jointly estimates and minimizes both migration and co-location interference among VMs, by designing a simple multi-resource demand-supply model. Extensive experiments and complementary large-scale simulations are conducted to validate the performance gain and runtime overhead of iAware in terms of I/O and network throughput, CPU consumption, and scalability, compared to the traditional interference-unaware VM migration approaches. Moreover, we demonstrate that iAware is flexible enough to cooperate with existing VM scheduling or consolidation policies in a complementary manner, such that the load balancing or power saving can still be achieved without sacrificing performance.
Fei Xu 0009, Fangming Liu, Linghui Liu, Hai Jin 0001, Bo Li 0001, Baochun Li
IEEE Trans. Computers1
2011 Enhancing the Reliability of SIP Service in Large-Scale P2P-SIP Networks
Fei Xu 0009, Hai Jin 0001, Xiaofei Liao, Fei Qiu
GPC1