Minxian Xu

dblp:130/0196 · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
40since 2021 · last 2026
0000-0002-0046-5153ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 5 first-author · 16 since 2021Software engineering, systems software and programming languages · 17 · 5 first-author · 14 since 2021Computer networks · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SD-MoE: Scenario-Driven MoE Forecasting for Intelligent Elastic Scaling in Cloud Clusters
Xianzhao Guo, Weipeng Cao, Minxian Xu, Dachuan Li, Chuanfei Xu, Zhong Ming 0001
CCGrid3
2026 AQESF: An adaptive QoS-enhanced scheduling framework for online batch of task scheduling
Huikang Huang, Weiwei Lin 0001, Minxian Xu, Keqin Li 0001
Future Gener. Comput. Syst.3
2026 BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
abstract
ABSTRACT Objective Large Language Models (LLMs) are increasingly deployed in modern AI infrastructure, creating a strong demand for high‐throughput and resource‐efficient serving systems. Disaggregated LLM serving, which decouples prompt prefill from auto‐regressive decode to accommodate their heterogeneous compute and memory characteristics, has emerged as a promising architecture. However, existing disaggregated serving systems suffer from three fundamental limitations: static resource allocation that fails to adapt to highly dynamic workloads, severe load imbalance between compute‐bound prefill and memory‐bound decode stages, and prefix‐cache‐aware routing that skews load distribution and creates performance hotspots. These issues collectively limit resource utilization, scalability, and the ability to meet service level objectives (SLOs) under real‐world workloads. Methods To address these challenges, we propose BanaServe, a dynamic orchestration framework for disaggregated LLM serving that continuously rebalances both computational and memory resources across prefill and decode instances. BanaServe introduces three key mechanisms: (i) layer‐level weight migration to enable coarse‐grained redistribution of computation, (ii) attention‐level Key–Value (KV) cache migration for fine‐grained memory load balancing, and (iii) a Global KV Cache Store with layer‐wise overlapped transmission to decouple routing decisions from cache placement. Together, these mechanisms eliminate cache‐induced hotspots and allow routers to perform purely load‐aware scheduling with minimal latency overhead. BanaServe is implemented on top of state‐of‐the‐art LLM serving frameworks, including vLLM and DistServe. Results We evaluate BanaServe under diverse and challenging workloads, including long‐context inference, bursty request arrivals, and mixed prompt–generation patterns. Experimental results show that, compared to vLLM, BanaServe improves throughput by 1.2–3.9× and reduces total processing time by 3.9%–78.4%. In comparison with DistServe, BanaServe achieves 1.1–2.8× higher throughput while reducing latency by 1.4%–70.1%. These gains are consistent across workload variations, demonstrating BanaServe's robustness under highly dynamic serving conditions. Conclusion BanaServe demonstrates that dynamic, multi‐granularity resource rebalancing and cache‐decoupled routing are essential for efficient disaggregated LLM serving. By jointly addressing resource elasticity, stage imbalance, and cache‐induced load skew, BanaServe substantially improves throughput, latency, and resource utilization in real‐world deployments. This work provides a practical and scalable foundation for next‐generation LLM serving systems operating under dynamic and heterogeneous workloads.
Yiyuan He, Minxian Xu, Jingfeng Wu, Jianmin Hu, Chong Ma 0005, Cheng-Zhong Xu 0001, Lin Qu, Kejiang Ye
Softw. Pract. Exp.2
2026 C-Koordinator: Interference-Aware Management for Large-Scale and Co-Located Microservice Clusters
abstract
ABSTRACT Objective Microservices transform traditional monolithic applications into lightweight, loosely coupled application components and have been widely adopted in many enterprises. Cloud platform infrastructure providers enhance the resource utilization efficiency of microservices systems by co‐locating different microservices. However, this approach also introduces resource competition and interference among microservices. Designing interference‐aware strategies for large‐scale, co‐located microservice clusters is crucial for enhancing resource utilization and mitigating competition‐induced interference. These challenges are further exacerbated by unreliable metrics, application diversity, and node heterogeneity. Methods In this paper, we first analyze the characteristics of large‐scale and co‐located microservices clusters at Alibaba and further discuss why cycle per instruction (CPI) is adopted as a metric for interference measurement in large‐scale production clusters, as well as how to achieve accurate prediction of CPI through multi‐dimensional metrics. Based on CPI interference prediction and analysis, we also present the design of the C‐Koordinator platform, an open‐source solution utilized in Alibaba cluster, which incorporates co‐location and interference mitigation strategies. Results The interference prediction models consistently achieve over 90.3% accuracy, enabling precise prediction and rapid mitigation of interference in operational environments. As a result, application latency is reduced and stabilized across all percentiles (P50, P90, P99) response time (RT), achieving improvements ranging from 16.7% to 36.1% under various system loads compared with state‐of‐the‐art system. Conclusion These results demonstrate the system's ability to maintain smooth application performance in co‐located environments.
Shengye Song, Minxian Xu, Chengxi Gao, Fansong Zeng, Kejiang Ye, Cheng-Zhong Xu 0001
Softw. Pract. Exp.2
2026 Trust-Enabled Decentralized Task Offloading for Collaborative Edge Computing Using Blockchain and Deep Reinforcement Learning
abstract
ABSTRACT Objective Collaborative edge computing (CEC) addresses the service quality issues that arise from the limited resources of a single node in traditional edge computing architectures by integrating resources from multiple edge nodes. However, ensuring reliable task offloading in this collaborative environment remains a significant challenge. Existing solutions often struggle to balance the intelligence and trustworthiness of offloading decisions effectively. This imbalance can lead to poor performance and reduced task success rates, especially if tasks are offloaded to malicious nodes. Methods To tackle these challenges, this paper proposes a trust‐enabled decentralized task offloading scheme that combines blockchain technology and deep reinforcement learning (DRL). First, we introduce a blockchain‐based reputation mechanism within the CEC architecture to facilitate trusted collaboration among nodes, utilizing smart contracts for reputation management. Next, we propose a beta distribution‐based three‐factor reputation update (BTRU) algorithm to enhance the accuracy of reputation evaluation. Finally, we present a decentralized and trust‐enabled task offloading (DTTO) algorithm based on DRL, which uses on‐chain reputation data to guide agents in learning trustworthy task offloading policies, thereby maximizing offloading trustworthiness and task success rates. Result To thoroughly assess the effectiveness and practicality of our proposed scheme, we develop a testbed for CEC task offloading based on Kubernetes and Ethereum. Experimental results demonstrate that the BTRU algorithm effectively distinguishes malicious nodes, reducing their average reputation by 97.54%, with an improvement of 9.94% compared to competitive algorithms. Meanwhile, the DTTO algorithm significantly enhances the efficiency and reliability of task offloading, raising the task success rate by at least 3.04%, especially when the proportion of malicious nodes reaches 40%, its task success rate is at least 5.41% higher than that of competitive algorithms. Conclusion The proposed trust‐enabled decentralized task offloading scheme successfully combines blockchain‐based reputation management with DRL to achieve both intelligent and trustworthy task offloading in the CEC environments. The experimental validation confirms the scheme's effectiveness in identifying malicious nodes and improving task success rates under various system conditions.
Genyuan Yang, Wenjuan Li 0002, Qifei Zhang 0001, Minxian Xu, Chengjie Pan
Softw. Pract. Exp.4
2026 BrownoutServe: SLO-Aware Inference Serving Under Bursty Workloads for MoE-Based LLMs
abstract
In recent years, the Mixture-of-Experts (MoE) architecture has been widely applied to large language models (LLMs), providing a promising solution that activates only a subset of the model’s parameters during computation, thereby reducing overall memory requirements and allowing for faster inference compared to dense models. Despite these advantages, existing systems still face issues of low efficiency due to static model placement and lack of dynamic workloads adaptation. This leads to suboptimal resource utilization and increased latency, especially during bursty requests periods.To address these challenges, this paper introduces Brownout-Serve, a novel serving framework designed to optimize inference efficiency and maintain service reliability for MoE-based LLMs under dynamic computational demands and traffic conditions. BrownoutServe introduces “united experts” that integrate knowledge from multiple experts, reducing the times of expert access and inference latency. Additionally, it proposes a dynamic brownout mechanism to adaptively adjust the processing of certain tokens, optimizing inference performance while guaranteeing service level objectives (SLOs) are met. Our evaluations show the effectiveness of BrownoutServe under various workloads: it achieves up to 2.46× throughput improvement compared to state-of- the-art systems and reduces SLO violations by up to 90.28%, showcasing its robustness under bursty traffic while maintaining acceptable inference accuracy.
Jianmin Hu, Minxian Xu, Kejiang Ye, Cheng-Zhong Xu 0001
IEEE Trans. Computers2
2026 DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
abstract
To meet strict Service-Level Objectives (SLO), contemporary Large Language Models (LLMs) decouple the prefill and decoding stages and place them on separate GPUs to mitigate the distinct bottlenecks inherent to each stage. However, the heterogeneity of LLM workloads causes producer-consumer imbalance between the two instance types in such disaggregated architecture. To address this problem, we propose DOPD (Dynamic Optimal Prefill/Decoding), a dynamic LLM inference system that adjusts instance allocations to achieve an optimal prefill-to-decoding (P/D) ratio based on real-time load monitoring. Combined with an appropriate request-scheduling policy, DOPD effectively resolves imbalances between prefill and decoding instances and mitigates resource allocation mismatches due to mixed-length requests under high concurrency. Experimental evaluations show that, compared with vLLM and DistServe (representative aggregation-based and disaggregation-based approaches), DOPD improves overall system goodput by up to$1.5\times$, decreases P90 time-to-first-token (TTFT) by up to 67.5%, and decreases P90 time-per-output-token (TPOT) by up to 22.8%. Furthermore, our dynamic P/D adjustment technique performs proactive reconfiguration based on historical load, achieving over 99% SLO attainment while using fewer additional resources.
Junhan Liao, Minxian Xu, Wanyi Zheng, Yan Wang 0037, Kejiang Ye, Rajkumar Buyya, Cheng-Zhong Xu 0001
IEEE Trans. Serv. Comput.2
2025 Gravity-GNN: Deep Reinforcement Learning Guided Space Gravity-based Graph Neural Network
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable capabilities in handling graph data. Typically, GNNs recursively aggregate node information, including node features and local topological information, through a message-passing scheme. However, most existing GNNs are highly sensitive to neighborhood aggregation, and irrelevant information in the graph topology can lead to inefficient or even invalid node embeddings. To overcome these challenges, we propose a novel Space Gravity-based Graph Neural Network (Gravity-GNN) guided by Deep Reinforcement Learning (DRL). In particular, we introduce a novel similarity measure called ''node gravity'', inspired by the gravitational force between particles in space, to compare nodes within graph data. Furthermore, we employ DRL technology to learn and select the most suitable number of adjacent nodes for each node. Our experimental results on various real-world datasets demonstrate that Gravity-GNN outperforms state-of-the-art methods regarding node classification accuracy, while exhibiting greater robustness against disturbances.
Huaming Wu, Chaogang Tang, Pengfei Jiao, Minxian Xu, Huijun Tang
CIKM5
2025 GreenK8s: Green-aware Scheduling for Sustainable Kubernetes Cluster Management
abstract
With the rise of large-scale data centers and increasing demand for energy-efficient operations, there is a growing need to optimize the use of green energy in cloud computing environments. However, current schedulers focus solely on performance, lacking awareness of energy types and opportunities to promote green, low-carbon operations. This paper presents a Green-Aware Scheduling Framework for Kubernetes, named GreenK8s, aimed at minimizing the use of brown energy and maximizing the utilization of renewable energy sources, specifically solar power. Our framework integrates real-time power consumption monitoring with predictive solar energy models to intelligently schedule workloads based on energy availability. The proposed solution incorporates an AI-based solar power prediction model, Pod oversubscription strategies, and a novel scheduler, enabling Kubernetes to dynamically adapt to both the type and availability of green energy. Extensive experiments using the real-world Google Borg dataset and a realistic Kubernetes testbed demonstrate that GreenK8s reduces total energy consumption by up to 39 % and increases the average share of green energy in total consumption to 50.65 %, compared to state-of-the-art baselines. This work provides a promising approach to improve operational efficiency and sustainability in data centers.
Minxian Xu, Adel Nadjaran Toosi
CLUSTER2
2025 SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
abstract
As securities trading systems transition to a microservices architecture, optimizing system performance presents challenges such as inefficient resource scheduling and high service response delays. Existing container orchestration platforms lack tailored performance optimization mechanisms for trading scenarios, making it difficult to meet the stringent 50ms response time requirement imposed by exchanges. This paper introduces SealOS+, a Sealos-based performance optimization approach for securities trading, incorporating an adaptive resource scheduling algorithm leveraging deep reinforcement learning, a three-level caching mechanism for trading operations, and a Long Short-Term Memory (LSTM) based load prediction model. Real-world deployment at a securities exchange demonstrates that the optimized system achieves an average CPU utilization of 78%, reduces transaction response time to 105ms, and reaches a peak processing capacity of 15,000 transactions per second, effectively meeting the rigorous performance and reliability demands of securities trading.
Haojie Jia, Minxian Xu, Kejiang Ye
ICCCN4
2025 TD3-Sched: Learning to Orchestrate Container-Based Cloud-Edge Resources via Distributed Reinforcement Learning
Shengye Song, Minxian Xu, Kan Hu, Wenxia Guo, Kejiang Ye
PDCAT2
2025 LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
abstract
ABSTRACT Objective The microservices architecture has become a dominant paradigm in cloud computing due to its advantages in development, deployment, modularity, and scalability. Ensuring Quality of Service (QoS) through efficient Service Level Objective (SLO) resource allocation is a critical challenge. Current frameworks for microservice autoscaling based on SLOs often rely on heavy and complex models that are time‐consuming and resource‐intensive, making them unsuitable for rapidly changing environments and highly dynamic workloads. Methods This study proposes LSRAM (Lightweight SLO Resource Allocation Management), a novel framework designed to overcome the limitations of existing SLO‐based autoscaling methods. LSRAM operates in two stages: 1). Lightweight SLO Resource Allocation Model: Computes optimal SLO resource allocation for each microservice using a gradient descent method, ensuring rapid computation and minimal computational overhead. 2). SLO Resource Update Model: Adapts resource allocation dynamically in response to changes in the cluster environment, such as varying loads and application types, without requiring extensive retraining. Results LSRAM effectively addresses scenarios involving bursty traffic and fluctuating workloads. Compared to state‐of‐the‐art SLO allocation frameworks, LSRAM achieves the following: 1). Reduces resource usage by 17%. 2). Maintains QoS guarantees for users, even under dynamic conditions. 3). Demonstrates faster adaptability to changes in the system environment due to its lightweight design. Conclusion LSRAM offers a scalable, efficient, and adaptive solution for SLO‐based resource allocation in microservices architectures. By reducing resource usage while maintaining QoS, it provides a robust framework for managing dynamic and unpredictable workloads in cloud environments. Its lightweight design ensures practical applicability and superior performance compared to traditional, resource‐intensive methods.
Kan Hu, Minxian Xu, Kejiang Ye, Cheng-Zhong Xu 0001
Softw. Pract. Exp.2
2025 Cloudnativesim: A Toolkit for Modeling and Simulation of Cloud-Native Applications
abstract
ABSTRACT Background Cloud‐native applications are increasingly becoming popular in modern software design. Employing a microservice‐based architecture into these applications is a prevalent strategy that enhances system availability and flexibility. However, cloud‐native applications introduce new challenges, including frequent inter‐service communication and the management of heterogeneous codebases and hardware, resulting in unpredictable complexity and dynamism. Furthermore, as applications scale, only limited research teams or enterprises possess the resources for large‐scale deployment and testing, which impedes progress in the cloud‐native domain. Aims To address these challenges, we propose CloudNativeSim, a simulator for cloud‐native applications with a microservice‐based architecture. Results CloudNativeSim offers several key benefits: (i) comprehensive and dynamic modeling for cloud‐native applications, (ii) an extended simulation framework with new policy interfaces for scheduling cloud‐native applications, and (iii) support for customized application scenarios and user feedback based on Quality of Service (QoS) metrics. Conclusion CloudNativeSim can be easily deployed on standard computers to manage a high volume of requests and services. Its performance was validated through a case study, demonstrating higher than 94.5% accuracy in terms of response time simulation. The study further highlights the feasibility of CloudNativeSim by illustrating the effects of various scaling policies.
Jingfeng Wu, Minxian Xu, Yiyuan He, Kejiang Ye, Cheng-Zhong Xu 0001
Softw. Pract. Exp.2
2025 StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
abstract
Microservice architecture has transformed traditional monolithic applications into lightweight components. Scaling these lightweight microservices is more efficient than scaling servers. However, scaling microservices still faces the challenges resulting from the unexpected spikes or bursts of requests, which are difficult to detect and can degrade performance instantaneously. To address this challenge and ensure the performance of microservice-based applications, we propose a status-aware and elastic scaling framework called StatuScale , which is based on load status detector that can select appropriate elastic scaling strategies for differentiated resource scheduling in vertical scaling. Additionally, StatuScale employs a horizontal scaling controller that utilizes comprehensive evaluation and resource reduction to manage the number of replicas for each microservice. We also present a novel metric named correlation factor to evaluate the resource usage efficiency. Finally, we use Kubernetes, an open source container orchestration and management platform, and realistic traces from Alibaba to validate our approach. The experimental results have demonstrated that the proposed framework can reduce the average response time in the Sock-Shop application by 8.59% to 12.34% and in the Hotel-Reservation application by 7.30% to 11.97%, decrease service level objective violations, and offer better performance in resource usage compared to baselines.
Linfeng Wen 0001, Minxian Xu, Sukhpal Singh, Muhammad Hafizhuddin Hilman, Satish Narayana Srirama, Kejiang Ye, Cheng-Zhong Xu 0001
ACM Trans. Auton. Adapt. Syst.2
2025 A Cross-Workload Power Prediction Method Based on Transfer Gaussian Process Regression in Cloud Data Centers
abstract
Nowadays, machine learning (ML)-based power prediction models for servers have shown remarkable performance, leveraging large volumes of labeled data for training. However, collecting extensive labeled power data from servers in cloud data centers incurs substantial costs. Additionally, varying resource demands across different workloads (e.g., CPU-intensive, memory-intensive, and I/O-intensive) lead to significant differences in power consumption behaviors, known as domain shift. Consequently, power data collected from one type of workload cannot effectively train power prediction models for other workloads, limiting the exploration of the collected power data. To tackle these challenges, we proposeTGCP, a cross-workload power prediction method based on multi-source transfer Gaussian process regression.TGCPtransfers knowledge from abundant power data across multiple source workloads to a target workload with limited power data. Furthermore, Continuous normalizing flows adjust the posterior prediction distribution of Gaussian process, making it locally non-Gaussian, enhancingTGCP's ability to handle real-world power data distribution. This method enhances prediction accuracy for the target workload while reducing the expense of acquiring power data for real cloud data centers. Experimental results on a realistic power consumption dataset demonstrate thatTGCPsurpasses four traditional ML methods and three transfer learning methods in cross-workload power prediction.
Ruichao Mo, Weiwei Lin 0001, Haocheng Zhong, Minxian Xu, Keqin Li 0001
IEEE Trans. Cloud Comput.4
2025 GMHA: Growable Meta-Heuristic Algorithm for Multi-Objective Optimization Problems and its Application in Cloud Scheduling
Minxian Xu, Jiashu Zhang, Rajkumar Buyya
IEEE Trans. Serv. Comput.2
2024 TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
abstract
Cloud native solutions are widely applied in various fields, placing higher demands on the efficient management and utilization of resource platforms. To achieve the efficiency, load forecasting and elastic scaling have become crucial technologies for dynamically adjusting cloud resources to meet user demands and minimizing resource waste. However, existing prediction-based methods lack comprehensive analysis and integration of load characteristics across different time scales. For instance, long-term trend analysis helps reveal long-term changes in load and resource demand, thereby supporting proactive resource allocation over longer periods, while short-term volatility analysis can examine short-term fluctuations in load and resource demand, providing support for real-time scheduling and rapid response. In response to this, our research introduces TempoScale, which aims to enhance the comprehensive understanding of temporal variations in cloud workloads, enabling more intelligent and adaptive decision-making for elastic scaling. TempoScale utilizes the Complete Ensemble Empirical Mode Decomposition with Adaptive Noise algorithm to decompose time-series load data into multiple Intrinsic Mode Functions (IMF) and a Residual Component (RC). First, we integrate the IMF, which represents both long-term trends and short-term fluctuations, into the time series prediction model to obtain intermediate results. Then, these intermediate results, along with the RC, are transferred into a fully connected layer to obtain the final result. Finally, this result is fed into the resource management system based on Kubernetes for resource scaling. Our proposed approach can reduce the Mean Square Error by 5.80% to 30.43% compared to the baselines, and reduce the average response time by 5.58% to 31.15%. The results demonstrate the effectiveness of our proposed method in reducing violations of service-level objectives and providing better performance in terms of resource utilization.
Linfeng Wen 0001, Minxian Xu, Adel Nadjaran Toosi, Kejiang Ye
CLOUD2
2024 Resource Management for GPT-Based Model Deployed on Clouds: Challenges, Solutions, and Future Directions
Yongkang Dang, Yiyuan He, Minxian Xu, Kejiang Ye
ICA3PP (2)3
2024 UELLM: A Unified and Efficient Approach for Large Language Model Inference Serving
Yiyuan He, Minxian Xu, Jingfeng Wu, Wanyi Zheng, Kejiang Ye, Cheng-Zhong Xu 0001
ICSOC (1)2
2024 MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
abstract
Service Level Objectives (SLOs) aim to set threshold for service time in cloud services to ensure acceptable quality of service (QoS) and user satisfaction. Currently, many studies consider SLOs as a system resource to be allocated, ensuring QoS that to meet the SLOs. Existing microservice auto-scaling frameworks that rely on SLO resources often utilize complex and computationally intensive models, requiring significant time and resources to determine appropriate resource allocation. This paper aims to rapidly allocate SLO resources and minimize resource costs while ensuring application QoS meets the SLO requirements in a dynamically changing microservice environment. We propose MSARS, a framework that leverages meta-learning to quickly derive SLO resource allocation strategies and employs reinforcement learning for adaptive scaling of microservice resources. It features three innovative components: first, MSARS uses graph convolutional networks to predict the most suitable SLO resource allocation scheme for the current environment. Second, MSARS utilizes meta-learning to enable the graph neural network to quickly adapt to environmental changes ensuring adaptability in highly dynamic microservice environments. Third, MSARS generates auto-scaling policies for each microservice based on an improved Twin Delayed Deep Deterministic Policy Gradient (TD3) model. The adaptive auto-scaling policy integrates the SLO resource allocation strategy into the scheduling algorithm to satisfy SLOs. Finally, we compare MSARS with state-of-the-art resource auto-scaling algorithms that utilize neural networks and reinforcement learning, MSARS takes 40% less time to adapt to new environments, 38% reduction of SLO violations, and 8% less resources cost.
Kan Hu, Linfeng Wen 0001, Minxian Xu, Kejiang Ye
ISPA3
2024 TollHelper: A Safe and Efficient Traffic Control Approach on Toll Plaza via Constrained Load Balancing
abstract
Traffic congestion at toll plazas is a critical issue in urban infrastructure, which is often exacerbated by surges in vehicle volume during peak hours. The congestion typically arises from imbalances in traffic demand and toll booth efficiency, often resulting in safety hazards and delays. Existing solutions, while addressing efficiency or safety aspects, often lack a comprehensive approach for efficient traffic management at toll plazas. To address this challenge, in this paper, we propose TollHelper, a framework designed to optimize vehicle scheduling and load balancing at toll plazas as well as improve safety. Our approach treats concurrently arriving vehicles as a single scheduling batch, guiding different vehicles from the same batch to different toll booths to enhance safety and reduce congestion. We address this scheduling constraint in both general and heterogeneous toll booth scenarios, introducing effective load balancing algorithms to minimize toll booth service loads and optimize user driving experiences. Based on empirical studies, we demonstrate that our methods achieve improvements in standard deviation compared to the baselines, ranging from 51.6% to 97.0% improvement in terms of load balancing effects.
Mengbing Zhou, Bocong Zhao, Minxian Xu, Yang Wang 0006
ISPA3
2024 Special issue on efficient management of microservice-based systems and applications
abstract
Special issue on efficient management of microservice-based systems and applicationsThe advent of microservice architecture marks a transition from conventional monolithic applications to a landscape of loosely linked, lightweight, and autonomous microservice components.The primary objective is to ensure strong environmental uniformity, portability across various operating systems, and robust resource isolation.Leading cloud service providers such as Amazon, Microsoft, Google, and Alibaba have widely embraced microservices within their infrastructures.This adoption is geared toward automating application management and optimizing system performance.Consequently, addressing the automation of tasks like deployment, maintenance, auto-scaling, and networking of microservices becomes pivotal.This underscores the importance of efficient management of systems and applications built on microservices as a critical research challenge.Efficient management methods must not only ensure the quality of service (QoS) across multiple microservices units (containers) but also provide greater control over individual components.However, the dynamic and varied nature of microservice applications and environments significantly amplifies the complexity of these management approaches.Each microservice unit can be deployed and operated independently, catering to distinct functionalities and business objectives.Furthermore, microservices can interact and combine through lightweight communication techniques to form a complete application.The expanding scale of microservice-based systems and their intricate interdependencies pose challenges in terms of load distribution and resource management at the infrastructure level.Furthermore, as cloud workloads surge in resource demands, bandwidth consumption, and QoS requirements, the traditional cloud computing environment extends to fog and edge infrastructures that are in close proximity to end users.As a result, current microservice management approaches need further enhancement to address the mounting resource diversity, application distribution, workload profiles, security prerequisites, and scalability demands across hybrid cloud infrastructures.Keeping this in mind, this special issue addressed some of the aspects related to efficient management of microservice-based systems and applications with the focus on various challenges faced, and promising solutions to address such challenges by using software engineering, machine learning and deep learning techniques.We have received 21 submissions in this issue, and we accepted six high-quality submissions for publication after a rigorous review process with at least three reviewers for each paper.The authors are from diverse countries, including the USA, China, UK, Germany, India, Brazil, etc.Each of the accepted papers is summarized as follows.In the first article, Batista et al. 1 presented two strategies for handling asynchronous workloads associated with tax integration in a multi-tenant microservice architecture specific to the company's context.The initial approach utilizes polling, employing a queue as a distributed lock.The second approach, named the single active consumer, adopts a push-based technique, leveraging the message broker's logic for message delivery.These methodologies are designed to optimize resource allocation in scenarios involving an increasing number of container replicas and tenants.In the second article, Kumar et al. 2 introduced a resource allocation model designed to enhance the QoS in microservice deployment.Utilizing a Fine-tuned Sunflower Whale Optimization Algorithm, the model strategically deploys container-based services on physical machines, optimizing their execution capacity by efficiently utilizing CPU and memory resources.The primary objective of this proposed technique is to achieve an efficient distribution of workload, preventing resource wastage and ultimately enhancing QoS parameters.In the third paper, Zhu et al. 3 introduced RADF, a semi-automatic approach for decomposing a monolith into serverless functions by analyzing the inherent business logic present in the application's interface.The proposed method employs a two-stage refactoring strategy, initially performing a coarse-grained decomposition followed by a fine-grained one.This approach streamlines the decomposition process into smaller, more manageable steps, providing adaptability to generate a solution at either the microservice or function level.In the fourth paper, Würz et al. 4 identified the principal tasks and subtasks of the application for the purpose of partitioning.Subsequently, they outlined the program flow to ascertain which application tasks could be transformed into functions and elucidated their interdependencies.In the concluding step, they precisely specified individual functions
Minxian Xu, Schahram Dustdar, Massimo Villari, Rajkumar Buyya
Softw. Pract. Exp.1
2024 Computation Energy Efficiency Maximization for NOMA-Based and Wireless-Powered Mobile Edge Computing With Backscatter Communication
abstract
In the Internet of Things (IoT) environment, a wide variety of mobile devices (MDs) have become part of it, leading to a dramatic increase in the amount of task data. However, due to the limited battery capacity and computing resources of MDs, a lot of effort is required to be taken on how to process more data with less energy. In this paper, we take into account the low utilization of spectrum resources and the short battery life of the equipment, and a backscatter communication-mobile edge computing (BC-MEC) network system based on Non-orthogonal multiple access (NOMA) communication mode is proposed. In order to maximize the computation energy efficiency (CEE) of the system, we jointly optimize the backscatter coefficient of each MD, the backscatter communication duration, the direct offloading duration, the MEC server processing time, the local processing time, the direct offloading power of each MD, the calculation frequency of the MEC server, and the local calculation frequency of each MD. We then formulate it as a joint fractional optimization problem, which is a non-convex optimization problem that is difficult to solve by heuristic algorithms with high computational complexity. To this end, we transform such a problem into a convex problem and apply the Lagrangian dual method to solve it efficiently. Furthermore, in order to meet different user requirements, two effective iterativeDinkelbach algorithms based onBackscatterCoefficientUpdates (DBCU) are proposed to solve this problem. Extensive simulation results demonstrate the superiority of our proposed approach, which improves the system CEE by at least 10% compared to state-of-the-art methods.
Junhui Du, Huaming Wu, Minxian Xu, Rajkumar Buyya
IEEE Trans. Mob. Comput.3
2024 DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters
abstract
Microservices have transformed monolithic applications into lightweight, self-contained, and isolated application components, establishing themselves as a dominant paradigm for application development and deployment in public clouds such as Google and Alibaba. Autoscaling emerges as an efficient strategy for managing resources allocated to microservices’ replicas. However, the dynamic and intricate dependencies within microservice chains present challenges to the effective management of scaled microservices. Additionally, the centralized autoscaling approach can encounter scalability issues, especially in the management of large-scale microservice-based clusters. To address these challenges and enhance scalability, we propose an innovative distributed resource provisioning approach for microservices based on the Twin Delayed Deep Deterministic Policy Gradient algorithm. This approach enables effective autoscaling decisions and decentralizes responsibilities from a central node to distributed nodes. Comparative results with state-of-the-art approaches, obtained from a realistic testbed and traces, indicate that our approach reduces the average response time by 15% and the number of failed requests by 24%, validating improved scalability as the number of requests increases.
Haoyu Bai, Minxian Xu, Kejiang Ye, Rajkumar Buyya, Cheng-Zhong Xu 0001
IEEE Trans. Serv. Comput.2
2024 Efficient Multi-Task Computation Offloading Game for Mobile Edge Computing
abstract
Mobile edge computing emerges to serve mobile users with low-latency computation offloading in edge networks, which are resource-constrained with massive users and workloads. However, existing communication and computing resource allocation schemes for offloaded tasks aren't efficient enough, where finished tasks still occupy resources, wasting constrained resources. Besides, the multi-user offloading is usually for scenarios of one task per user, ignoring real-worldmulti-taskoffloading scenarios where each user has multiple tasks, lack generality and flexibility. Meanwhile, local computing resource allocation schemes in multi-task scenarios ignore resource readjustment, causing low resource utilization. To solve these problems, we propose ECO-GAME, an efficient multi-task offloading scheme, which dynamically allocates bandwidth and computing resources to unfinished tasks, resulting in high resource utilization. We initially formulate the multi-task offloading problem as the game minimizing each user's cost, which is NP-hard. Thus we re-formulate the game utilizing potential games to optimize user's objective either locally or globally, and prove the existence of its Nash equilibrium. We then design an efficient multi-task offloading algorithm to obtain an approximate solution in polynomial time, together with computational complexity analysis. We further conduct performance evaluation on ECO-GAME utilizing price of anarchy. Numerical results demonstrate the efficiency of ECO-GAME, and show ECO-GAME reduces 49.2% cost over the state-of-the-art work, and scales well with the increasing number of tasks and users.
Shuhui Chu, Chengxi Gao, Minxian Xu, Kejiang Ye, Zhu Xiao, Cheng-Zhong Xu 0001
IEEE Trans. Serv. Comput.3
2024 Computation Energy Efficiency Maximization for Intelligent Reflective Surface-Aided Wireless Powered Mobile Edge Computing
abstract
A wide variety of Mobile Devices (MDs) are adopted in Internet of Things (IoT) environments, resulting in a dramatic increase in the volume of task data and greenhouse gas emissions. However, due to the limited battery power and computing resources of MD, it is critical to process more data with less energy. This paper studies the Wireless Power Transfer-based Mobile Edge Computing (WPT-MEC) network system assisted by Intelligent Reflective Surface (IRS) to enhance communication performance while improving the battery life of MD. In order to maximize the Computation Energy Efficiency (CEE) of the system and reduce the carbon footprint of the MEC server, we jointly optimize the CPU frequencies of MDs and MEC server, the transmit power of Power Beacon (PB), the processing time of MEC server, the offloading time and the energy harvesting time of MDs, the local processing time and the offloading power of MD and the phase shift coefficient matrix of Intelligent Reflecting Surface (IRS). Moreover, we transform this joint optimization problem into a fractional programming problem. We then propose the Dinkelbach Iterative Algorithm with Gradient Updates (DIA-GU) to solve this problem effectively. With the help of convex optimization theory, we can obtain closed-form solutions, revealing the correlation between different variables. Compared to other algorithms, the DIA-GU algorithm not only exhibits superior performance in enhancing the system's CEE but also demonstrates significant reductions in carbon emissions.
Junhui Du, Minxian Xu, Sukhpal Singh, Huaming Wu
IEEE Trans. Sustain. Comput.2
2024 ATOM: AI-Powered Sustainable Resource Management for Serverless Edge Computing Environments
abstract
Serverless edge computing decreases unnecessary resource usage on end devices with limited processing power and storage capacity. Despite its benefits, serverless edge computing's zero scalability is the major source of the cold start delay, which is yet unsolved. This latency is unacceptable for time-sensitive Internet of Things (IoT) applications like autonomous cars. Most existing approaches need containers to idle and use extra computing resources. Edge devices have fewer resources than cloud-based systems, requiring new sustainable solutions. Therefore, we propose an AI-powered, sustainable resource management framework called ATOM for serverless edge computing. ATOM utilizes a deep reinforcement learning model to predict exactly when cold start latency will happen. We create a cold start dataset using a heart disease risk scenario and deploy using Google Cloud Functions. To demonstrate the superiority of ATOM, its performance is compared with two different baselines, which use the warm-start containers and a two-layer adaptive approach. The experimental results showed that although the ATOM required more calculation time of 118.76 seconds, it performed better in predicting cold start than baseline models with an RMSE ratio of 148.76. Additionally, the energy consumption and$CO_{2}$emission amount of these models are evaluated and compared for the training and prediction phases.
Muhammed Golec, Sukhpal Singh, Félix Cuadrado, Ajith Kumar Parlikad, Minxian Xu, Huaming Wu, Steve Uhlig
IEEE Trans. Sustain. Comput.5
2023 ChainsFormer: A Chain Latency-Aware Resource Provisioning Approach for Microservices Cluster
Chenghao Song, Minxian Xu, Kejiang Ye, Huaming Wu, Sukhpal Singh, Rajkumar Buyya, Cheng-Zhong Xu 0001
ICSOC (1)2
2023 Practical model with strong interpretability and predictability: An explanatory model for individuals' destination prediction considering personal and crowd travel behavior
abstract
Abstract Real‐time individuals' destination prediction is of great significance for real‐time user tracking, service recommendation and other related applications. Traditional technology mainly used statistical methods based on the travel patterns mined from personal history travel data. However, it is not clear how to predict the destinations of individuals with only limited personal historical data. In this paper, taking the public transportation metro systems as example, we design a practical method called practical model with strong interpretability and predictability to predict each passenger's destination. Our main novelties are two aspects: (1) We propose to predict individuals' destination by combining personal and crowd behavior under certain context. (2) An explanatory model combining discrete choice model and neural network model is proposed to predict individuals' stochastic trip's destination, which can be applied to other transportation analysis scenarios about individuals' choice behavior such as travel mode choice or route choice. We validate our method based on extensive experiments, using smart card data collected by automatic fare collection system and weather data in Shenzhen, China. The experimental results demonstrate that our approach can achieve better performance than other baselines in terms of prediction accuracy.
Juanjuan Zhao 0001, Jiexia Ye, Minxian Xu, Cheng-Zhong Xu 0001
Concurr. Comput. Pract. Exp.3
2023 BlockFaaS: Blockchain-enabled Serverless Computing Framework for AI-driven IoT Healthcare Applications
Muhammed Golec, Sukhpal Singh, Mustafa Golec, Minxian Xu, Soumya K. Ghosh 0001, Salil S. Kanhere, Omer F. Rana, Steve Uhlig
J. Grid Comput.4
2023 Prepartition: Load Balancing Approach for Virtual Machine Reservations in a Cloud Data Center
Wenhong Tian, Minxian Xu, Kui Wu 0001, Cheng-Zhong Xu 0001, Rajkumar Buyya
J. Comput. Sci. Technol.2
2023 Flash: Joint Flow Scheduling and Congestion Control in Data Center Networks
abstract
Flow scheduling and congestion control are two important techniques to reduce flow completion time in data center networks. While existing works largely treat them independently, the interactions between flow scheduling and congestion control are in general overlooked which leads to sub-optimal solutions, especially given that the link capacity is increasing faster than the switch port buffer size. In this paper, we presentFlash, a simple yet effective scheme that integrates scheduling and congestion control. Specifically,Flashputs forward a congestion-aware scheduling scheme to determine the priority of flows based on the latest network congestion extent and the flow’s bytes sent. Besides,Flashproposes a priority-based packet dropping scheme in switch port buffers and implements a priority-aware congestion control scheme. Experiment results show thatFlashhas superior performance: (1) it has 35.8% lower tail latency than PIAS and performs similar with pFabric in a 10G network without knowing the flow size, (2) in 100G networks with shallow buffers, the information agnosticFlashhas 6.8% lower average FCT than the information-aware pFabric, (3) it outperforms pFabric by 13.5% in FCT if flow size is also known toFlash.
Chengxi Gao, Shuhui Chu, Hong Xu 0001, Minxian Xu, Kejiang Ye, Cheng-Zhong Xu 0001
IEEE Trans. Cloud Comput.4
2022 A multi-output prediction model for physical machine resource usage in cloud data centers
Yongde Zhang, Fagui Liu, Bin Wang 0048, Weiwei Lin 0001, Guoxiang Zhong, Minxian Xu, Keqin Li 0001
Future Gener. Comput. Syst.6
2022 BenchSubset: A framework for selecting benchmark subsets based on consensus clustering
abstract
The redundancy in the benchmark suite will increase the time for computer system performance evaluation and simulation. The most typical method to solve this problem is to select subsets based on clustering. However, it is a challenge to validate benchmark subsetting results for unlabeled benchmark suites when using the clustering method, and existing research has not considered this problem. Also, there is no quantitative evaluation method for subsetting which can reflect the universal and the diversity characteristics of the benchmark suite at the same time. To solve the above problems, we propose BenchSubset, a framework for selecting benchmark subsets based on consensus clustering, which includes Group Principal Components Analysis, consensus clustering, and a new evaluation method considering the universal and the diversity characteristics of the benchmark suite. We conducted SPEC CPU2017 subsetting experiments on Huawei's Taishan 200, then verified the effectiveness of BenchSubset in selecting a benchmark subset. Compared with the mainstream principal components analysis with hierarchical clustering (PCA-H) method, the benchmark subset selected by BenchSubset performs better in representing the universal and the diversity characteristics of SPEC CPU2017.
Hongping Zhan, Weiwei Lin 0001, Feiqiao Mao, Minxian Xu, Guangxin Wu, Guokai Wu, Jianzhuo Li
Int. J. Intell. Syst.4
2022 HUNTER: AI based holistic resource management for sustainable cloud computing
Shreshth Tuli, Sukhpal Singh, Minxian Xu, Peter Garraghan, Rami Bahsoon, Schahram Dustdar, Rizos Sakellariou, Omer F. Rana, Rajkumar Buyya, Giuliano Casale, Nicholas R. Jennings
J. Syst. Softw.3
2022 PDMA: Probabilistic service migration approach for delay-aware and mobility-aware mobile edge computing
abstract
Abstract As a key technology in the 5G era, mobile edge computing (MEC) has developed rapidly in recent years. MEC aims to reduce the service delay of mobile users, while alleviating the processing pressure on the core network. MEC can be regarded as an extension of cloud computing on the user side, which can deploy edge servers and bring computing resources closer to mobile users, and provide more efficient interactions. However, due to the user's dynamic mobility, the distance between the user and the edge server will change dynamically, which may cause fluctuations in Quality of Service. Therefore, when a mobile user moves in the MEC environment, certain approaches are needed to schedule services deployed on the edge server to ensure the user experience. In this article, we model service scheduling in MEC scenarios and propose a delay‐aware and mobility‐aware service management approach based on concise probabilistic methods. This approach has low computational complexity and can effectively reduce service delay and migration costs. Furthermore, we conduct experiments by utilizing multiple realistic datasets and use iFogSim to evaluate the performance of the algorithm. The results show that our proposed approach can optimize the performance on service delay, with 8%–20% improvement and reduce the migration cost by more than 75% compared with baselines during the rush hours.
Minxian Xu, Qiheng Zhou, Huaming Wu, Weiwei Lin 0001, Kejiang Ye, Cheng-Zhong Xu 0001
Softw. Pract. Exp.1
2022 CoScal: Multifaceted Scaling of Microservices With Reinforcement Learning
abstract
The emerging trend towards moving from monolithic applications to microservices has raised new performance challenges in cloud computing environments. Compared with traditional monolithic applications, the microservices are lightweight, fine-grained, and must be executed in a shorter time. Efficient scaling approaches are required to ensure microservices’ system performance under diverse workloads with strict Quality of Service (QoS) requirements and optimize resource provisioning. To solve this problem, we investigate the trade-offs between the dominant scaling techniques, including horizontal scaling, vertical scaling, and brownout in terms of execution cost and response time. We first present a prediction algorithm based on gradient recurrent units to accurately predict workloads assisting in scaling to achieve efficient scaling. Further, we propose a multi-faceted scaling approach using reinforcement learning called CoScal to learn the scaling techniques efficiently. The proposed CoScal approach takes full advantage of data-driven decisions and improves the system performance in terms of high communication cost and delay. We validate our proposed solution by implementing a containerized microservice prototype system and evaluated with two microservice applications. The extensive experiments demonstrate that CoScal reduces response time by 19%-29% and decreases the connection time of services by 16% when compared with the state-of-the-art scaling techniques for Sock Shop application. CoScal can also improve the number of successful transactions with 6%-10% for Stan’s Robot Shop application.
Minxian Xu, Chenghao Song, Shashikant Ilager, Sukhpal Singh, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001
IEEE Trans. Netw. Serv. Manag.1
2022 esDNN: Deep Neural Network Based Multivariate Workload Prediction in Cloud Computing Environments
abstract
Cloud computing has been regarded as a successful paradigm for IT industry by providing benefits for both service providers and customers. In spite of the advantages, cloud computing also suffers from distinct challenges, and one of them is the inefficient resource provisioning for dynamic workloads. Accurate workload predictions for cloud computing can support efficient resource provisioning and avoid resource wastage. However, due to the high-dimensional and high-variable features of cloud workloads, it is difficult to predict the workloads effectively and accurately. The current dominant work for cloud workload prediction is based on regression approaches or recurrent neural networks, which fail to capture the long-term variance of workloads. To address the challenges and overcome the limitations of existing works, we proposed an e fficient supervised learning-based D eep N eural Network ( esDNN ) approach for cloud workload prediction. First, we utilize a sliding window to convert the multivariate data into a supervised learning time series that allows deep learning for processing. Then, we apply a revised Gated Recurrent Unit (GRU) to achieve accurate prediction. To show the effectiveness of esDNN, we also conduct comprehensive experiments based on realistic traces derived from Alibaba and Google cloud data centers. The experimental results demonstrate that esDNN can accurately and efficiently predict cloud workloads. Compared with the state-of-the-art baselines, esDNN can reduce the mean square errors significantly, e.g., 15%. rather than the approach using GRU only. We also apply esDNN for machines auto-scaling, which illustrates that esDNN can reduce the number of active hosts efficiently, thus the costs of service providers can be optimized.
Minxian Xu, Chenghao Song, Huaming Wu, Sukhpal Singh, Kejiang Ye, Cheng-Zhong Xu 0001
ACM Trans. Internet Techn.1
2021 EEDTO: An Energy-Efficient Dynamic Task Offloading Algorithm for Blockchain-Enabled IoT-Edge-Cloud Orchestrated Computing
abstract
With the proliferation of compute-intensive and delay-sensitive mobile applications, large amounts of computational resources with stringent latency requirements are required on Internet-of-Things (IoT) devices. One promising solution is to offload complex computing tasks from IoT devices either to mobile-edge computing (MEC) or mobile cloud computing (MCC) servers. MEC servers are much closer to IoT devices and thus have lower latency, while MCC servers can provide flexible and scalable computing capability to support complicated applications. To address the tradeoff between limited computing capacity and high latency, and meanwhile, ensure the data integrity during the offloading process, we consider a blockchain scenario where edge computing and cloud computing can collaborate toward secure task offloading. We further propose a blockchain-enabled IoT-Edge-Cloud computing architecture that benefits both from MCC and MEC, where MEC servers offer lower latency computing services, while MCC servers provide stronger computation power. Moreover, we develop an energy-efficient dynamic task offloading (EEDTO) algorithm by choosing the optimal computing place in an online way, either on the IoT device, the MEC server or the MCC server with the goal of jointly minimizing the energy consumption and task response time. The Lyapunov optimization technique is applied to control computation and communication costs incurred by different types of applications and the dynamic changes of wireless environments. During the optimization, the best computing location for each task is chosen adaptively without requiring future system information as prior knowledge. Compared with previous offloading schemes with/without MEC and MCC cooperation, EEDTO can achieve energy-efficient offloading decisions with relatively lower computational complexity.
Huaming Wu, Katinka Wolter, Pengfei Jiao, Yubin Zhao, Minxian Xu
IEEE Internet Things J.6
2021 A Self-Adaptive Approach for Managing Applications and Harnessing Renewable Energy for Sustainable Cloud Computing
abstract
Rapid adoption of Cloud computing for hosting services and its success is primarily attributed to its attractive features such as elasticity, availability and pay-as-you-go pricing model. However, the huge amount of energy consumed by cloud data centers makes it to be one of the fastest growing sources of carbon emissions. Approaches for improving the energy efficiency include enhancing the resource utilization to reduce resource wastage and applying the renewable energy as the energy supply. This work aims to reduce the carbon footprint of the data centers by reducing the usage of brown energy and maximizing the usage of renewable energy. Taking advantage of microservices and renewable energy, we propose a self-adaptive approach for the resource management of interactive workloads and batch workloads. To ensure the quality of service of workloads, a brownout-based algorithm for interactive workloads and a deferring algorithm for batch workloads are proposed. We have implemented the proposed approach in a prototype system and evaluated it with web services under real traces. The results illustrate our approach can reduce the brown energy usage by 21 percent and improve the renewable energy usage by 10 percent.
Minxian Xu, Adel Nadjaran Toosi, Rajkumar Buyya
IEEE Trans. Sustain. Comput.1
2020 Energy Efficient Algorithms based on VM Consolidation for Cloud Computing: Comparisons and Evaluations
abstract
Cloud Computing paradigm has revolutionized IT industry and be able to offer computing as the fifth utility. With the pay-as-you-go model, cloud computing enables to offer the resources dynamically for customers anytime. Drawing the attention from both academia and industry, cloud computing is viewed as one of the backbones of the modern economy. However, the high energy consumption of cloud data centers contributes to high operational costs and carbon emission to the environment. Therefore, Green cloud computing is required to ensure energy efficiency and sustainability, which can be achieved via energy efficient techniques. One of the dominant approaches is to apply energy efficient algorithms to optimize resource usage and energy consumption. Currently, various virtual machine consolidation-based energy efficient algorithms have been proposed to reduce the energy of cloud computing environment. However, most of them are not compared comprehensively under the same scenario, and their performance is not evaluated with the same experimental settings. This makes users hard to select the appropriate algorithm for their objectives. To provide insights for existing energy efficient algorithms and help researchers to choose the most suitable algorithm, in this paper, we compare several state-of-the-art energy efficient algorithms in depth from multiple perspectives, including architecture, modelling and metrics. In addition, we also implement and evaluate these algorithms with the same experimental settings in CloudSim toolkit. The experimental results show the performance comparison of these algorithms with comprehensive results. Finally, detailed discussions of these algorithms are provided.
Qiheng Zhou, Minxian Xu, Sukhpal Singh, Chengxi Gao, Wenhong Tian, Cheng-Zhong Xu 0001, Rajkumar Buyya
CCGRID2
2020 A Reinforcement Learning Based Approach to Identify Resource Bottlenecks for Multiple Services Interactions in Cloud Computing Environments
Lingxiao Xu, Minxian Xu, Richard Semmes, Hong Mu, Shuangquan Gui, Wenhong Tian, Kui Wu 0001, Rajkumar Buyya
CollaborateCom (2)2
2020 Collaborate Edge and Cloud Computing With Distributed Deep Learning for Smart City Internet of Things
abstract
City Internet-of-Things (IoT) applications are becoming increasingly complicated and thus require large amounts of computational resources and strict latency requirements. Mobile cloud computing (MCC) is an effective way to alleviate the limitation of computation capacity by offloading complex tasks from mobile devices (MDs) to central clouds. Besides, mobile-edge computing (MEC) is a promising technology to reduce latency during data transmission and save energy by providing services in a timely manner. However, it is still difficult to solve the task offloading challenges in heterogeneous cloud computing environments, where edge clouds and central clouds work collaboratively to satisfy the requirements of city IoT applications. In this article, we consider the heterogeneity of edge and central cloud servers in the offloading destination selection. To jointly optimize the system utility and the bandwidth allocation for each MD, we establish a hybrid offloading model, including the collaboration of MCC and MEC. A distributed deep learning-driven task offloading (DDTO) algorithm is proposed to generate near-optimal offloading decisions over the MDs, edge cloud server, and central cloud server. Experimental results demonstrate the accuracy of the DDTO algorithm, which can effectively and efficiently generate near-optimal offloading decisions in the edge and cloud computing environments. Furthermore, it achieves high performance and greatly reduces the computational complexity when compared with other offloading schemes that neglect the collaboration of heterogeneous clouds. More precisely, the DDTO scheme can improve computational performance by 63%, compared with the local-only scheme.
Huaming Wu, Ziru Zhang, Chang Guan, Katinka Wolter, Minxian Xu
IEEE Internet Things J.5
2020 Managing renewable energy and carbon footprint in multi-cloud computing environments
Minxian Xu, Rajkumar Buyya
J. Parallel Distributed Comput.1
2019 Optimized Renewable Energy Use in Green Cloud Data Centers
Minxian Xu, Adel Nadjaran Toosi, Behrooz Bahrani, Reza Razzaghi, Martin Singh
ICSOC1
2019 BrownoutCon: A software system based on brownout and containers for energy-efficient cloud computing
Minxian Xu, Rajkumar Buyya
J. Syst. Softw.1
2019 iBrownout: An Integrated Approach for Managing Energy and Brownout in Container-Based Clouds
abstract
Energy consumption of Cloud data centers has been a major concern of many researchers, and one of the reasons for huge energy consumption of Clouds lies in the inefficient utilization of computing resources. Besides energy consumption, another challenge of data centers is the unexpected loads, which leads to the overloads and performance degradation. Compared with VM consolidation and Dynamic Voltage Frequency Scaling that cannot function well when the whole data center is overloaded, brownout has shown to be a promising technique to handle both overloads and energy consumption through dynamically deactivating application optional components, which are also identified as containers/microservices. In this work, we propose an integrated approach to manage energy consumption and brownout in container-based cloud data centers. We also evaluate our proposed scheduling policies with real traces in a prototype system. The results show that our approach reduces about 40, 20, and 10 percent energy than the approach without power-saving techniques, brownout-overbooking approach and auto-scaling approach, respectively, while ensuring Quality of Service.
Minxian Xu, Adel Nadjaran Toosi, Rajkumar Buyya
IEEE Trans. Sustain. Comput.1
2017 Energy Efficient Scheduling of Application Components via Brownout and Approximate Markov Decision Process
Minxian Xu, Rajkumar Buyya
ICSOC1
2017 A survey on load balancing algorithms for virtual machines placement in cloud computing
abstract
Summary The emergence of cloud computing based on virtualization technologies brings huge opportunities to host virtual resource at low cost without the need of owning any infrastructure. Virtualization technologies enable users to acquire, configure, and be charged on pay‐per‐use basis. However, cloud data centers mostly comprise heterogeneous commodity servers hosting multiple virtual machines (VMs) with potential various specifications and fluctuating resource usages, which may cause imbalanced resource utilization within servers that may lead to performance degradation and service level agreements violations. So as to achieve efficient scheduling, these challenges should be addressed and solved by using load balancing strategies, which have been proved to be nondeterministic polynomial time (NP)‐hard problem. From multiple perspectives, this work identifies the challenges and analyzes existing algorithms for allocating VMs to hosts in infrastructure clouds, especially focuses on load balancing. A detailed classification targeting load balancing algorithms for VM placement in cloud data centers is investigated, and the surveyed algorithms are classified according to the classification. The goal of this paper is to provide a comprehensive and comparative understanding of existing literature and aid researchers by providing an insight for potential future enhancements.
Minxian Xu, Wenhong Tian, Rajkumar Buyya
Concurr. Comput. Pract. Exp.1
2016 Energy Efficient Scheduling of Cloud Application Components with Brownout
abstract
It is common for cloud data centers meeting unexpected loads like request bursts, which may lead to overloaded situation and performance degradation. Dynamic Voltage Frequency Scaling and VM consolidation have been proved effective to manage overloads. However, they cannot function when the whole data center is overloaded. Brownout provides a promising direction to avoid overloads through configuring applications to temporarily degrade user experience. Additionally, brownout can also be applied to reduce data center energy consumption. As a complementary option for Dynamic Voltage Frequency Scaling and VM consolidation, our combined brownout approach reduces energy consumption through selectively and dynamically deactivating application optional components, which can also be applied to self-contained microservices. The results show that our approach can save more than 20 percent energy consumption and there are trade-offs between energy saving and discount offered to users.
Minxian Xu, Amir Vahid Dastjerdi, Rajkumar Buyya
IEEE Trans. Sustain. Comput.1
2015 A Toolkit for Modeling and Simulation of Real-Time Virtual Machine Allocation in a Cloud Data Center
abstract
Resource scheduling in infrastructure as a service (IaaS) is one of the keys for large-scale Cloud applications. Extensive research on all issues in real environment is extremely difficult because it requires developers to consider network infrastructure and the environment, which may be beyond the control. In addition, the network conditions cannot be predicted or controlled. Therefore, performance evaluation of workload models and Cloud provisioning algorithms in a repeatable manner under different configurations and requirements is difficult. There is still lack of tools that enable developers to compare different resource scheduling algorithms in IaaS regarding both computing servers and user workloads. To fill this gap in tools for evaluation and modeling of Cloud environments and applications, we propose CloudSched. CloudSched can help developers identify and explore appropriate solutions considering different resource scheduling algorithms. Unlike traditional scheduling algorithms considering only one factor such as CPU, which can cause hotspots or bottlenecks in many cases, CloudSched treats multidimensional resource such as CPU, memory and network bandwidth integrated for both physical machines and virtual machines (VMs) for different scheduling objectives (algorithms). In this paper, two existing simulation systems at application level for Cloud computing are studied, a novel lightweight simulation system is proposed for real-time VM scheduling in Cloud data centers, and results by applying the proposed simulation system are analyzed and discussed.
Wenhong Tian, Yong Zhao 0009, Minxian Xu, Yuanliang Zhong, Xiashuang Sun
IEEE Trans Autom. Sci. Eng.3
2014 Prepartition: A new paradigm for the load balance of virtual machine reservations in data centers
abstract
It is significant to apply load-balancing strategy to improve the performance and reliability of resource in data centers. One of the challenging scheduling problems in Cloud data centers is to take the allocation and migration of reconfigurable virtual machines (VMs) as well as the integrated features of hosting physical machines (PMs) into consideration. In the reservation model, workload of data centers has fixed process interval characteristics. In general, load-balance scheduling is NP-hard problem as proved in many open literatures. Traditionally, for offline load balance without migration, one of the best approaches is LPT (Longest Process Time first), which is well known to have approximation ratio 4/3. With virtualization, reactive (post) migration of VMs after allocation is one popular way for load balance and traffic consolidation. However, reactive migration has difficulty to reach predefined load balance objectives, and may cause interruption and instability of service and other associated costs. In view of this, we propose a new paradigm-Prepartition: it proactively sets process-time bound for each request on each PM and prepares in advance to migrate VMs to achieve the predefined balance goal. Prepartition can reduce process time by preparing VM migration in advance and therefore reduce instability and achieve better load balance as desired. Trace-driven and synthetic simulation results show that Prepartition has 10%-20% better performance than the well known load balancing algorithms with regard to average CPU utilization, makespan as well as capacity makespan.
Wenhong Tian, Minxian Xu, Yong Zhao 0009
ICC2