Shashikant Ilager

dblp:192/5685 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0003-1178-6582ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 9 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ARKV: Adaptive and Resource-Efficient KV Cache Management Under Limited Memory Budget for Long-Context Inference in LLMs
Jianlong Lei, Shashikant Ilager
CCGrid2
2026 Characterizing LLM Inference Energy-Performance Tradeoffs Across Workloads and GPU Scaling
abstract
LLM inference exhibits substantial variability across queries and execution phases, yet inference configurations are often applied uniformly. We present a measurement-driven characterization of workload heterogeneity and energy-performance behavior of LLM inference under GPU dynamic voltage and frequency scaling (DVFS). We evaluate five decoder-only LLMs (1B-32B parameters) across four NLP benchmarks using a controlled offline setup. We show that lightweight semantic features predict inference difficulty better than input length, with 44.5% of queries achieving comparable quality across model sizes. At the hardware level, the decode phase dominates inference time (77-91%) and is largely insensitive to GPU frequency. Consequently, reducing GPU frequency from 2842 MHz to 180 MHz achieves an average of 42% energy savings with only a 1-6% latency increase. We further provide a use case with an upper-bound analysis of the potential benefits of combining workload-aware model selection with phase-aware DVFS, motivating future energy-efficient LLM inference systems.
Paul Joe Maliakel, Shashikant Ilager, Ivona Brandic
CCGrid2
2026 A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
abstract
Edge computing provides resources for IoT workloads at the network edge. Monitoring systems are vital for efficiently managing resources and application workloads by collecting, storing, and providing relevant information about the state of the resources. However, traditional monitoring systems have a centralized architecture for both data plane and control plane, which increases latency, creates a failure bottleneck, and faces challenges in providing quick and trustworthy data in volatile edge environments, especially where infrastructures are often built upon failure-prone, unsophisticated computing and network resources. Thus, we propose DEMon, a decentralized, self-adaptive monitoring system for edge. DEMon leverages the stochastic gossip communication protocol at its core. It develops efficient protocols for information dissemination, communication, and retrieval, avoiding a single point of failure and ensuring fast and trustworthy data access. Its decentralized control enables self-adaptive management of monitoring parameters, addressing the tradeoffs between the quality of service of monitoring and resource consumption. We implement the proposed system as a lightweight and portable container-based system and evaluate it through experiments. We also present a use case demonstrating its feasibility. The results show that DEMon efficiently disseminates and retrieves the monitoring information, addressing the challenges of edge monitoring.
Shashikant Ilager, Jakob Fahringer, Alessandro Tundo, Ivona Brandic
ACM Trans. Auton. Adapt. Syst.1
2025 Green-Code: Learning to Optimize Energy Efficiency in Llm-Based Code Generation
abstract
Large Language Models (LLMs) are becoming integral to daily life, showcasing their vast potential across various Natural Language Processing (NLP) tasks. Beyond NLP, LLMs are increasingly used in software development tasks, such as code completion, modification, bug fixing, and code translation. Software engineers widely use tools like GitHub Copilot and Amazon Q, streamlining workflows and automating tasks with high accuracy. While the resource and energy intensity of LLM training is often highlighted, inference can be even more resourceintensive over time, as it's a continuous process with a high number of invocations. Therefore, developing resource-efficient alternatives for LLM inference is crucial for sustainability. This work proposes GREEN-CODE, a framework for energy-aware code generation in LLMs. GREEN-CODE performs dynamic early exit during LLM inference. We train a Reinforcement Learning (RL) agent that learns to balance the trade-offs between accuracy, latency, and energy consumption. Our approach is evaluated on two open-source LLMs, Llama 3.2 3B and OPT 2.7 B, using the JavaCorpus and PY150 datasets. Results show that our method reduces the energy consumption between 2350 % on average for code generation tasks without significantly affecting accuracy.
Shashikant Ilager, Lukas Florian Briem, Ivona Brandic
CCGrid1
2025 A Framework for Carbon-Aware Real-Time Workload Management in Clouds Using Renewables-Driven Cores
abstract
Cloud platforms commonly exploit workload temporal flexibility to reduce their carbon emissions. They suspend/resume workload execution for when and where the energy is greenest. However, increasingly prevalent delay-intolerant real-time workloads challenge this approach. To this end, we present a framework to harvest green renewable energy for real-time workloads in cloud systems. We useRenewables-driven coresin servers to dynamically switch CPU cores between real-time and low-power profiles, matching renewable energy availability. We then develop a VM Execution Model to guarantee running VMs are allocated with cores in thereal-time power profile. If such cores are insufficient, we conduct criticality-aware VM evictions as needed. Furthermore, we develop a VM Packing Algorithm to utilize available cores across the servers. We introduce theGreen Coresconcept in our algorithm to convert renewable energy usage into a server inventory attribute. Based on this, we jointly optimize for renewable energy utilization and reduction of VM eviction incidents. We implement a prototype of our framework in OpenStack asopenstack-gc. Using an experimentalopenstack-gccloud and a large-scale simulation testbed, we expose our framework to VMs running RTEval, a real-time evaluation program, and a 14-day Azure VM arrival trace. Our results show: i) a 6.52× reduction in coefficient of variation of real-time latency over an existing workload temporal flexibility-based solution, and ii) a joint 79.64% reduction in eviction incidents with a 34.83% increase in energy harvest over the state-of-the-art packing algorithms. We open sourceopenstack-gcat https://github.com/tharindu-b-hewage/openstack-gc.
Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read, Rajkumar Buyya
IEEE Trans. Computers2
2025 FRESCO: Fast and Reliable Edge Offloading With Reputation-Based Hybrid Smart Contracts
abstract
Mobile devices offload latency-sensitive application tasks to edge servers to satisfy applications' Quality of Service (QoS) deadlines. Consequently, ensuring reliable offloading without QoS violations is challenging in distributed and unreliable edge environments with diverse resource and reliability levels. We propose FRESCO, a fast and reliable edge offloading framework that utilizes a blockchain-based reputation system, which enhances the reliability of offloading in the distributed edge. The distributed reputation system tracks the historical performance of edge servers, while blockchain through a consensus mechanism ensures that sensitive reputation information is secured against tampering. However, blockchain consensus typically has high latency, and therefore we employ a Hybrid Smart Contract (HSC) as areputation state managerthat automatically computes and stores reputation securely on-chain (i.e., on the blockchain) while allowing fast offloading decisions off-chain (i.e., outside of blockchain). Theoffloading decision engineuses a reputation score from HSC to derive fast offloading decisions, which are based on Satisfiability Modulo Theory (SMT). The SMT can formally guarantee a feasible solution that is valuable for latency-sensitive applications that require high reliability. With a combination of an on-chain HSC reputation state manager and an off-chain SMT decision engine, FRESCO offloads tasks to reliable servers without being hindered by blockchain consensus. In our experiment, FRESCO reduces response time by up to 7.86 times and saves energy by up to 5.4% compared to all baselines while minimizing QoS violations to 0.4% and achieving an average decision time of just 5.05 milliseconds.
Josip Zilic, Vincenzo De Maio, Shashikant Ilager, Ivona Brandic
IEEE Trans. Serv. Comput.3
2024 iContinuum: An Emulation Toolkit for Intent-Based Computing Across the Edge-to-Cloud Continuum
abstract
The Internet of Things (IoT) has led to a surge in smart devices, generating vast volumes of data. Cloud computing offers scalability but does not suffice for many real-time and privacy-sensitive IoT applications. This limitation has prompted a blend of both edge and cloud resources, creating the need for seamless integration, known as the “compute continuum“. Testing applications and resource management techniques within this continuum is vital but can be very complex. Simulation and emulation are preferred methods, with emulation providing more accurate representations of real-world environments. In this paper, we introduce iContinuum, a novel emulation toolkit facilitating an intent-based platform for edge-to-cloud testing and experimentation. Leveraging Software-Defined Networking (SDN) and containerization, iContinuum enables experimentation and performance evaluation while aligning application requirements with actual performance. We present our detailed architecture, implementation, and evaluation of iContinuum, showcasing how our proposed toolkit bridges the gap between simulation and real-world deployment within compute continuum environments, and further demonstrate the effectiveness of Intent-Based Scheduling through a specific use case.
Negin Akbari, Adel Nadjaran Toosi, John C. Grundy, Hourieh Khalajzadeh, Mohammad Sadegh Aslanpour, Shashikant Ilager
CLOUD6
2024 Machine Learning in Energy and Thermal-aware Resource Management of Cloud Data Centers: A Taxonomy and Future Directions
abstract
Cloud data centres (CDCs) are the backbone infrastructures of modern digital society, but they also consume huge amounts of energy and generate heat.To manage CDC resources efficiently, we must consider the complex interactions between diverse workloads and data centre components.However, most existing resource management systems rely on simple and static rules that fail to capture these complex interactions.Therefore, we require new data-driven Machine learning-based resource management approaches that can efficiently capture the interdependencies between parameters and guide resource management systems.This review describes the in-depth analysis of the existing resource management approaches in CDCs for energy and thermal efficiency.It mainly focuses on learning-based resource management systems in data centres and also identifies the need for integrated computing and cooling systems management.A taxonomy on energy and thermal efficient resource management in data centres is proposed.Furthermore, based on this taxonomy, existing resource management approaches from server level, data centre level, and cooling system level are discussed.Finally, key future research directions for sustainable Cloud computing services are proposed.
Shashikant Ilager, Rajkumar Buyya
FedCSIS1
2024 Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
abstract
HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.
Xiaoyu Chu, Daniel Hofstätter, Shashikant Ilager, Sacheendra Talluri, Duncan Kampert, Damian Podareanu, Dmitry Duplyakin, Ivona Brandic, Alexandru Iosup
ICPADS3
2024 ABBA-VSM: Time Series Classification Using Symbolic Representation on the Edge
Meerzhan Kanatbekova, Shashikant Ilager, Ivona Brandic
ICSOC (1)2
2024 CloudSim express: A novel framework for rapid low code simulation of cloud computing environments
abstract
Abstract Cloud computing environment simulators enable cost‐effective experimentation of novel infrastructure designs and management approaches by avoiding significant costs incurred from repetitive deployments in real Cloud platforms. However, widely used Cloud environment simulators compromise on usability due to complexities in design and configuration, along with the added overhead of programming language expertise. Existing approaches attempting to reduce this overhead, such as script‐based simulators and graphical user interface (GUI) based simulators, often compromise on the extensibility of the simulator. Simulator extensibility allows for customization at a fine‐grained level, thus reducing it significantly affects flexibility in creating simulations. To address these challenges, we propose an architectural framework to enable human‐readable script‐based simulations in existing Cloud environment simulators while minimizing the impact on simulator extensibility. We implement the proposed framework for the widely used Cloud environment simulator, the CloudSim toolkit, and compare it against state‐of‐the‐art baselines using a practical use case. The resulting framework, called CloudSim Express, achieves extensible simulations while surpassing baselines with over a % reduction in code complexity and an 89.42% reduction in lines of code.
Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read, Rajkumar Buyya
Softw. Pract. Exp.2
2023 DEMOTS: A Decentralized Task Scheduling Algorithm for Micro-Clouds with Dynamic Power-Budgets
abstract
The Internet of Things (IoT) driven latency-critical applications are deployed on lightweight Micro-Clouds at the network's edge. Renting physical space from geographically distributed colocation datacenters connected via a Wide Area Network (WAN) is a cost-effective way of deploying Micro-Clouds, despite WANs' dynamic communication latency from traffic congestion. However, this deployment approach can limit Micro-Clouds to operate within a soft power budget, as colocation datacenter providers utilize it to add more servers and lower capital costs through oversubscribing power infrastructure. As a result, Micro-Clouds use extreme energy reduction measures like power capping and task throttling to address power overdraw events, where power consumption exceeds soft power budget limits, which reduces the performance of latency-critical applications. We propose a solution where a dynamic power budget can be achieved by adding renewable energy sources to the existing soft power budget without upgrading power delivery systems. To take advantage of this, we propose a dynamic, decentralized task-scheduling algorithm called DEMOTS. DEMOTS effectively utilizes the available dynamic power budget in a WAN with varying degrees of network traffic congestion, thereby avoiding the need for extreme energy reduction measures. We implement DEMOTS on a simulation test-bed. Compared to state-of-the-art baseline using MCOP for decentralized task-scheduling in Micro-Clouds, DEMOTS reduces Power Overdraw Impact up to 19%, Task Latency Increase Impact up to 47 %, and Task Schedule Time Impact up to 49%.
Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read, Patricia Arroba, Rajkumar Buyya
CLOUD2
2023 SymED: Adaptive and Online Symbolic Representation of Data on the Edge
Daniel Hofstätter, Shashikant Ilager, Ivan Lujic, Ivona Brandic
Euro-Par2
2023 An Energy-Aware Approach to Design Self-Adaptive AI-based Applications on the Edge
abstract
The advent of edge devices dedicated to machine learning tasks enabled the execution of AI-based applications that efficiently process and classify the data acquired by the resource-constrained devices populating the Internet of Things. The proliferation of such applications (e.g., critical monitoring in smart cities) demands new strategies to make these systems also sustainable from an energetic point of view. In this paper, we present an energy-aware approach for the design and deployment of self-adaptive AI-based applications that can balance application objectives (e.g., accuracy in object detection and frames processing rate) with energy consumption. We address the problem of determining the set of configurations that can be used to self-adapt the system with a meta-heuristic search procedure that only needs a small number of empirical samples. The final set of configurations are selected using weighted gray relational analysis, and mapped to the operation modes of the self-adaptive application. We validate our approach on an AI-based application for pedestrian detection. Results show that our self-adaptive application can outperform non-adaptive baseline configurations by saving up to 81% of energy while loosing only between 2% and 6 % in accuracy.
Alessandro Tundo, Marco Mobilio, Shashikant Ilager, Ivona Brandic, Ezio Bartocci, Leonardo Mariani
ASE3
2022 Workload forecasting and energy state estimation in cloud data centres: ML-centric approach
Tahseen Khan, Wenhong Tian, Shashikant Ilager, Rajkumar Buyya
Future Gener. Comput. Syst.3
2022 Machine learning (ML)-centric resource management in cloud computing: A review and future directions
Tahseen Khan, Wenhong Tian, Shashikant Ilager, Mingming Gong, Rajkumar Buyya
J. Netw. Comput. Appl.4
2022 Dynamic Scheduling for Stochastic Edge-Cloud Computing Environments Using A3C Learning and Residual Recurrent Neural Networks
abstract
The ubiquitous adoption of Internet-of-Things (IoT) based applications has resulted in the emergence of the Fog computing paradigm, which allows seamlessly harnessing both mobile-edge and cloud resources. Efficient scheduling of application tasks in such environments is challenging due to constrained resource capabilities, mobility factors in IoT, resource heterogeneity, network hierarchy, and stochastic behaviors. Existing heuristics and Reinforcement Learning based approaches lack generalizability and quick adaptability, thus failing to tackle this problem optimally. They are also unable to utilize the temporal workload patterns and are suitable only for centralized setups. However, asynchronous-advantage-actor-critic (A3C) learning is known to quickly adapt to dynamic scenarios with less data and residual recurrent neural network (R2N2) to quickly update model parameters. Thus, we propose an A3C based real-time scheduler for stochastic Edge-Cloud environments allowing decentralized learning, concurrently across multiple agents. We use the R2N2 architecture to capture a large number of host and task parameters together with temporal patterns to provide efficient scheduling decisions. The proposed model is adaptive and able to tune different hyper-parameters based on the application requirements. We explicate our choice of hyper-parameters through sensitivity analysis. The experiments conducted on real-world data set show a significant improvement in terms of energy consumption, response time, Service-Level-Agreement and running cost by 14.4, 7.74, 31.9, and 4.64 percent, respectively when compared to the state-of-the-art algorithms.
Shreshth Tuli, Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya
IEEE Trans. Mob. Comput.2
2022 CoScal: Multifaceted Scaling of Microservices With Reinforcement Learning
abstract
The emerging trend towards moving from monolithic applications to microservices has raised new performance challenges in cloud computing environments. Compared with traditional monolithic applications, the microservices are lightweight, fine-grained, and must be executed in a shorter time. Efficient scaling approaches are required to ensure microservices’ system performance under diverse workloads with strict Quality of Service (QoS) requirements and optimize resource provisioning. To solve this problem, we investigate the trade-offs between the dominant scaling techniques, including horizontal scaling, vertical scaling, and brownout in terms of execution cost and response time. We first present a prediction algorithm based on gradient recurrent units to accurately predict workloads assisting in scaling to achieve efficient scaling. Further, we propose a multi-faceted scaling approach using reinforcement learning called CoScal to learn the scaling techniques efficiently. The proposed CoScal approach takes full advantage of data-driven decisions and improves the system performance in terms of high communication cost and delay. We validate our proposed solution by implementing a containerized microservice prototype system and evaluated with two microservice applications. The extensive experiments demonstrate that CoScal reduces response time by 19%-29% and decreases the connection time of services by 16% when compared with the state-of-the-art scaling techniques for Sock Shop application. CoScal can also improve the number of successful transactions with 6%-10% for Stan’s Robot Shop application.
Minxian Xu, Chenghao Song, Shashikant Ilager, Sukhpal Singh, Juanjuan Zhao 0001, Kejiang Ye, Cheng-Zhong Xu 0001
IEEE Trans. Netw. Serv. Manag.3
2021 Special issue: Elastic computing from edge to the cloud environments
abstract
We are pleased to present a special issue that focuses on state-of-the-art research on Elastic Computing from Edge to the Cloud Environments. Today, a huge amount of data is being generated by the Internet of Things (IoT) devices such as smartphones, sensors, cameras, cars, and robots.1 In order to process the generated data, there exist Big Data platforms (such as Hadoop and Spark). Conventionally, they are deployed in centralized Data Centers, which, however, fall short of addressing time-critical requirements of the applications due to high latency between the Edge, where the data are generated and the Data Centers where they are processed.2 The emerging Edge and Fog computing paradigms promise to solve this problem by seamlessly integrating hardware and software resources across multiple computing tiers, from the Edge to the Data Center/Cloud. Since computing resources at the Edge may be power and capacity constrained, it is necessary to invent new lightweight platforms and techniques that seamlessly interact, sense, execute and produce results with very low latency, while at the same time address other high-level requirements of applications, such as security and privacy. Regarding these problems, there are many challenges that must be addressed with the invention of new architectures, methods, algorithms, and solutions. This special issue features six papers covering a range of topics including IoT application deployment frameworks in edge-cloud environments, cost optimization models, and lightweight virtualization model. The first paper in this special issue titled “Edge-adaptable serverless acceleration for machine learning Internet of Things applications”3 presents STOIC (serverless teleoperable hybrid cloud), an IoT application deployment and offloading system that extends the serverless model. The authors have developed a dynamic feedback control mechanism to precisely predict latency and dispatch workloads uniformly across edge and cloud systems using a distributed serverless framework. STOIC leverages hardware acceleration (e.g., GPU resources) for serverless function execution when available from the underlying cloud system. Finally, it configures the system to overcome deployment variability associated with public clouds. The system is evaluated using real-world machine learning applications and multitier IoT deployments (edge and cloud) and shown that it reduces overall execution time and achieves placement accuracy in the range of 92%–97%. The second paper in this special issue titled “Server Configuration Optimisation in Mobile Edge Computing: A Cost-Performance Tradeoff Perspective”4 studies the problem of server configuration optimization in Mobile edge computing environments. The authors use M/M/m queuing models and establish the performance and cost models for the system. The article considers cost-constrained performance optimization, and performance constrained cost optimization based on multiple numerical algorithms. The numerical simulation-based experiments have shown this approach is able to balance the trade-off between investment cost and service quality. The third paper in this special issue titled “EFFORT: Energy-efficient framework for offload. Communication in mobile cloud computing”5 proposes a task offloading mechanism from mobile to the remote cloud. The author's solution aims to solve the energy consumption of communication-intensive applications from mobile devices such as smartphones. The experimental evaluation is done by implementing a demonstration application in Android mobile OS. The results have shown that the proposed solution reduces the energy consumption of smartphones while executing applications and simultaneously reducing the communication cost. The fourth paper in this special issue titled “Human Microservices: A framework for turning humans into service providers”6 presents a framework facilitating the deployment of Application Programming Interface on companion devices (smartphones and IoT devices). The author's framework proposes a new approach aiming to integrate humans in the IoT loop and facilitate computation units' deployment in the devices that are closer to users, instead of remote clouds. The solution leverages existing standards for the rapid development of applications and improves software quality. Through a detailed case study on monitoring people's activity, the authors developed a software application following framework specification and demonstrated the feasibility of their proposed solution. The fifth paper in this special issue titled “Function delivery network: Extending serverless computing for heterogeneous platforms”7 introduces the extension to Function-as-a-Service model. The authors consider heterogeneous clusters and support heterogeneous functions through a network of distributed heterogeneous target platforms called Function Delivery Network (FDN). The article proposes Function-Delivery-as-a-Service, a service model to deliver the function to the required to be targeted platform. It also ensures that Service Level Objective (SLO) energy efficiency requirements are met when scheduling functions. The new benchmarking system FDNInspecto is implemented which benchmarks the different distributed target platforms. The results have found that functions scheduling on edge reduce energy consumption by 17× without violating the SLO requirements when compared to a high-end target platform. Finally, the last paper of this special issue titled “A Lightweight Virtualisation Model to Enable Edge Computing in Deeply Embedded Systems”8 builds a virtualization model for resource-constrained embedded devices. The existing containers cannot be used in many deeply embedded systems (DES) due to an underlying operating system's resource requirements including storage, memory, and processing power. To mitigate this issue, the authors present the Hellfire hypervisor, a lightweight virtualization model that enables separation and improves the security of IoT applications on DES. The Hellfire simplifies essential components of virtualization including compute, storage, and I/O among others. The experimental results conducted with the coremark benchmark show Hellfire has a small footprint of 23 KB while keeping a low average virtualization overhead of 0.62% for multiple virtual machines execution. To summarize, the articles in this special issue provide new definitions, architectural principles, systems, and approaches to solve the challenges posed by emerging distributed systems such as edge and cloud computing, mainly driven by IoT workloads. They address important problems including lightweight platforms, energy efficiency, reliability and easy to integrate frameworks. We hope that you will find this special issue truly useful and enjoyable. We sincerely thank the editor-in-chief for helping us to organize this special issue. We thank the editorial office staffs for their continuous support. We are also thankful to all the authors for submitting their research works. More importantly, thanks to the reviewers for their valuable contributions in providing thoughtful comments and improving the quality of articles.
Shashikant Ilager, Vlado Stankovski, Shrideep Pallickara, Rajkumar Buyya
Softw. Pract. Exp.1
2021 Thermal Prediction for Efficient Energy Management of Clouds Using Machine Learning
abstract
Thermal management in the hyper-scale cloud data centers is a critical problem. Increased host temperature creates hotspots which significantly increases cooling cost and affects reliability. Accurate prediction of host temperature is crucial for managing the resources effectively. Temperature estimation is a non-trivial problem due to thermal variations in the data center. Existing solutions for temperature estimation are inefficient due to their computational complexity and lack of accurate prediction. However, data-driven machine learning methods for temperature prediction is a promising approach. In this regard, we collect and study data from a private cloud and show the presence of thermal variations. We investigate several machine learning models to accurately predict the host temperature. Specifically, we propose a gradient boosting machine learning model for temperature prediction. The experiment results show that our model accurately predicts the temperature with the average RMSE value of 0.05 or an average prediction error of 2.38 °C, which is 6 °C less as compared to an existing theoretical model. In addition, we propose a dynamic scheduling algorithm to minimize the peak temperature of hosts. The results show that our algorithm reduces the peak temperature by 6.5 °C and consumes 34.5 percent less energy as compared to the baseline algorithm.
Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya
IEEE Trans. Parallel Distributed Syst.1
2020 A Data-Driven Frequency Scaling Approach for Deadline-aware Energy Efficient Scheduling on Graphics Processing Units (GPUs)
abstract
Modern computing paradigms, such as cloud computing, are increasingly adopting GPUs to boost their computing capabilities primarily due to the heterogeneous nature of AI/ML/deep learning workloads. However, the energy consumption of GPUs is a critical problem. Dynamic Voltage Frequency Scaling (DVFS) is a widely used technique to reduce the dynamic power of GPUs. Yet, configuring the optimal clock frequency for essential performance requirements is a non-trivial task due to the complex nonlinear relationship between the application's runtime performance characteristics, energy, and execution time. It becomes more challenging when different applications behave distinctively with similar clock settings. Simple analytical solutions and standard GPU frequency scaling heuristics fail to capture these intricacies and scale the frequencies appropriately. In this regard, we propose a data-driven frequency scaling technique by predicting the power and execution time of a given application over different clock settings. We collect the data from application profiling and train the models to predict the outcome accurately. The proposed solution is generic and can be easily extended to different kinds of workloads and GPU architectures. Furthermore, using this frequency scaling by prediction models, we present a deadline-aware application scheduling algorithm to reduce energy consumption while simultaneously meeting their deadlines. We conduct real extensive experiments on NVIDIA GPUs using several benchmark applications. The experiment results have shown that our prediction models have high accuracy with the average RMSE values of 0.38 and 0.05 for energy and time prediction, respectively. Also, the scheduling algorithm consumes 15.07% less energy as compared to the baseline policies.
Shashikant Ilager, Rajeev Muralidhar, Kotagiri Ramamohanarao, Rajkumar Buyya
CCGRID1
2019 ETAS: Energy and thermal-aware dynamic virtual machine consolidation in cloud data center with proactive hotspot mitigation
abstract
Summary Data centers consume an enormous amount of energy to meet the ever‐increasing demand for cloud resources. Computing and Cooling are the two main subsystems that largely contribute to energy consumption in a data center. Dynamic Virtual Machine (VM) consolidation is a widely adopted technique to reduce the energy consumption of computing systems. However, aggressive consolidation leads to the creation of local hotspots that has adverse effects on energy consumption and reliability of the system. These issues can be addressed through efficient and thermal‐aware consolidation methods. We propose an Energy and Thermal‐Aware Scheduling (ETAS) algorithm that dynamically consolidates VMs to minimize the overall energy consumption while proactively preventing hotspots. ETAS is designed to address the trade‐off between time and the cost savings and it can be tuned based on the requirement. We perform extensive experiments by using the real‐world traces with precise power and thermal models. The experimental results and empirical studies demonstrate that ETAS outperforms other state‐of‐the‐art algorithms by reducing overall energy without any hotspot creation.
Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya
Concurr. Comput. Pract. Exp.1