VLDB 2026 Research / reviewers in the wild / expert
Rodrigo N. Calheiros
dblp:78/6512 · also Rodrigo Neves Calheiros
· DBLP profile ↗
78ranked-venue papers
14as first author
17since 2021 · last 2026
0000-0001-7435-2445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 7 first-author · 7 since 2021Computer networks · 14 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 12 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CroSatFL: Energy-Efficient Federated Learning with Cross-Aggregation for Satellite Edge Computing
Bahman Javadi, Rodrigo N. Calheiros, David Boland, Philip Leong |
CCGrid | 3 |
| 2025 | Minimizing the Number of Drone-Based Repeaters in Deploying Quantum NetworksabstractDeploying quantum networks involves the placement of quantum repeaters, which generate entangled qubits for quantum computers to perform quantum teleportation. Besides placing repeaters on ground devices, there have been efforts towards the use of satellites or drones as repeaters. This paper considers the scenario of using drones due to their cost-efficiency and support for flexible network formation. While several issues in this scenario, such as enhancing network connectivity, have been studied, this paper addresses a new problem: how to minimize the number of drones placed in the air to cover all quantum computers on the ground and make the entire quantum network connected. Given that this problem is NP-hard, we propose a suboptimal but polynomial-time approach for it. Our approach consists of two stages. Stage 1 aims to minimize the number of drones needed to cover the ground computers, and Stage 2 aims to minimize the number of drones needed for connecting the drones returned by Stage 1. Both stages use algorithms of no more than quadratic time complexity. We conduct experiments to determine the best algorithm for the setting of quantum networks from several high-performing existing algorithms and design optimization techniques for existing algorithms. Our experiments confirm that our approach reduces the number of drones placed effectively. Romtham Sripotchanart, Weisheng Si, Rodrigo N. Calheiros, Tie Qiu 0001 |
ICC | 3 |
| 2025 | Federated Learning with Reliability-Aware Workload Allocation in Distributed Edge ComputingabstractFederated Learning (FL) enables collaborative model training across various distributed devices without sharing raw data. Client failures, variable energy availability, and outages of edge servers contribute to unreliable training participation, incomplete model updates, and failures at the system level during the aggregation process. In this study, we introduce a reliability-aware workload allocation FL framework (FedRAW) aimed at improving system reliability in failure-prone edge computing systems. Our approach dynamically modifies client workloads based on their failure history and integrates a lightweight backup mechanism to maintain aggregation continuity during edge server failures by backup servers handling. Additionally, we employ Bayesian optimization to fine-tune workload parameters, achieving improved energy efficiency. Experimental results reveal that our proposed method improves model accuracy while reducing energy consumption compared to recent federated learning algorithms. Fatemeh Mirhakimi, Bahman Javadi, Rodrigo N. Calheiros, Adel Nadjaran Toosi |
MSWiM | 3 |
| 2025 | Resource-Efficient Multiview Perception: Integrating Semantic Masking with Masked AutoencodersabstractMultiview systems have become a key technology in modern computer vision, offering advanced capabilities in scene understanding and analysis. However, these systems face critical challenges in bandwidth limitations and computational constraints, particularly for resource-limited camera nodes. This paper presents a novel approach for communication-efficient distributed multiview detection and tracking using masked autoencoders (MAEs). We introduce a semantic-guided masking strategy that leverages pre-trained segmentation models and a tunable power function to prioritize informative image regions. This approach, combined with an MAE, reduces communication overhead while preserving essential visual information. We evaluate our method on both virtual and real-world multiview datasets, demonstrating comparable performance in terms of detection and tracking performance metrics compared to state-of-the-art techniques, even at high masking ratios. Our selective masking algorithm outperforms random masking, maintaining higher accuracy and precision as the masking ratio increases. Furthermore, our approach achieves a significant reduction in transmission data volume compared to baseline methods, thereby balancing multiview tracking performance with communication efficiency. Kosta Dakic, Kanchana Thilakarathna, Rodrigo N. Calheiros, Teng Joon Lim |
PerCom | 3 |
| 2024 | DigiNet: Scaling up Provisioning of Network Digital TwinabstractThe pursuit of self-driving networks is increasing pressure on adopting intelligent, edge-based networking services. However, deploying autonomous network models within operational and large-scale infrastructures entails substantial risks that require rigorous verification and validation procedures. In this context, the application of a Network Digital Twin (NDT) is emerging as a viable approach towards intelligent network decision-making based on high-fidelity models built upon digital representations of physical network devices (i.e., Digital Twins). In this paper, we take the first steps towards efficiently provisioning NDT models. To that end, we introduce the Digital Twin Network Provisioning Problem (DigiNet), which encompasses the optimal placement of NDT models and the efficient collection of telemetry data for synchronizing NDT models with their physical counterparts. We theoretically formalize DigiNet as a Mixed-Integer Linear Programming (MILP) model and present a polynomial-time heuristic. Our results show that DigiNet outperforms baseline approaches by up to 10x regarding the number of NDT models provisioned. Marcelo Caggiani Luizelli, Francisco Germano Vogt, Paulo Silas Severo de Souza, Arthur Francisco Lorenzon, Roberto Irajá Tavares da Costa Filho, Fábio D. Rossi, Rodrigo N. Calheiros, Christian Esteve Rothenberg |
NetSoft | 7 |
| 2024 | A two-step linear programming approach for repeater placement in large-scale quantum networksabstractThanks to the applications such as Quantum Key Distribution and Distributed Quantum Computing, the deployment of quantum networks is gaining great momentum. A major component in quantum networks is repeaters, which are essential for reducing the error rate of qubit transmission for long-distance links. However, repeaters are expensive devices, so minimizing the number of repeaters placed in a quantum network while satisfying performance requirements becomes an important problem. Existing solutions typically solve this problem optimally by formulating an Integer Linear Program (ILP). However, the number of variables in their ILPs is O ( n 2 ) , where n is the number of nodes in a network. This incurs infeasible running time when the network scale is large. To overcome this drawback, this paper proposes to solve the repeater placement problem by two steps, with each step using a linear program of a much smaller scale with O ( n ) variables. Although this solution is not optimal, it dramatically reduces the time complexity, making it practical for large-scale networks. Moreover, it constructs networks that have higher node connectivity than those by existing solutions, since it deploys slightly more number of repeaters into networks. Our extensive experiments on both synthetic and real-world network topologies verified our claims. Romtham Sripotchanart, Weisheng Si, Rodrigo N. Calheiros, Qing Cao 0001, Tie Qiu 0001 |
Comput. Networks | 3 |
| 2024 | A neural network framework for optimizing parallel computing in cloud servers
Everton Camargo de Lima, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Arthur Francisco Lorenzon |
J. Syst. Archit. | 4 |
| 2024 | Evaluating machine learning prediction techniques and their impact on proactive resource provisioning for cloud environments
Dionatra F. Kirchoff, Vinícius Meyer, Rodrigo N. Calheiros, César A. F. De Rose |
J. Supercomput. | 3 |
| 2024 | FedOrbit: Energy Efficient Federated Learning for Orbital Edge Computing Using Block Minifloat ArithmeticabstractLow Earth Orbit (LEO) satellite constellations have diverse applications, including earth observation, communication services, navigation, and positioning. These constellations have evolved into a valuable data source; however, their use in a ground station (GS) for analysis via machine learning algorithms presents challenges due to constraints on power consumption, communication bandwidth, and onboard computing capabilities. While the combination of Federated Learning (FL) and Orbital Edge Computing has been employed to address these challenges, its heavy reliance on the GS for model aggregation and edge resource limitations remains a research challenge. This article presents FedOrbit, a novel energy-efficient and decentralised FL method to optimise communication with the GS and reduce power consumption. FedOrbit utilises reinforcement learning for cluster formation, satellite visiting patterns for master satellite selection, and block minifloat arithmetic for power reduction. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets reveals that FedOrbit can maintain high accuracy while significantly reduce communication demand, power consumption and training time in comparison to state-of-the-art FL approaches. The proposed technique can also reduce the training time by 5× compared with the centralised FL approaches. In addition, the utilisation of block minifloat representation as low-precision arithmetic enhanced the energy consumption by 3.5× compared with the single-precision (FP32) format. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland |
IEEE Trans. Serv. Comput. | 4 |
| 2023 | On-Board Federated Learning in Orbital Edge ComputingabstractLow Earth Orbit (LEO) satellite constellations are used for a wide range of applications including earth observation, communication services, navigation, and positioning. They have emerged as a new source of data but transferring this data to a ground station (GS) for analysis and machine learning requires extensive bandwidth and incurs high latency. Limited battery capacity, communication and computing capabilities are other factors affecting the training process. Federated Learning (FL) is being used to address these challenges, although it heavily relies on the GS for model aggregation. In this paper, we consider Orbital Edge Computing (OEC) as an architecture for LEO satellite constellations and propose an on-board Federated Learning to reduce communication with the GS. We present a novel decentralised FL algorithm, called FedOrbit, based on reinforcement learning cluster formation and satellite visiting patterns to utilise intra and inter-satellite communications for model aggregation. Extensive performance evaluation under Walker Delta-based LEO constellation configurations and different datasets including MNIST, CIFAR-10, and EuroSat revealed that FedOrbit can significantly reduce communication rounds, power consumption and training time in comparison to state-of-the-art FL approaches while maintaining a high accuracy. FedOrbit demonstrates a significant decrease in power consumption, specifically by 8.8% and 79.1% for the MNIST dataset, when compared to decentralised and centralised FL approaches, respectively. The proposed technique can also reduce the training time by 5× and 48× compared with the decentralised and centralised FL approaches, respectively. Mohammad Reza Jabbarpour, Bahman Javadi, Philip H. W. Leong, Rodrigo N. Calheiros, David Boland, Chris Butler |
ICPADS | 4 |
| 2023 | NeurOPar, A Neural Network-Driven EDP Optimization Strategy for Parallel WorkloadsabstractThe pursuit of energy efficiency has been driving the development of techniques to optimize hardware resource usage in high-performance computing (HPC) servers. On multicore architectures, thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are three popular methods applied to improve the trade-off between performance and energy consumption, represented by the energy-delay product (EDP). However, the complexity of selecting the optimal configuration (TLP degree, DVFS, and UFS) for each application poses a challenge to software developers and end-users due to the massive number of possible configurations. To tackle this challenge, we propose NeurOpar, an optimization strategy for parallel workloads driven by an artificial neural network (ANN). It uses representative hardware and software metrics to build and train an ANN model that predicts combinations of thread count and core/uncore frequency levels that provide optimal EDP results. Through experiments on four multicore processors using twenty-five applications, we demonstrate that NeurOPar predicts combinations that yield EDP values close to the best ones achieved by an exhaustive search and improve the overall EDP by 42% compared to the default execution of HPC applications. We also show that NeurOPar can enhance the execution of parallel applications without incurring the performance and energy penalties associated with online methods by comparing it with two state-of-the-art strategies. Cristiano A. Künas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon |
SBAC-PAD | 4 |
| 2023 | EdgeSimPy: Python-based modeling and simulation of edge computing resource management policies
Paulo Silas Severo de Souza, Tiago Ferreto, Rodrigo N. Calheiros |
Future Gener. Comput. Syst. | 3 |
| 2022 | Scheduling Algorithms for Efficient Execution of Stream Workflow Applications in Multicloud EnvironmentsabstractBig data processing applications are becoming more and more complex. They are no more monolithic in nature but instead they are composed of decoupled analytical processes in the form of a workflow. One type of such workflow applications is stream workflow application, which integrates multiple streaming big data applications to support decision making. Each analytical component of these applications runs continuously and processes data streams whose velocity will depend on several factors such as network bandwidth and processing rate of parent analytical component. As a consequence, the execution of these applications on cloud environments requires advanced scheduling techniques that adhere to end user’s requirements in terms of data processing and deadline for decision making. In this article, we propose two multicloud scheduling and resource allocation techniques for efficient execution of stream workflow applications on multicloud environments while adhering to workflow application and user performance requirements and reducing execution cost. Results showed that the proposed genetic algorithm is an adequate and effective for all experiments. Mutaz Barika, Saurabh Kumar Garg 0001, Andrew H. C. Chan, Rodrigo N. Calheiros |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | Hybrid Workflow Provisioning and Scheduling on Cooperative Edge Cloud ComputingabstractThe dramatic growth of IoT-based applications in many domains such as real-time monitoring, interactive reporting, and smart manufacturing brings challenges for adoption of cloud-based solutions for integration of latency-sensitive and resource-intensive applications. We refer to this integration as a hybrid-workflow. This paper provides a resource estimation and task scheduling framework to run hybrid workflows on edge and cloud computing systems. We propose an adaptive resource estimation technique with an online gradient descent approximation to handle the complexity of hybrid workflows. In addition, a scheduling technique to execute workflow tasks on a cooperative edge cloud system to resolve the issues of latency-sensitive application as well as to improve resource utilization at the edge layer is proposed. Experimental results show the capability of the cooperative model in reducing the time and cost of running complex and large scale hybrid workflows. Raed Alsurdeh, Rodrigo N. Calheiros, Kenan M. Matawie, Bahman Javadi |
CCGRID | 2 |
| 2021 | BigDataSDNSim: A simulator for analyzing big data applications in software-defined cloud data centersabstractAbstract The integration and crosscoordination of big data processing and software‐defined networking (SDN) are vital for improving the performance of big data applications. Various approaches for combining big data and SDN have been investigated by both industry and academia. However, empirical evaluations of solutions that combine big data processing and SDN are extremely costly and complicated. To address the problem of effective evaluation of solutions that combine big data processing with SDN, we present a new, self‐contained simulation tool named BigDataSDNSim that enables the modeling and simulation of the big data management system YARN, its related programming models MapReduce, and SDN‐enabled networks in a cloud computing environment. BigDataSDNSim supports cost‐effective and easy to conduct experimentation in a controllable, repeatable, and configurable manner. The article illustrates the simulation accuracy and correctness of BigDataSDNSim by comparing the behavior and results of a real environment that combines big data processing and SDN with an equivalent simulated environment. Finally, the article presents two uses cases of BigDataSDNSim, which exhibit its practicality and features, illustrate the impact of data replication mechanisms of MapReduce in Hadoop YARN, and show the superiority of SDN over traditional networks to improve the performance of MapReduce applications. Khaled Alwasel, Rodrigo N. Calheiros, Saurabh Kumar Garg 0001, Rajkumar Buyya, Mukaddim Pathan, Dimitrios Georgakopoulos 0001, Rajiv Ranjan 0001 |
Softw. Pract. Exp. | 2 |
| 2021 | Guest Editorial: Special issue on blockchain and decentralized applications
Zibin Zheng, Shangguang Wang, Rodrigo N. Calheiros |
Softw. Pract. Exp. | 3 |
| 2021 | SLA-Based Profit Optimization Resource Scheduling for Big Data Analytics-as-a-Service Platforms in Cloud Computing EnvironmentsabstractThe value that can be extracted from big data greatly motivates users to explore data analytics technologies for better decision making and problem solving in various application domains. Analytical solutions can be expensive due to the demand for large-scale and high-performance computing resources. To provision online big data Analytics-as-a-Service (AaaS) to users in various domains, a general purpose AaaS platform is required to deliver on-demand services at low cost and in an easy to use manner. Our research focuses on proposing efficient and automatic admission control and resource scheduling algorithms for AaaS platforms in cloud environments. In this paper, we propose scalable and automatic admission control and profit optimization resource scheduling algorithms, which effectively admit data analytics requests, dynamically provision resources, and maximize profit for AaaS providers, while satisfying QoS requirements of queries with Service Level Agreement (SLA) guarantees. Moreover, the proposed algorithms enable users to trade-off accuracy for faster response times and less resource costs for query processing on large datasets. We evaluate the algorithm performance by adopting a data splitting method to process smaller data samples as representatives of the original big datasets. We conduct extensive experiments to evaluate the proposed admission control and profit optimization scheduling algorithms. Experimental evaluation shows the algorithms perform significantly better compared to the state-of-the-art algorithms in enhancing profits, reducing resource costs, increasing query admission rates, and decreasing query response times. Yali Zhao, Rodrigo N. Calheiros, Graeme Gange, James Bailey 0001, Richard O. Sinnott |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | Smart Food Scanner System Based on Mobile Edge ComputingabstractSmart applications, including Internet of Things (IoT) and Big Data analytics, are traditionally hosted by cloud infrastructures, which can result in high latency and cost beyond users expectation. Edge computing has emerged as a paradigm that can alleviate the pressure on clouds by delegating parts of the computation to devices in the edge of the network, at closer proximity to end users and IoT devices. In this paper, we discuss a smart application, built on top of mobile edge computing concept, to enables users to measure and analyse their food intake and support nutritional decision-making. The approach utilizes mobile edge computing to offload application computations and communications to the edge, thus saving battery life, increasing the processing capacity, and improving user comfort. In order to develop this system, we propose a loosely coupled architecture for a smart food scanner and then implement it using various IoT sensors. The performance evaluation results reveal that the implemented system can be used as an interactive appliance by users with minimum dependency and usage of their mobile phones. Bahman Javadi, Quoc Lap Trieu, Kenan M. Matawie, Rodrigo N. Calheiros |
IC2E | 4 |
| 2020 | Hybrid Workflow Provisioning and Scheduling on Edge Cloud Computing Using a Gradient Descent Search ApproachabstractThe dramatic growth of the Internet of Things (IoT) technology in many application domains, ranging from intelligent video surveillance, smart retail to the Internet-of-Vehicles brings new computation challenges for rationalized utilization of computing resources. IoT application execution refers to hybrid processing model of stream and batch to achieve data analytics objectives. Hybrid workflow execution combines the challenges of latency-sensitive and resource-intensive processing. To resolve these challenges, we proposed a two stages hybrid workflow scheduling framework on edge cloud computing. In the first stage, we proposed a resource estimation algorithm based on a linear optimization approach, the gradient descent search (GDS) and in the second stage, we adopted a cluster-based provisioning and scheduling technique on heterogeneous edge cloud resources. This work provides a multi-objective optimization model for execution time and monetary cost under constraints of deadline and throughput. Results demonstrated the framework performance in controlling the execution of hybrid workflows by an efficient tuning for stream processing parameters, such as arrival rate and processing throughput. Under working constraints, the proposed scheduler provides significant improvement for large hybrid workflows in terms of execution time and monetary cost with an average of 8% and 35%, respectively. Raed Alsurdeh, Rodrigo N. Calheiros, Kenan M. Matawie, Bahman Javadi |
ISPDC | 2 |
| 2020 | Data-intensive application scheduling on Mobile Edge Cloud Computing
Mohammad Alkhalaileh, Rodrigo N. Calheiros, Quang Vinh Nguyen 0002, Bahman Javadi |
J. Netw. Comput. Appl. | 2 |
| 2019 | SLA-Aware and Deadline Constrained Profit Optimization for Cloud Resource Management in Big Data Analytics-as-a-Service PlatformsabstractDiscovering optimal data analytics solutions to extract value from data for better and faster decision making is essential for many application domains, especially in the big data era. Big data analytics typically requires a tremendous amount of computational resources to process large data volumes that can be very expensive and time consuming. Our research focuses on providing optimization solutions for Analytics-as-a-Service (AaaS) platforms that automatically and elastically provision cloud resources to execute queries guaranteeing Service Level Agreements (SLAs) across a range of Quality of Service (QoS) requirements. We propose admission control and resource scheduling algorithms for AaaS platforms to maximize profits while providing time-minimized query execution plans to meet user demands and expectations. To enable timely responses as required for many domains, the algorithms utilize data splitting-based query admission and resource scheduling offering parallel processing on the split datasets. Extensive experiments are conducted to evaluate the algorithm performance compared to state-of-the-art optimization algorithms. Experimental results show that our algorithms perform significantly better from a range of perspectives, including increasing query admission rates and creating higher profits, whilst supporting efficient resource configurations that are able to support big data processing demands under tight deadlines. Yali Zhao, Rodrigo N. Calheiros, Athanasios V. Vasilakos, James Bailey 0001, Richard O. Sinnott |
CLOUD | 2 |
| 2019 | Towards Balancing Energy Savings and Performance for Volunteer Computing through Virtualized Approach
Fábio D. Rossi, Tiago Ferreto, Marcelo Da Silva Conterato, Paulo Silas Severo de Souza, Wagner dos Santos Marques, Rodrigo N. Calheiros, Guilherme da Cunha Rodrigues |
CLOSER | 6 |
| 2019 | ProactiveCache: On Reducing Degraded Read Latency of Erasure Coded Cloud StorageabstractErasure coding is gaining attraction in cloud storage systems because it improves data reliability with huge cost savings in terms of storage. However, data recovery in erasure codes includes high disk I/O, network traffic and complex decoding that impacts degraded read latency, in case of failures. Data access latency is one of the most important metrics to determine Quality of Service. Reducing degraded latency in erasure coding is vital to improve user performance. To reduce degraded read latency of erasure codes, in this paper, we have proposed a cache based technique called ProactiveCache. This proactively copies objects in failure predicted machine into a cache tier. To deploy ProactiveCache, cloud storage system should employ various failure prediction methods to predict hardware failures. On accurate failure predictions, ProactiveCache eliminates degraded read latency. For evaluation, ProactiveCache is implemented on Ceph object storage. Experimental results show that Proactive-Cache reduces degraded read latency up to 38% and improves throughput by 37%. Rekha Nachiappan, Bahman Javadi, Rodrigo N. Calheiros, Kenan M. Matawie |
CloudCom | 3 |
| 2019 | Dynamic Resource Allocation in Hybrid Mobile Cloud Computing for Data-Intensive Applications
Mohammad Alkhalaileh, Rodrigo N. Calheiros, Quang Vinh Nguyen 0002, Bahman Javadi |
GPC | 2 |
| 2019 | IoTSim-Stream: Modelling stream graph application in cloud simulation
Mutaz Barika, Saurabh Kumar Garg 0001, Andrew H. C. Chan, Rodrigo N. Calheiros, Rajiv Ranjan 0001 |
Future Gener. Comput. Syst. | 4 |
| 2018 | Adaptive Bandwidth-Efficient Recovery Techniques in Erasure-Coded Cloud Storage
Rekha Nachiappan, Bahman Javadi, Rodrigo N. Calheiros, Kenan M. Matawie |
Euro-Par | 3 |
| 2018 | Cloud Resource Provisioning for Combined Stream and Batch WorkflowsabstractThe increasing adoption of Internet of Thing (IoT) technology in many application domains generates a new need for rationalized utilization of computing resources supporting such computations. IoT applications can be represented as workflows in which stream and batch applications are integrated to accomplish data analytics objectives in many application domains such as smart home, health care, bioinformatics, astronomy, education, etc. The main challenge of this combination is the differentiation of service quality constraints between the two computation paradigms. Stream processing is highly sensitive to real-time constraint while batch processes are usually resource-intensive. In this work we propose a resource provisioning framework for combined workflows which aims to find an optimal workflow configuration plan to minimize execution time and monetary cost. The framework has functions of execution plan generation, task clustering, and resource provisioning. Results show that framework is capable to control the execution of combined-workflows by efficient tunning several parameters including stream arrival rate and processing throughput. Raed Alsurdeh, Rodrigo N. Calheiros, Kenan M. Matawie, Bahman Javadi |
IPCCC | 2 |
| 2018 | Evaluating container-based virtualization overhead on the general-purpose IoT platformabstractVirtualization has become a key technology that provides several advantages (e.g., flexibility, migration, isolation) for a plethora of computing infrastructures. However, traditional virtualization models are not suitable for embedded IoT platforms due to the virtualization layer verhead. New virtualization proposals such as container-based approaches arise as an option where performance is not impacted. However, when working on general-purpose embedded platforms, some studies have demonstrated that applications on container-based virtualization on embedded devices present considerable performance overhead. Since most performance evaluations on platforms using containers were run on servers, this study expands the testbed scenario by analyzing several metrics that measure the overhead of container-based virtualization layer on embedded IoT devices. Results demonstrated improvements up to 23% in terms of performance and up to 32% in terms of EDP. Wagner dos Santos Marques, Paulo Silas Severo de Souza, Fábio D. Rossi, Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Marcelo Da Silva Conterato, Tiago Ferreto |
ISCC | 5 |
| 2018 | Unfolding the Mutual Relation Between Timeliness and Scalability in Cloud MonitoringabstractCloud computing is a suitable solution for professionals, companies, and institutions that need to have access to computational resources on demand. Clouds rely on proper management to provide such computational resources with adequate quality of service, which is established by Service Level Agreements (SLAs), to customers. In this context, cloud monitoring is a critical function to achieve such proper management. Cloud monitoring systems have to accomplish requirements to perform its functions properly, and currently, there are plenty of requirements which includes: timeliness, adaptability, comprehensiveness, and scalability. However, such requirements usually have mutual influence, which is positive or negative, among themselves, and it has prevented the development of complete cloud monitoring solutions. This paper presents a mathematical model to predict the mutual influence between timeliness and scalability, which is a step forward in cloud monitoring because it paves the way for the development of complete monitoring solutions. It complements our previous work that identified the monitoring parameters (e.g., frequency sampling, amount of monitoring data) that influence timeliness and scalability. Evaluations present the effectiveness of the mathematical model based on a comparison of the results provided by the mathematical model and the results obtained via simulation. Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Glederson Lessa dos Santos, Vinicius Tavares Guimaraes, Lisandro Z. Granville, Liane Margarida Rockenbach Tarouco, Rajkumar Buyya |
ISCC | 2 |
| 2018 | A Stepwise Auto-Profiling Method for Performance Optimization of Streaming ApplicationsabstractData stream management systems (DSMSs) are scalable, highly available, and fault-tolerant systems that aggregate and analyze real-time data in motion. To continuously perform analytics on the fly within the stream, state-of-the-art DSMSs host streaming applications as a set of interconnected operators, with each operator encapsulating the semantic of a specific operation. For parallel execution on a particular platform, these operators need to be appropriately replicated in multiple instances that split and process the workload simultaneously. Because the way operators are partitioned affects the resulting performance of streaming applications, it is essential for DSMSs to have a method to compare different operators and make holistic replication decisions to avoid performance bottlenecks and resource wastage. To this end, we propose a stepwise profiling approach to optimize application performance on a given execution platform. It automatically scales distributed computations over streams based on application features and processing power of provisioned resources and builds the relationship between provisioned resources and application performance metrics to evaluate the efficiency of the resulting configuration. Experimental results confirm that the proposed approach successfully fulfills its goals with minimal profiling overhead. Xunyun Liu, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Chenhao Qu, Rajkumar Buyya |
ACM Trans. Auton. Adapt. Syst. | 3 |
| 2018 | An Online Algorithm for Task Offloading in Heterogeneous Mobile CloudsabstractMobile cloud computing is emerging as a promising approach to enrich user experiences at the mobile device end. Computation offloading in a heterogeneous mobile cloud environment has recently drawn increasing attention in research. The computation offloading decision making and tasks scheduling among heterogeneous shared resources in mobile clouds are becoming challenging problems in terms of providing global optimal task response time and energy efficiency. In this article, we address these two problems together in a heterogeneous mobile cloud environment as an optimization problem. Different from conventional distributed computing system scheduling problems, our joint offloading and scheduling optimization problem considers unique contexts of mobile clouds such as wireless network connections and mobile device mobility, which makes the problem more complex. We propose a context-aware mixed integer programming model to provide off-line optimal solutions for making the offloading decisions and scheduling the offloaded tasks among the shared computing resources in heterogeneous mobile clouds. The objective is to minimize the global task completion time (i.e., makespan). To solve the problem in real time, we further propose a deterministic online algorithm—the Online Code Offloading and Scheduling (OCOS) algorithm—based on the rent/buy problem and prove the algorithm is 2-competitive. Performance evaluation results show that the OCOS algorithm can generate schedules that have around two times shorter makespan than conventional independent task scheduling algorithms. Also, it can save around 30% more on makespan of task execution schedules than conventional offloading strategies, and scales well as the number of users grows. Bowen Zhou 0007, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Rajkumar Buyya |
ACM Trans. Internet Techn. | 3 |
| 2017 | Dynamic Network Bandwidth Resizing for Big Data ApplicationsabstractBig Data concerns processing of large volumes of digital data with high velocity and variety. Big Data technologies allow the analysis of data in real time, which is critical for various eScience applications. In order to meet the growing demand of Big Data applications, the infrastructures must be flexible enough to adapt to the characteristics of the applications. Most of the solutions presented in the literature to support Big Data applications focus on scaling processors and memory to handle a variable demand from applications. In a complementary way, this article targets the problem of adapting the network bandwidth to the amount of data to be transferred to and from the applications in order to improve the performance of the applications. For this purpose, we propose the use link aggregation protocol along with Software-Defined Network capabilities for management of the network flow. Results showed that the proposed approach improves the application's performance by up to 33%. Fábio D. Rossi, Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Marcelo Da Silva Conterato |
eScience | 3 |
| 2017 | On the effectiveness of isolation-based anomaly detection in cloud data centersabstractSummary The high volume of monitoring information generated by large‐scale cloud infrastructures poses a challenge to the capacity of cloud providers in detecting anomalies in the infrastructure. Traditional anomaly detection methods are resource‐intensive and computationally complex for training and/or detection, what is undesirable in very dynamic and large‐scale environment such as clouds. Isolation‐based methods have the advantage of low complexity for training and detection and are optimized for detecting failures. In this work, we explore the feasibility of Isolation Forest, an isolation‐based anomaly detection method, to detect anomalies in large‐scale cloud data centers. We propose a method to code time‐series information as extra attributes that enable temporal anomaly detection and establish its feasibility to adapt to seasonality and trends in the time‐series and to be applied online and in real‐time. Rodrigo N. Calheiros, Kotagiri Ramamohanarao, Rajkumar Buyya, Christopher Leckie, Steve Versteeg |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Mitigating impact of short-term overload on multi-cloud web applications through geographical load balancingabstractSummary Managed by an auto‐scaler in the clouds, applications may still be overloaded by sudden flash crowds or resource failures as the auto‐scaler takes time to make scaling decisions and provision resources. With more cloud providers building geographically dispersed data centers, applications are commonly deployed in multiple data centers to better serve customers worldwide. In this case, instead of sufficiently over‐provisioning each data center to prepare for occasional overloads, it is more cost‐efficient to over‐provision each data center a small amount of capacity and to balance the extra load among them when resources in any data center are suddenly saturated. In this paper, we present a decentralized system that timely detects short‐term overload situations and autonomously handles them using geographical load balancing and admission control to minimize the resulted performance degradation. Our approach also includes a new algorithm that optimally distributes the excessive load to remote data centers causing minimum increase of overall response times. We developed a prototype and evaluated it on Amazon Web Services. The results show that our approach is able to maintain acceptable quality of service while greatly increase the number of requests served during overloading periods. Chenhao Qu, Rodrigo N. Calheiros, Rajkumar Buyya |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Modeling and simulation of global and sleep states in ACPI-compliant energy-efficient cloud environmentsabstractSummary The more large‐scale data centers infrastructure costs increase, the more simulation‐based evaluations are needed to understand better the trade‐off between energy and performance and support the development of new energy‐aware resource allocation policies. Specifically, in the cloud computing field, various simulators are able to predict and measure the behavior of applications on different architectures using different resource allocation policies. Yet, only a few of them have the ability to simulate energy‐saving strategies, and none of them support the complete advanced configuration and power interface (ACPI) specification. ACPI defines a terminology for all possible power states of a machine and their associated power rate. The hardware industry has relied on ACPI to provide up‐to‐date standard interfaces for hardware discovery, configuration, power management, and monitoring, enabling a better understanding of the energy consumption level of different hardware states, referred to as ACPI G‐states, S‐states, and P‐states. In this paper, we improve the modeling and simulation of the ACPI G/S‐states and show not only that these states offer different energy‐saving levels but also that state transitions consume energy. In addition, we model the latency to transit between two states and the effects on the turnaround time when the transitions are not performed conservatively. Furthermore, the equations provide essential information to quantify the trade‐off between energy consumption and performance and assist in the analysis/decision on which strategy fits better in the environment and how it could be refined. Our expanded energy model was implemented in CloudSim and validated with simulation‐based experiments with a very high level of accuracy, with a standard deviation of at most 6%. Copyright © 2016 John Wiley & Sons, Ltd. Miguel G. Xavier, Fábio D. Rossi, César A. F. De Rose, Rodrigo N. Calheiros, Danielo Goncalves Gomes |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | An algorithm for network and data-aware placement of multi-tier applications in cloud data centers
Md. Hasanul Ferdaus, M. Manzur Murshed, Rodrigo N. Calheiros, Rajkumar Buyya |
J. Netw. Comput. Appl. | 3 |
| 2017 | Cloud storage reliability for Big Data applications: A state of the art survey
Rekha Nachiappan, Bahman Javadi, Rodrigo N. Calheiros, Kenan M. Matawie |
J. Netw. Comput. Appl. | 3 |
| 2017 | E-eco: Performance-aware energy-efficient cloud data center orchestration
Fábio D. Rossi, Miguel G. Xavier, César A. F. De Rose, Rodrigo N. Calheiros, Rajkumar Buyya |
J. Netw. Comput. Appl. | 4 |
| 2017 | ContainerCloudSim: An environment for modeling and simulation of containers in cloud data centersabstractSummary Containers are increasingly gaining popularity and becoming one of the major deployment models in cloud environments. To evaluate the performance of scheduling and allocation policies in containerized cloud data centers, there is a need for evaluation environments that support scalable and repeatable experiments. Simulation techniques provide repeatable and controllable environments, and hence, they serve as a powerful tool for such purpose. This paper introduces ContainerCloudSim , which provides support for modeling and simulation of containerized cloud computing environments. We developed a simulation architecture for containerized clouds and implemented it as an extension of CloudSim. We described a number of use cases to demonstrate how one can plug in and compare their container scheduling and provisioning policies in terms of energy efficiency and SLA compliance. Our system is highly scalable as it supports simulation of large number of containers, given that there are more containers than virtual machines in a data center. Copyright © 2016 John Wiley & Sons, Ltd. Sareh Fotuhi Piraghaj, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Rajkumar Buyya |
Softw. Pract. Exp. | 3 |
| 2017 | mCloud: A Context-Aware Offloading Framework for Heterogeneous Mobile CloudabstractMobile cloud computing (MCC) has become a significant paradigm for bringing the benefits of cloud computing to mobile devices' proximity. Service availability along with performance enhancement and energy efficiency are primary targets in MCC. This paper proposes a code offloading framework, called mCloud, which consists of mobile devices, nearby cloudlets and public cloud services, to improve the performance and availability of the MCC services. The effect of the mobile device context (e.g., network conditions) on offloading decisions is studied by proposing a context-aware offloading decision algorithm aiming to provide code offloading decisions at runtime on selecting wireless medium and appropriate cloud resources for offloading. We also investigate failure detection and recovery policies for our mCloud system. We explain in details the design and implementation of the mCloud prototype framework. We conduct real experiments on the implemented system to evaluate the performance of the algorithm. Results indicate the system and embedded decision algorithm are able to provide decisions on selecting wireless medium and cloud resources based on different context of the mobile devices, and achieve significant reduction on makespan and energy, with the improved service availability when compared with existing offloading schemes. Bowen Zhou 0007, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Satish Narayana Srirama, Rajkumar Buyya |
IEEE Trans. Serv. Comput. | 3 |
| 2017 | SLA-Aware and Energy-Efficient Dynamic Overbooking in SDN-Based Cloud Data CentersabstractPower management of cloud data centers has received great attention from industry and academia as they are expensive to operate due to their high energy consumption. While hosts are dominant to consume electric power, networks account for 10 to 20 percent of the total energy costs in a data center. Resource overbooking is one way to reduce the usage of active hosts and networks by placing more requests to the same amount of resources. Network resource overbooking can be facilitated by Software Defined Networking (SDN) that can consolidate traffics and control Quality of Service (QoS) dynamically. However, the existing approaches employ fixed overbooking ratio to decide the amount of resources to be allocated, which in reality may cause excessive Service Level Agreements (SLA) violation with workloads being unpredictable. In this paper, we propose dynamic overbooking strategy which jointly leverages virtualization capabilities and SDN for VM and traffic consolidation. With the dynamically changing workload, the proposed strategy allocates more precise amount of resources to VMs and traffics. This strategy can increase overbooking in a host and network while still providing enough resources to minimize SLA violations. Our approach calculates resource allocation ratio based on the historical monitoring data from the online analysis of the host and network utilization without any pre-knowledge of workloads. We implemented it in simulation environment in large scale to demonstrate the effectiveness in the context of Wikipedia workloads. Our approach saves energy consumption in the data center while reducing SLA violations. Jungmin Son, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Rajkumar Buyya |
IEEE Trans. Sustain. Comput. | 3 |
| 2016 | SLA-based profit optimization for resource management of big data analytics-as-a-service platforms in cloud computing environmentsabstractThe value that can be extracted from big data greatly motivates organizations to explore data analytics technologies for better decision making and problem solving in a wide range of application domains. Cloud computing greatly eases and benefits big data analytics by offering on-demand and scalable computing infrastructures, platforms, and applications as services. Big data Analytics-as-a-Service (AaaS) platforms aim to deliver data analytics as consumable services in cloud computing environments in a pay as you go model with Service Level Agreement (SLA) guarantees. Resource scheduling for AaaS platforms is significant as big data analytics requires large-scale computing, which can consume huge amounts of resources and incur high resource costs. Our research focuses on proposing automatic and scalable resource scheduling algorithms to maximize the profits for AaaS platforms while delivering AaaS services to users with SLA guarantees on budgets and deadlines to allow timely responses with controllable costs. In this paper, we model and formulate the profit optimization resource scheduling problem and propose an optimization scheduling algorithm that maximizes profits for AaaS platforms and guarantees SLAs for query requests. Experimental evaluations show that the profit optimization scheduling algorithm performs significantly better in cost saving and profit enhancement compared to the state-of-the-art scheduling algorithms. Yali Zhao, Rodrigo N. Calheiros, James Bailey 0001, Richard O. Sinnott |
IEEE BigData | 2 |
| 2016 | iGiraph: A Cost-Efficient Framework for Processing Large-Scale Graphs on Public CloudsabstractLarge-scale graph analytics has gained attention during the past few years. As the world is going to be more connected by appearance of new technologies and applications such as social networks, Web portals, mobile devices, Internet of things, etc, a huge amount of data are created and stored every day in the form of graphs consisting of billions of vertices and edges. Many graph processing frameworks have been developed to process these large graphs since Google introduced its graph processing framework called Pregel in 2010. On the other hand, cloud computing which is a new paradigm of computing that overcomes restrictions of traditional problems in computing by enabling some novel technological and economical solutions such as distributed computing, elasticity and pay-as-you-go models has improved service delivery features. In this paper, we present iGiraph, a cost-efficient Pregel-like graph processing framework for processing large-scale graphs on public clouds. iGiraph uses a new dynamic re-partitioning approach based on messaging pattern to minimize the cost of resource utilization on public clouds. We also present the experimental results on the performance and cost effects of our method and compare them with basic Giraph framework. Our results validate that iGiraph remarkably decreases the cost and improves the performance by scaling the number of workers dynamically. Safiollah Heidari, Rodrigo N. Calheiros, Rajkumar Buyya |
CCGrid | 2 |
| 2016 | Virtual Machine Customization and Task Mapping Architecture for Efficient Allocation of Cloud Data Center ResourcesabstractEnergy usage of large-scale data centers has become a major concern for cloud providers. There has been an active effort in techniques for the minimization of the energy consumed in the data centers. However, most approaches lack the analysis and application of real cloud backend traces. In existing approaches, the variation of cloud workloads and its effect on the performance of the solutions are not investigated. Furthermore, the focus of existing approaches is on virtual machine migration and placement algorithms, with little regard to tailoring virtualmachine configuration to workload characteristics, which can further reduce the energy consumption and resource wastage in a typical data center. To address these weaknesses and challenges, we propose a new architecture for cloud resource allocation that maps groups of tasks to customized virtual machine types. This mapping is based on the task usage patterns obtained from the analysis of the historical data extracted from utilization traces. In our work, the energy consumption is decreased via efficient resource allocation based on the actual resource usage of tasks. Experimental results show that, when resources are allocated based on the discovered usage patterns, significant energy saving can be achieved. Sareh Fotuhi Piraghaj, Rodrigo N. Calheiros, Jeffrey Chan, Amir Vahid Dastjerdi, Rajkumar Buyya |
Comput. J. | 2 |
| 2016 | Dynamic resource demand prediction and allocation in multi-tenant service cloudsabstractSummary Cloud computing is emerging as an increasingly popular computing paradigm, allowing dynamic scaling of resources available to users as needed. This requires a highly accurate demand prediction and resource allocation methodology that can provision resources in advance, thereby minimizing the virtual machine downtime required for resource provisioning. In this paper, we present a dynamic resource demand prediction and allocation framework in multi‐tenant service clouds. The novel contribution of our proposed framework is that it classifies the service tenants as per whether their resource requirements would increase or not; based on this classification, our framework prioritizes prediction for those service tenants in which resource demand would increase, thereby minimizing the time needed for prediction. Furthermore, our approach adds the service tenants to matched virtual machines and allocates the virtual machines to physical host machines using a best‐fit heuristic approach. Performance results demonstrate how our best‐fit heuristic approach could efficiently allocate virtual machines to hosts so that the hosts are utilized to their fullest capacity. Copyright © 2016 John Wiley & Sons, Ltd. Manish Verma, G. R. Gangadharan, Nanjangud C. Narendra, Vadlamani Ravi, Vidyadhar Inamdar, Lakshmi Ramachandran, Rodrigo N. Calheiros, Rajkumar Buyya |
Concurr. Comput. Pract. Exp. | 7 |
| 2016 | A reliable and cost-efficient auto-scaling system for web applications using heterogeneous spot instances
Chenhao Qu, Rodrigo N. Calheiros, Rajkumar Buyya |
J. Netw. Comput. Appl. | 2 |
| 2015 | SLO-Aware Deployment of Web Applications Requiring Strong Consistency Using Multiple CloudsabstractGeographically dispersed cloud data centers (DCs) enable web application providers to improve their services' response time and availability by deploying application replicas in multiple DCs. To allow applications requiring strong consistency to be deployed in multiple clouds, industry and academia have developed various scalable database systems that can guarantee strong inter-DC consistency with alleviated network overhead. For applications using these database systems, it is essential to take both the network latencies to the end users and the communication overhead of the databases into account when selecting the hosting DCs. In this paper, we study how to identify the satisfactory deployment plan (hosting DCs and request routing) considering SLO satisfaction, migration cost, and operational cost for applications using these databases. The proposed approach involves two steps. First, it searches the deployment plan with minimum amount of SLO violations using genetic algorithm when the application is first migrated to the clouds. Then it continuously optimizes the deployment in a certain time interval according to the changing workload and the current deployment plan. We illustrate how our approach works for the applications using two databases (Cassandra and Galera Cluster), and demonstrate the effectiveness of our approach through simulation studies using settings of two example applications (TPC-W and Twissandra). Our solution is extensible to applications using other database systems that have similar properties. Chenhao Qu, Rodrigo N. Calheiros, Rajkumar Buyya |
CLOUD | 2 |
| 2015 | A Context Sensitive Offloading Scheme for Mobile Cloud Computing ServiceabstractMobile cloud computing (MCC) has drawn significant research attention as the popularity and capability of mobile devices have been improved in recent years. In this paper, we propose a prototype MCC offloading system that considers multiple cloud resources such as mobile ad-hoc network, cloudlet and public clouds to provide an adaptive MCC service. We propose a context-aware offloading decision algorithm aiming to provide code offloading decisions at runtime on selecting wireless medium and which potential cloud resources as the offloading location based on the device context. We also conduct real experiments on the implemented system to evaluate the performance of the algorithm. Results indicate the system and embedded decision algorithm can select suitable wireless medium and cloud resources based on different context of the mobile devices, and achieve significant performance improvement. Bowen Zhou 0007, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Satish Narayana Srirama, Rajkumar Buyya |
CLOUD | 3 |
| 2015 | CloudSimSDN: Modeling and Simulation of Software-Defined Cloud Data CentersabstractSoftware-Defined Networking not only addresses the shortcoming of traditional network technologies in dealing with frequent and immediate changes in cloud data centers but also made network resource management open and innovation-friendly. To further accelerate the innovation pace, accessible and easy-to-learn testbeds are required which estimate and measure the performance of network and host capacity provisioning approaches simultaneously within a data center. This is a challenging task and is often costly if accomplished in a physical environment. Thus, a lightweight and scalable simulation environment is necessary to evaluate the network allocation capacity policies while avoiding such a complicated and expensive facility. This paper introduces CloudSimSDN, a simulation framework for SDN-enabled cloud environments based on CloudSim. This paper develops and presents the overall architecture and features of the framework and provides several use cases. Moreover, we empirically validate the accuracy and effectiveness of CloudSimSDN through a number of simulations of a cloud-based three-tier web application. Jungmin Son, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Xiaohui Ji, Young Yoon, Rajkumar Buyya |
CCGRID | 3 |
| 2015 | Big Data Analytics-Enhanced Cloud Computing: Challenges, Architectural Elements, and Future DirectionsabstractThe emergence of cloud computing has made dynamic provisioning of elastic capacity to applications on-demand. Cloud data centers contain thousands of physical servers hosting orders of magnitude more virtual machines that can be allocated on demand to users in a pay-as-you-go model. However, not all systems are able to scale up by just adding more virtual machines. Therefore, it is essential, even for scalable systems, to project workloads in advance rather than using a purely reactive approach. Given the scale of modern cloud infrastructures generating real time monitoring information, along with all the information generated by operating systems and applications, this data poses the issues of volume, velocity, and variety that are addressed by Big Data approaches. In this paper, we investigate how utilization of Big Data analytics helps in enhancing the operation of cloud computing environments. We discuss diverse applications of Big Data analytics in clouds, open issues for enhancing cloud operations via Big Data analytics, and architecture for anomaly detection and prevention in clouds along with future research directions. Rajkumar Buyya, Kotagiri Ramamohanarao, Christopher Leckie, Rodrigo N. Calheiros, Amir Vahid Dastjerdi, Steve Versteeg |
ICPADS | 4 |
| 2015 | SLA-Based Resource Scheduling for Big Data Analytics as a Service in Cloud Computing EnvironmentsabstractData analytics plays a significant role in gaining insight of big data that can benefit in decision making and problem solving for various application domains such as science, engineering, and commerce. Cloud computing is a suitable platform for Big Data Analytic Applications (BDAAs) that can greatly reduce application cost by elastically provisioning resources based on user requirements and in a pay as you go model. BDAAs are typically catered for specific domains and are usually expensive. Moreover, it is difficult to provision resources for BDAAs with fluctuating resource requirements and reduce the resource cost. As a result, BDAAs are mostly used by large enterprises. Therefore, it is necessary to have a general Analytics as a Service (AaaS) platform that can provision BDAAs to users in various domains as consumable services in an easy to use way and at lower price. To support the AaaS platform, our research focuses on efficiently scheduling Cloud resources for BDAAs to satisfy Quality of Service (QoS) requirements of budget and deadline for data analytic requests and maximize profit for the AaaS platform. We propose an admission control and resource scheduling algorithm, which not only satisfies QoS requirements of requests as guaranteed in Service Level Agreements (SLAs), but also increases the profit for AaaS providers by offering a cost-effective resource scheduling solution. We propose the architecture and models for the AaaS platform and conduct experiments to evaluate the proposed algorithm. Results show the efficiency of the algorithm in SLA guarantee, profit enhancement, and cost saving. Yali Zhao, Rodrigo N. Calheiros, Graeme Gange, Kotagiri Ramamohanarao, Rajkumar Buyya |
ICPP | 2 |
| 2015 | The interplay between timeliness and scalability in cloud monitoring systemsabstractCloud computing is a groundbreaking solution to acquire computational resources on demand. To deliver high quality cloud services and provide features such as reduced costs and availability to customers, a cloud, like any other computational system, needs to be properly managed in accordance with its characteristics (e.g., scalability, elasticity, timeliness). In this scenario, cloud monitoring is a key to achieve it. To properly work, cloud monitoring systems need to meet several requirements such as scalability, accuracy, and timeliness. This paper aims to unveil the trade-off between timeliness and scalability. Evaluations demonstrate the mutual influence between scalability and timeliness based on monitoring parameters (e.g., monitoring topologies, frequency sampling). Results show that non-deep monitoring topologies and decreasing the frequency sampling assist to reduce the mutual influence between timeliness and scalability. Guilherme da Cunha Rodrigues, Rodrigo N. Calheiros, Marcio Barbosa de Carvalho, Carlos Raniery Paula dos Santos, Lisandro Z. Granville, Liane Margarida Rockenbach Tarouco, Rajkumar Buyya |
ISCC | 2 |
| 2015 | Efficient Virtual Machine Sizing for Hosting Containers as a Service (SERVICES 2015)abstractThere has been a growing effort in decreasing energy consumption of large-scale cloud data centers via maximization of host-level utilization and load balancing techniques. However, with the recent introduction of Container as a Service (CaaS) by cloud providers, maximizing the utilization at virtual machine (VM) level becomes essential. To this end, this paper focuses on finding efficient virtual machine sizes for hosting containers in such a way that the workload is executed with minimum wastage of resources on VM level. Suitable VM sizes for containers are calculated, and application tasks are grouped and clustered based on their usage patterns obtained from historical data. Furthermore, tasks are mapped to containers and containers are hosted on their associated VM types. We analyzed clouds' trace logs from Google cluster and consider the cloud workload variances, which is crucial for testing and validating our proposed solutions. Experimental results showed up to 7.55% improvement in the average energy consumption compared to baseline scenarios where the virtual machine sizes are fixed. In addition, comparing to the baseline scenarios, the total number of VMs instantiated for hosting the containers is also improved by 68% on average. Sareh Fotuhi Piraghaj, Amir Vahid Dastjerdi, Rodrigo N. Calheiros, Rajkumar Buyya |
SERVICES | 3 |
| 2015 | Big Data computing and clouds: Trends and future directions
Marcos Dias de Assunção, Rodrigo N. Calheiros, Silvia Bianchi, Marco Aurélio Stelmar Netto, Rajkumar Buyya |
J. Parallel Distributed Comput. | 2 |
| 2015 | Workload Prediction Using ARIMA Model and Its Impact on Cloud Applications' QoSabstractAs companies shift from desktop applications to cloud-based software as a service (SaaS) applications deployed on public clouds, the competition for end-users by cloud providers offering similar services grows. In order to survive in such a competitive market, cloud-based companies must achieve good quality of service (QoS) for their users, or risk losing their customers to competitors. However, meeting the QoS with a cost-effective amount of resources is challenging because workloads experience variation overtime. This problem can be solved with proactive dynamic provisioning of resources, which can estimate the future need of applications in terms of resources and allocate them in advance, releasing them once they are not required. In this paper, we present the realization of a cloud workload prediction module for SaaS providers based on the autoregressive integrated moving average (ARIMA) model. We introduce the prediction based on the ARIMA model and evaluate its accuracy of future workload prediction using real traces of requests to Web servers. We also evaluate the impact of the achieved accuracy in terms of efficiency in resource utilization and QoS. Simulation results show that our model is able to achieve an average accuracy of up to 91 percent, which leads to efficiency in resource utilization with minimal impact on the QoS. Rodrigo N. Calheiros, Enayat Masoumi, Rajiv Ranjan 0001, Rajkumar Buyya |
IEEE Trans. Cloud Comput. | 1 |
| 2014 | Energy-Efficient Scheduling of Urgent Bag-of-Tasks Applications in Clouds through DVFSabstractThe broad adoption of cloud services led to an increasing concentration of servers in a few data centers. Reports estimate the energy consumptions of these data centers to be between 1.1% and 1.5% of the worldwide electricity consumption. This extensive energy consumption precludes massive CO2 emissions, as a significant number of data centers are backed by "brown" power plants. While most researchers have focused on reducing energy consumption of cloud data centers via server consolidation, we propose an approach for reducing the power required to execute urgent, CPU-intensive Bag-of-Tasks applications on cloud infrastructures. It exploits intelligent scheduling combined with the Dynamic Voltage and Frequency Scaling (DVFS) capability of modern CPU processors to keep the CPU operating at the minimum voltage level (and consequently minimum frequency and power consumption) that enables the application to complete before a user-defined deadline. Experiments demonstrate that our approach reduces energy consumption with the extra feature of not requiring virtual machines to have knowledge about its underlying physical infrastructure, which is an assumption of previous works. Rodrigo N. Calheiros, Rajkumar Buyya |
CloudCom | 1 |
| 2014 | Virtual Machine Consolidation in Cloud Data Centers Using ACO Metaheuristic
Md. Hasanul Ferdaus, M. Manzur Murshed, Rodrigo N. Calheiros, Rajkumar Buyya |
Euro-Par | 3 |
| 2014 | Meeting Deadlines of Scientific Workflows in Public Clouds with Tasks ReplicationabstractThe elasticity of Cloud infrastructures makes them a suitable platform for execution of deadline-constrained workflow applications, because resources available to the application can be dynamically increased to enable application speedup. Existing research in execution of scientific workflows in Clouds either try to minimize the workflow execution time ignoring deadlines and budgets or focus on the minimization of cost while trying to meet the application deadline. However, they implement limited contingency strategies to correct delays caused by underestimation of tasks execution time or fluctuations in the delivered performance of leased public Cloud resources. To mitigate effects of performance variation of resources on soft deadlines of workflow applications, we propose an algorithm that uses idle time of provisioned resources and budget surplus to replicate tasks. Simulation experiments with four well-known scientific workflows show that the proposed algorithm increases the likelihood of deadlines being met and reduces the total execution time of applications as the budget available for replication increases. Rodrigo N. Calheiros, Rajkumar Buyya |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Scaling MapReduce Applications Across Hybrid Clouds to Meet Soft DeadlinesabstractCloud platforms make available a virtually infinite amount of computing resources, which are managed by third parties and are accessed by users on demand in a pay-per-use manner, with Quality of Service guarantees. This enables computing infrastructures to be scaled up and down accordingly to the amount of data to be processed. MapReduce is among the most popular models for development of Cloud applications. As the utilization of such programming model spreads across multiple application domains, the need for timely execution of these applications arises. While existing approaches focus in meeting deadlines via admission control or preemption of lower priority applications, we propose a policy for dynamic provisioning of Cloud resources to speed up execution of deadline-constrained MapReduce applications, by enabling concurrent execution of tasks, in order to meet a deadline for completion of the Map phase of the application. We describe the proposed algorithm and an actual implementation of it in the Aneka Cloud Platform. Experiments on such prototype implementation show that our proposed approach can effectively meet the soft deadlines while minimizing the budget for application execution. Michael Mattess, Rodrigo N. Calheiros, Rajkumar Buyya |
AINA | 2 |
| 2013 | EMUSIM: an integrated emulation and simulation environment for modeling, evaluation, and validation of performance of Cloud computing applicationsabstractSUMMARY Cloud computing allows the deployment and delivery of application services for users worldwide. Software as a Service providers with limited upfront budget can take advantage of Cloud computing and lease the required capacity in a pay‐as‐you‐go basis, which also enables flexible and dynamic resource allocation according to service demand. One key challenge potential Cloud customers have before renting resources is to know how their services will behave in a set of resources and the costs involved when growing and shrinking their resource pool. Most of the studies in this area rely on simulation‐based experiments, which consider simplified modeling of applications and computing environment. In order to better predict service's behavior on Cloud platforms, we developed an integrated architecture that is based on both simulation and emulation. The proposed architecture, named EMUSIM, automatically extracts information from application behavior via emulation and then uses this information to generate the corresponding simulation model. We performed experiments using an image processing application as a case study and found that EMUSIM was able to accurately model such application via emulation and use the model to supply information about its potential performance in a Cloud provider. We also discuss our experience using EMUSIM for deploying applications in a real public Cloud provider. EMUSIM is based on an open source software stack and therefore it can be extended for analysis behavior of several other applications. Copyright © 2012 John Wiley & Sons, Ltd. Rodrigo N. Calheiros, Marco Aurélio Stelmar Netto, César A. F. De Rose, Rajkumar Buyya |
Softw. Pract. Exp. | 1 |
| 2012 | Design and Development of an Adaptive Workflow-Enabled Spatial-Temporal Analytics FrameworkabstractCloud computing is a suitable platform for execution of complex computational tasks and scientific simulations that are described in the form of workflows. Such applications are managed by Workflow Management System (WfMS). Because existing WfMSs are not able to autonomically provision resources to real-time applications and schedule them while supporting fault tolerance and data privacy, we present a highly-scalable workflow-enabled analytics system that manages inter-dependable analytics tasks adaptively with varying operational requirements on a common platform and enables visualization of multidimensional datasets of real world phenomena. In this paper, we present the architecture of such a WfMS and evaluate it in terms of performance for execution of workflows in Clouds. A real world application of climate-associated dengue fever prediction was evaluated on public, private, and hybrid Clouds and experienced effective speedup in all the environments. Xiaorong Li, Rodrigo N. Calheiros, Sifei Lu, Long Wang 0005, Henry Novianus Palit, Qin Zheng 0002, Rajkumar Buyya |
ICPADS | 2 |
| 2012 | Cost-Effective Provisioning and Scheduling of Deadline-Constrained Applications in Hybrid Clouds
Rodrigo N. Calheiros, Rajkumar Buyya |
WISE | 1 |
| 2012 | A coordinator for scaling elastic applications across multiple clouds
Rodrigo N. Calheiros, Adel Nadjaran Toosi, Christian Vecchiola, Rajkumar Buyya |
Future Gener. Comput. Syst. | 1 |
| 2012 | The Aneka platform and QoS-driven resource provisioning for elastic applications on hybrid Clouds
Rodrigo N. Calheiros, Christian Vecchiola, Dileban Karunamoorthy, Rajkumar Buyya |
Future Gener. Comput. Syst. | 1 |
| 2012 | Towards autonomic detection of SLA violations in Cloud infrastructures
Vincent C. Emeakaroha, Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, Ivona Brandic, Rajkumar Buyya, César A. F. De Rose |
Future Gener. Comput. Syst. | 3 |
| 2012 | Deadline-driven provisioning of resources for scientific applications in hybrid clouds with Aneka
Christian Vecchiola, Rodrigo N. Calheiros, Dileban Karunamoorthy, Rajkumar Buyya |
Future Gener. Comput. Syst. | 2 |
| 2011 | Resource Provisioning Policies to Increase IaaS Provider's Profit in a Federated Cloud EnvironmentabstractCloud Federation is a recent paradigm that helps Infrastructure as a Service (IaaS) providers to overcome resource limitation during spikes in demand for Virtual Machines (VMs) by outsourcing requests to other federation members. IaaS providers also have the option of terminating spot VMs, i.e, cheaper VMs that can be canceled to free resources for more profitable VM requests. By both approaches, providers can expect to reject less profitable requests. For IaaS providers, pricing and profit are two important factors, in addition to maintaining a high Quality of Service (QoS) and utilization of their resources to remain in the business. For this, a clear understanding of the usage pattern, types of requests, and infrastructure costs are necessary while making decisions to terminate spot VMs, outsourcing or contributing to the federation. In this paper, we propose policies that help in the decision-making process to increase resources utilization and profit. Simulation results indicate that the proposed policies enhance the profit, utilization, and QoS (smaller number of rejected VM requests) in a Cloud federation environment. Adel Nadjaran Toosi, Rodrigo N. Calheiros, Ruppa K. Thulasiram, Rajkumar Buyya |
HPCC | 2 |
| 2011 | Virtual Machine Provisioning Based on Analytical Performance and QoS in Cloud Computing EnvironmentsabstractCloud computing is the latest computing paradigm that delivers IT resources as services in which users are free from the burden of worrying about the low-level implementation or system administration details. However, there are significant problems that exist with regard to efficient provisioning and delivery of applications using Cloud-based IT resources. These barriers concern various levels such as workload modeling, virtualization, performance modeling, deployment, and monitoring of applications on virtualized IT resources. If these problems can be solved, then applications can operate more efficiently, with reduced financial and environmental costs, reduced under-utilization of resources, and better performance at times of peak load. In this paper, we present a provisioning technique that automatically adapts to workload changes related to applications for facilitating the adaptive management of system and offering end-users guaranteed Quality of Services (QoS) in large, autonomous, and highly dynamic environments. We model the behavior and performance of applications and Cloud-based IT resources to adaptively serve end-user requests. To improve the efficiency of the system, we use analytical performance (queueing network system model) and workload information to supply intelligent input about system requirements to an application provisioner with limited information about the physical infrastructure. Our simulation-based experimental results using production workload models indicate that the proposed provisioning technique detects changes in workload intensity (arrival pattern, resource demands) that occur over time and allocates multiple virtualized IT resources accordingly to achieve application QoS targets. Rodrigo N. Calheiros, Rajiv Ranjan 0001, Rajkumar Buyya |
ICPP | 1 |
| 2011 | Server consolidation with migration control for virtualized data centers
Tiago Ferreto, Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, César A. F. De Rose |
Future Gener. Comput. Syst. | 3 |
| 2011 | CloudSim: a toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithmsabstractAbstract Cloud computing is a recent advancement wherein IT infrastructure and applications are provided as ‘services’ to end‐users under a usage‐based payment model. It can leverage virtualized services even on the fly based on requirements (workload patterns and QoS) varying with time. The application services hosted under Cloud computing model have complex provisioning, composition, configuration, and deployment requirements. Evaluating the performance of Cloud provisioning policies, application workload models, and resources performance models in a repeatable manner under varying system and user configurations and requirements is difficult to achieve. To overcome this challenge, we propose CloudSim: an extensible simulation toolkit that enables modeling and simulation of Cloud computing systems and application provisioning environments. The CloudSim toolkit supports both system and behavior modeling of Cloud system components such as data centers, virtual machines (VMs) and resource provisioning policies. It implements generic application provisioning techniques that can be extended with ease and limited effort. Currently, it supports modeling and simulation of Cloud computing environments consisting of both single and inter‐networked clouds (federation of clouds). Moreover, it exposes custom interfaces for implementing policies and provisioning techniques for allocation of VMs under inter‐networked Cloud computing scenarios. Several researchers from organizations, such as HP Labs in U.S.A., are using CloudSim in their investigation on Cloud resource provisioning and energy‐efficient management of data center resources. The usefulness of CloudSim is demonstrated by a case study involving dynamic provisioning of application services in the hybrid federated clouds environment. The result of this case study proves that the federated Cloud computing model significantly improves the application QoS requirements under fluctuating resource and service demand patterns. Copyright © 2010 John Wiley & Sons, Ltd. Rodrigo N. Calheiros, Rajiv Ranjan 0001, Anton Beloglazov, César A. F. De Rose, Rajkumar Buyya |
Softw. Pract. Exp. | 1 |
| 2010 | CloudAnalyst: A CloudSim-Based Visual Modeller for Analysing Cloud Computing Environments and ApplicationsabstractAdvances in Cloud computing opens up many new possibilities for Internet applications developers. Previously, a main concern of Internet applications developers was deployment and hosting of applications, because it required acquisition of a server with a fixed capacity able to handle the expected application peak demand and the installation and maintenance of the whole software infrastructure of the platform supporting the application. Furthermore, server was underutilized because peak traffic happens only at specific times. With the advent of the Cloud, deployment and hosting became cheaper and easier with the use of pay-peruse flexible elastic infrastructure services offered by Cloud providers. Because several Cloud providers are available, each one offering different pricing models and located in different geographic regions, a new concern of application developers is selecting providers and data center locations for applications. However, there is a lack of tools that enable developers to evaluate requirements of large-scale Cloud applications in terms of geographic distribution of both computing servers and user workloads. To fill this gap in tools for evaluation and modeling of Cloud environments and applications, we propose CloudAnalyst. It was developed to simulate large-scale Cloud applications with the purpose of studying the behavior of such applications under various deployment configurations. CloudAnalyst helps developers with insights in how to distribute applications among Cloud infrastructures and value added services such as optimization of applications performance and providers incoming with the use of Service Brokers. Bhathiya Wickremasinghe, Rodrigo N. Calheiros, Rajkumar Buyya |
AINA | 2 |
| 2010 | InterCloud: Utility-Oriented Federation of Cloud Computing Environments for Scaling of Application Services
Rajkumar Buyya, Rajiv Ranjan 0001, Rodrigo N. Calheiros |
ICA3PP (1) | 3 |
| 2010 | Building an automated and self-configurable emulation testbed for grid applicationsabstractAbstract Distributed systems, such as grids, are composed of geographically distributed computing elements that belong to multiple administrative domains and are controlled by multiple entities. It is unlikely that testers are able to acquire repeatedly the same resources, for the same amount of time, and under the same network conditions, which are paramount requirements for enabling reproducible and controlled tests in software under development. An alternative to experiments in real testbeds is the use of emulation tools, which allow the software to run in an environment that behaves like a distributed system. Although advances in virtualization technology allowed the development of efficient emulators, few efforts were put in making operation of such emulators easier. This paper presents the design and the development of the Automated Emulation Framework that allows automatic mapping of virtual machines to hosts, virtual machine deployment, network configuration, and proactive management and reconfiguration of the virtual infrastructure. Copyright © 2010 John Wiley & Sons, Ltd. Rodrigo N. Calheiros, Rajkumar Buyya, César A. F. De Rose |
Softw. Pract. Exp. | 1 |
| 2009 | A Heuristic for Mapping Virtual Machines and Links in Emulation TestbedsabstractDistributed system emulators provide a paramount platform for testing of network protocols and distributed applications in clusters and networks of workstations. However, to allow testers to benefit from these systems, it is necessary an efficient and automatic mapping of hundreds, or even thousands, of virtual nodes to physical hosts-and the mapping of the virtual links between guests to physical paths in the physical environment. In this paper we present a heuristic to map both virtual machines to hosts and virtual links between virtual machines to paths in the real system. We define the problem we are addressing, present the solution for it and evaluate it in different usage scenarios. Rodrigo N. Calheiros, Rajkumar Buyya, César A. F. De Rose |
ICPP | 1 |
| 2009 | Towards self-managed adaptive emulation of grid environmentsabstractDistributed systems emulators built with the aid of virtualization tools allow testing of systems in a testbed whose number of real elements are orders of magnitude smaller than the number of virtual elements being tested. However, to allow testers to benefit from these systems, operation of the virtual environment should be hidden from them and performed automatically by the emulator. Moreover, testers may be unsure on the exact needs of their environment, and thus can request an environment that does not fit the experiment. In this paper we present our achievements in providing an emulation framework able to provide environment reconfiguration if the requested one does not comply with experiment's demands. Also, it supplies services such as execution log, environment monitoring, and automatic management of applications running in the virtual environment. Rodrigo N. Calheiros, Everton Alexandre, Andriele B. do Carmo, César A. F. De Rose, Rajkumar Buyya |
ISCC | 1 |
| 2008 | Applying Virtualization and System Management in a Cluster to Implement an Automated Emulation Testbed for Grid ApplicationsabstractAlthough grid systems have evolved in such a way that they are largely used both in industry and academy, techniques to test and evaluate them, such as simulation and emulation, have limitations on both their applicability and their reliability. We are investigating the utilization of paravirtualization techniques merged with systems management tools to build an automated emulation framework for grid experiments. This framework accesses standard network resources to manage communication among virtual nodes, allowing virtual machines to behave like a real grid environment. The development of this framework involves the mapping of virtual machines to physical hosts, automatic deployment and management of virtual machines, automatic configuration of virtual network and experiment control. In this paper, we address these issues and present results demonstrating the feasibility and advantages of our approach. Rodrigo N. Calheiros, Mauro Storch, Everton Alexandre, César A. F. De Rose, Marcus Breda |
SBAC-PAD | 1 |
| 2008 | Allocation strategies for utilization of space-shared resources in Bag of Tasks grids
César A. F. De Rose, Tiago Ferreto, Rodrigo N. Calheiros, Walfredo Cirne, Lauro Beltrão Costa, Daniel Fireman |
Future Gener. Comput. Syst. | 3 |
| 2005 | Transparent Resource Allocation to Exploit Idle Cluster Nodes in Computational GridsabstractClusters of workstations are one of the most suitable resources to assist e-scientists in the execution of large-scale experiments that demand processing power. The utilization rate of these machines is usually far from 100%, and hence this should motivate administrators to share their clusters to grid communities. However, exploiting these resources in computational grids is challenging and brings several problems. This paper presents a transparent resource allocation strategy to harness idle cluster resources aimed at executing grid applications. This novel approach does not make use of a formal allocation request to cluster resource managers. Moreover, it does not interfere with local cluster users, being non-intrusive, and hence motivating cluster administrators to publish their resources to grid communities. We present experimental results regarding the effects of the proposed strategy on the attendance time of both cluster and grid requests and we also analyze its effectiveness in clusters with different utilization rates. Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, Rafael K. S. Silva, César A. F. De Rose, Caio Northfleet, Walfredo Cirne |
e-Science | 2 |