César A. F. De Rose

dblp:95/5199 · also César Augusto Fonticielha De Rose · DBLP profile ↗
← Back
68ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-0070-0157ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 2 since 2021Computer networks · 5Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Improving Cost-Performance Efficiency of Scientific Cloud Workflows through GPU Sharing
Matheus M. Costa, Tiago Ferreto, César A. F. De Rose, Odej Kao, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
CLOSER3
2025 DynaMap: A Map Equation-based Parallel Algorithm for Detecting Communities on Dynamic Graphs
abstract
Community detection is a common graph workload used in various domains. With the rapid increase in volumes of data, a lot of research has been done on accelerating different community detection algorithms through parallelization. However, the vast majority of such works focus on static graph structures. In recent years, dynamic graphs have been gaining a lot of attention, since various applications require the ability to change their data and execute new analytics. Most of the work on dynamic community detection has been done on sequential modularity-based approaches. In this paper, we present a Map Equation-based approach to dynamic community detection. The Map Equation is used by the Infomap algorithm and achieves better community structures on static graphs compared to modularity. We design our approach to be easy to parallelize, increasing its applicability to real-world graphs. We show that a parallel implementation of our approach can be faster than modularity-based implementations and as fast as parallel naive ones, with a minimal impact on accuracy, providing a positive impact in the efficiency and efficacy of dynamic graph workflows.
Gabriel G. Dos Santos, Kartik Lakhotia, César A. F. De Rose
SBAC-PAD3
2025 Toward a Dynamic Allocation Strategy for Deadline-Oriented Resource and Job Management in HPC Systems
abstract
ABSTRACT As high‐performance computing (HPC) becomes a tool used in many different workflows, quality of service (QoS) becomes increasingly important. In many cases, this includes the reliable execution of an HPC job and the generation of the results by a certain deadline. The resource and job management system (RJMS) or simply RMS is responsible for receiving the job requests and executing the jobs with a deadline‐oriented policy to support the workflows. In this article, we evaluate how well static resource management policies cope with deadline‐constrained HPC jobs and explore two variations of a dynamic policy in this context. As the Hilbert curve‐based approach used by the SLURM workload manager represents the state‐of‐the‐art in production environments, it was selected as one of the static allocation strategies. The Manhattan median approach as a second allocation strategy was introduced as a research work that aims to minimize the communication overhead of the parallel programs by providing compact partitions more than the Hilbert curve approach. In contrast to the static partitions provided by the Hilbert curve approach and the Manhattan median approach, the leak approach focuses on supporting dynamic runtime behavior of the jobs and assigning nodes of the HPC system on demand at runtime. Since the contiguous leak version also relies on a compact set of nodes, the noncontiguous leak can provide additional nodes at a greater distance from the nodes already used by the job. Our preliminary results clearly show that a dynamic policy is needed to meet the requirements of a modern deadline‐oriented RMS scenario.
Barry Linnert, César A. F. De Rose, Hans-Ulrich Heiß
Concurr. Comput. Pract. Exp.2
2024 Towards a Scalable Parallel Infomap Algorithm for Community Detection
abstract
Identifying Community structures is a fundamental problem in graph analysis. To detect communities in massive contemporary graphs, researchers have extensively explored shared- and distributed-memory parallel algorithms for several methods including Louvain Modularity Optimization and Label Propagation. The widely used Infomap algorithm based on Map Equation Framework (MEF) is known to provide better quality results than other approaches. However, research on parallel community detection using MEF or Infomap is extremely sparse when compared to other methods. We present a comprehensive characterization of Infomap and some of its known parallel implementations to facilitate research into parallel algorithms based on MEF. Most implementations take simple parallelization approaches, leaving strategies used to parallelize similar algorithms such as Louvain untouched. We highlight the scalability limitations of current implementations and implement and eval-uate optimizations for MEF based parallel community detection that achieved up to 119% improvement on the overall speedup across the tested datasets.
Gabriel G. Dos Santos, Kartik Lakhotia, César A. F. De Rose
PDP3
2024 Evaluating machine learning prediction techniques and their impact on proactive resource provisioning for cloud environments
Dionatra F. Kirchoff, Vinícius Meyer, Rodrigo N. Calheiros, César A. F. De Rose
J. Supercomput.4
2023 Message from the General Chairs
abstract
On behalf of the organizing committee, we welcome you to the 35th International Symposium on Computer Architecture and High-Performance Computing (SBAC-PAD 2023), in Porto Alegre, Rio Grande do Sul, Brazil. Since 1987, SBAC-PAD has continuously presented an overview of new developments, applications, and trends in parallel and distributed computing technologies. SBAC-PAD is open to faculty members, researchers, specialists, and graduate students. We have striven to continue this tradition and have considered papers across a whole range of computer system design related topics.
Tiago Ferreto, Dalvan Griebler, César A. F. De Rose
SBAC-PAD3
2022 Estimating the Impact of Communication Schemes for Distributed Graph Processing
abstract
Extreme scale graph analytics is imperative for several real-world Big Data applications with the underlying graph structure containing millions or billions of vertices and edges. Since such huge graphs cannot fit into the memory of a single computer, distributed processing of the graph is required. Several frameworks have been developed for performing graph processing on distributed systems. The frameworks focus primarily on choosing the right computation model and the partitioning scheme under the assumption that such design choices will automatically reduce the communication overheads. For any computational model and partitioning scheme, communication schemes — the data to be communicated and the virtual interconnection network among the nodes — have significant impact on the performance. To analyze this impact, in this work, we identify widely used communication schemes and estimate their performance. Analyzing the trade-offs between the number of compute nodes and communication costs of various schemes on a distributed platform by brute force experimentation can be prohibitively expensive. Thus, our performance estimation models provide an economic way to perform the analyses given the partitions and the communication scheme as input. We validate our model on a local HPC cluster as well as the cloud hosted NSF Chameleon cluster. Using our estimates as well as the actual measurements, we compare the communication schemes and provide conditions under which one scheme should be preferred over the others.
Tian Ye 0002, Sanmukh R. Kuppannagari, César A. F. De Rose, Sasindu Wijeratne, Rajgopal Kannan, Viktor Prasanna 0001
ISPDC3
2022 IntP: Quantifying cross-application interference via system-level instrumentation
abstract
Large-scale container datacenters host tens of thousands of diverse container-wrapped applications each day improving resource usage and maintenance costs. However, resource contention-related interference between co-located applications can severely degrade performance, affecting the quality of service at the user level and compromising experience. Understanding the sources of noise that generates this interference and better managing how to consolidate applications to physical hosts can significantly improve resource usage and overall performance reducing costs for providers and users. This paper presents IntP-an open-source system-level monitoring tool, which analyses selected architectural counters and operating systems data structures to estimate the stress an application puts on each hardware's subsystem and consequently infer the potential interference it could generate in other applications hosted in the same physical machine. Different from state-of-the-art tools that apply a more high-level approach using micro benchmarks and application metrics, IntPs low level instrumentation enables a more accurate prediction of the performance degradation that results from contention on shared resources, with less monitoring overhead. This information can be used to optimize scheduling strategies, which will make datacenter more resource-efficient and cost-effective. To show examples on how to use this tool and validate its results we present three cases studies that applied IntP in their interference-aware methodologies to improve resource utilization in distributed architectures that were able to achieve an increase up to 35% in resource efficiency and up to 25% in user level performance.
Miguel G. Xavier, Carlos H. C. Cano, Vinícius Meyer, César A. F. De Rose
SBAC-PAD4
2022 IADA: A dynamic interference-aware cloud scheduling architecture for latency-sensitive workloads
Vinícius Meyer, Matheus L. da Silva, Dionatra F. Kirchoff, César A. F. De Rose
J. Syst. Softw.4
2021 Resource Sharing and Security Implications on Machine Learning Inference Accelerators
abstract
Due to the increasing adoption of Machine Learning (ML) and in particular Deep Learning (DL), many specialized energy efficient accelerators are being proposed by academia and industry. A number of these accelerators are designed to run a single application at a time in exclusive access mode. This approach gives applications maximum performance but reduces resource efficiency, resulting in increased costs over time. Sharing the device among multiple jobs increases resource utilization and amplifies return on investment. This study is driven by a broad investigation of various spatial resource sharing strategies in machine learning hardware accelerators and performance evaluation in a novel memristor-based accelerator called PUMA [1]. Two methods of spatial sharing are discussed: Model Packing and Logical Allocation. Simulations showed that both methods can be implemented on the PUMA accelerator and have advantages in terms of increased resource utilization. The former spatial sharing strategy achieves higher level of parallelism, fitting more models per device (7 models on 11 tiles), but has higher interference overhead (up to 49%), still being in most cases better than the overhead found for GPUs. The latter spatial sharing strategy achieves better isolation with almost no interference overhead (<1%) with the cost of leaving resources unused (same 7 models consumed 16 tiles). Finally, we discuss security implications of resource sharing for ML and other concerns, presenting a novel ML model integrity check and model bias verification.
Plínio Silveira, César A. F. De Rose, Avelino Francisco Zorzo, Miguel G. Xavier, Dejan S. Milojicic, Sai Rahul Chalamalasetti, Sergey Serebryakov
COMPSAC2
2021 ML-driven classification scheme for dynamic interference-aware resource scheduling in cloud infrastructures
Vinícius Meyer, Dionatra F. Kirchoff, Matheus L. da Silva, César A. F. De Rose
J. Syst. Archit.4
2020 Towards Interference-Aware Dynamic Scheduling in Virtualized Environments
Vinícius Meyer, Uillian L. Ludwig, Miguel G. Xavier, Dionatra F. Kirchoff, César A. F. De Rose
JSSPP5
2020 MemMAP: Compact and Generalizable Meta-LSTM Models for Memory Access Prediction
Ajitesh Srivastava, Ta-Yang Wang, Pengmiao Zhang, César A. F. De Rose, Rajgopal Kannan, Viktor Prasanna 0001
PAKDD (2)4
2020 Modeling and Simulation of QoS-Aware Power Budgeting in Cloud Data Centers
abstract
Power budgeting is a commonly employed solution to reduce the negative consequences of high power consumption of large scale data centers. While various power budgeting techniques and algorithms have been proposed at different levels of data center infrastructures to optimize the power allocation to servers and hosted applications, testing them has been challenging with no available simulation platform that enables such testing for different scenarios and configurations. To facilitate evaluation and comparison of such techniques and algorithms, we introduce a simulation model for Quality-of-Service aware power budgeting and its implementation in CloudSim. We validate the proposed simulation model against a deployment on a real testbed, showcase simulator capabilities, and evaluate its scalability.
Jakub Krzywda, Vinícius Meyer, Miguel G. Xavier, Ahmed Ali-Eldin, Per-Olov Östberg, César A. F. De Rose, Erik Elmroth
PDP6
2020 An Interference-Aware Application Classifier Based on Machine Learning to Improve Scheduling in Clouds
abstract
To maximize resource utilization and system throughput in cloud platforms, hardware resources are often shared across multiple virtualized services or applications. In such a consolidated scenario, performance of applications running concurrently in the same physical host can be negatively affected due to interference caused by resource contention. This should be taken into account for efficient scheduling of such applications and performance prediction at user level. Nevertheless, resource scheduling in cloud computing is usually based solely on resource capacity, implemented by heuristics such as bin-packing. Our previous work has introduced an interference-aware scheduling model for web-applications considering their resource utilization profile, and to classify applications we applied fixed interference intervals based on common utilization patters. Although this resulted in placements with better overall results, we observed that some applications with more dynamic workload patterns were wrongly classified with intervals. In this paper, we propose an alternative to the use of intervals and present an interference-aware application classifier for cloud-based applications that deals better with dynamic workloads. Our classifier defines automatically interference levels ranges combining two well-known machine learning techniques: Support Vector Machines and K-Means. Preliminary experiments evaluated the applied machine learning techniques in three quality metrics: Accuracy, F1-Score and Rand Index, observing rates over 80%. The proposed solution creates a workload-aware fine-grained classification that was compared with previous work over different workload scenarios. The results demonstrate that our classification approach improves the placement efficiency by 23% on average.
Vinícius Meyer, Dionatra F. Kirchoff, Matheus L. da Silva, César A. F. De Rose
PDP4
2020 Evaluating the performance and improving the usability of parallel and distributed Word Embeddings tools
abstract
The representation of words by means of vectors, also called Word Embeddings (WE), has been receiving great attention from the Natural Language Processing (NLP) field. WE models are able to express syntactic and semantic similarities, as well as relationships and contexts of words within a given corpus. Although the most popular implementations of WE algorithms present low scalability, there are new approaches that apply High-Performance Computing (HPC) techniques. This is an opportunity for an analysis of the main differences among the existing implementations, based on performance and scalability metrics. In this paper, we present a study which addresses resource utilization and performance aspects of known WE algorithms found in the literature. To improve scalability and usability we propose a wrapper library for local and remote execution environments that contains a set of optimizations such as the pWord2vec, pWord2vec MPI, Wang2vec and the original Word2vec algorithm. Utilizing these optimizations it is possible to achieve an average performance gain of 15x for multicores and 105x for multinodes compared to the original version. There is also a big reduction in the memory footprint compared to the most popular python versions.
Matheus L. da Silva, Vinícius Meyer, Dionatra F. Kirchoff, Joaquim Francisco Santos Neto, Renata Vieira, César A. F. De Rose
PDP6
2020 RECEIPT: REfine CoarsE-grained IndePendent Tasks for Parallel Tip decomposition of Bipartite Graphs
abstract
Tip decomposition is a crucial kernel for mining dense subgraphs in bipartite networks, with applications in spam detection, analysis of affiliation networks etc. It creates a hierarchy of vertex-induced subgraphs with varying densities determined by the participation of vertices in butterflies (2, 2-bicliques). To build the hierarchy, existing algorithms iteratively follow a delete-update (peeling) process: deleting vertices with the minimum number of butterflies and correspondingly updating the butterfly count of their 2-hop neighbors. The need to explore 2-hop neighborhood renders tip-decomposition computationally very expensive. Furthermore, the inherent sequentiality in peeling only minimum butterfly vertices makes derived parallel algorithms prone to heavy synchronization. In this paper, we propose a novel parallel tip-decomposition algorithm - REfine CoarsE-grained Independent Tasks (RECEIPT) that relaxes the peeling order restrictions by partitioning the vertices into multiple independent subsets that can be concurrently peeled. This enables RECEIPT to simultaneously achieve a high degree of parallelism and dramatic reduction in synchronizations. Further, RECEIPT employs a hybrid peeling strategy along with other optimizations that drastically reduce the amount of wedge exploration and execution time. We perform detailed experimental evaluation of RECEIPT on a shared-memory multicore server. It can process some of the largest publicly available bipartite datasets orders of magnitude faster than the state-of-the-art algorithms - achieving up to 1100× and 64× reduction in the number of thread synchronizations and traversed wedges, respectively. Using 36 threads, RECEIPT can provide up to 17.1× self-relative speedup.
Kartik Lakhotia, Rajgopal Kannan, Viktor Prasanna 0001, César A. F. De Rose
Proc. VLDB Endow.4
2020 Orthogonal persistence in nonvolatile memory architectures: A persistent heap design and its implementation for a Java Virtual Machine
abstract
Summary Current computer systems separate main memory from storage, and programming languages typically reflect this distinction using different representations for data in memory and storage. However, moving data back and forth between these different layers and representations compromise both programming and execution efficiency. To remedy this, the concept of orthogonal persistence (OP) was proposed in the early 1980s advocating that, from a programmer's standpoint, there should be no differences in the way that short‐term and long‐term data are manipulated. However, at that time, the underlying implementations still had to cope with the complexity of moving data across memory and storage. Today, recent nonvolatile memory (NVM) technologies, such as resistive RAM and phase‐change memory, allow main memory and storage to be collapsed into a single layer of persistent memory, opening the way for more efficient programming abstractions for handling persistence. In this work, we revisit OP concepts in the context of NVM architectures and propose a persistent heap design for languages with automatic memory management. We demonstrate how it can significantly increase programmer and execution efficiency, removing the impedance mismatch of crossing semantic boundaries. To validate and demonstrate the presented concepts, we present JaphaVM, an implementation of the proposed design based on JamVM, an open‐source Java Virtual Machine. Our results show that JaphaVM, in most cases, executes the same operations between one and two orders of magnitude faster than regular database‐based and file‐based implementations, while requiring significantly less lines of code.
Taciano Perez, Marcelo Veiga Neves, Diego Medaglia, Pedro H. G. Monteiro, César A. F. De Rose
Softw. Pract. Exp.5
2019 Performance and Cost Analysis between Elasticity Strategies over Pipeline-structured Applications
abstract
With the advances in eScience-related areas and the growing complexity of scientific analysis, more and more scientists are interested in workflow systems. There is a class of scientific workflows that has become a standard to stream processing, called pipeline-structured application. Due to the amount of data these applications need to process nowadays, Cloud Computing has been explored in order to accelerate such processing. However, applying elasticity over stage-dependent applications is not a trivial task since there are some issues that must be taken into consideration, such as the workload proportionality among stages and ensuring the processing flow. There are studies which explore elasticity on pipeline-structured applications but none of them compare or adopt different strategies. In this paper, we present a comparison between two elasticity approaches which consider not only CPU load but also workload processing time information to reorganize resources. We have conducted a number of experiments in order to evaluate the performance gain and cost reduction when applied our strategies. As results, we have reached an average of 72% in performance gain and 73% in cost reduction when comparing non-elastic and elastic executions.
Vinícius Meyer, Miguel G. Xavier, Dionatra F. Kirchoff, Rodrigo da Rosa Righi, César A. F. De Rose
CLOSER5
2019 A Preliminary Study of Machine Learning Workload Prediction Techniques for Cloud Applications
abstract
Cloud computing has transformed the means of computing in recent years with several benefits over traditional systems, like scalability and high availability. However, there are still some opportunities, especially in the area of resource provisioning and scaling [13]. Since workload may fluctuate a lot in certain environments, over-provisioning is a common practice to avoid abrupt Quality of Service (QoS) drops that may result in Service Level Agreement (SLA) violations, but at the price of an increase in provisioning costs and energy consumption. Workload prediction is one of the strategies by which efficiency and operational cost of a cloud can be improved [13]. Knowing demand in advance allows the previous allocation of sufficient resources to maintain QoS and avoid SLA violations [1]. This paper presents the advantages and disadvantages of three workload prediction techniques when applied in the context of cloud computing. Our preliminary results compare ARIMA, MLP, and GRU under different cloud configurations to help administrators choose the more appropriate and efficient predictive model for their specific problem.
Dionatra F. Kirchoff, Miguel G. Xavier, Juliana Mastella, César A. F. De Rose
PDP4
2019 Optimizing multi-tier application performance with interference and affinity-aware placement algorithms
abstract
Summary Cloud providers are constantly seeking to become more cost effective, where a common strategy is to consolidate multiple applications in physical machines using virtualization techniques. This consolidation, however, may result in performance related problems such as resource interference. Moreover, if the workload is composed of multi‐tier applications, an increasingly popular method of application development, especially for web and mobile, in which tiers need to communicate through the network, we have another possible source of performance degradation, which we refer as network affinity. In order to reduce the effects of such problems, placement techniques are used to better distribute the applications in the physical machines. Several of these placement techniques consider resource interference or network affinity in order to decide the best placement, however, none of them apply both criteria at the same time. In our previous work, we identified that a combined approach could result in better solutions for this problem and proposed a set of placement policies that explore this tradeoff. In this paper, we propose placement algorithms based on these policies and evaluate the proposed solutions for different workload scenarios using a visual simulation tool we developed called CIAPA. CIAPA introduces a performance degradation model, a cost function, and heuristics to find a placement with the minimum cost for a specific workload of multi‐tier applications. In our preliminary experiments, we compared the solution generated by CIAPA with other placement strategies from related work, and have verified that, for the tested scenarios, it delivers placement decisions with better cost and, consequently, improved performance. We observed a reduction in response time of 10% when compared to interference strategies and up to 18% when considering only affinity strategies.
Uillian L. Ludwig, Miguel G. Xavier, Dionatra F. Kirchoff, Ian B. Cezar, César A. F. De Rose
Concurr. Comput. Pract. Exp.5
2019 Foreword to the special issue of the workshop on high performance computing systems (XVIII Simpósio em Sistemas Computacionais de Alto Desempenho, WSCAD 2017)
abstract
This special issue of Concurrency and Computation Practice and Experience gathers extended versions of six selected research articles that were previously presented at the Brazilian Workshop on High Performance Computing Systems (“XVIII Simpósio em Sistemas Computacionais de Alto Desempenho”, WSCAD 2017), held in conjunction with the 29th International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2017, in Campinas, SP, Brazil, from the 17th to the 20th of October 2017. Since 2000, this workshop has presented important and interesting research in the fields of Computer Architecture, High Performance Computing and Distributed Systems. The scope of the current special issue is broad and representative of the multidisciplinary nature of High Performance Computing and Computer Architecture research domains. The set of accepted research articles was organized under three key themes: Parallel Algorithms and Optimizations, Scheduling and Placement, and Parallel Architecture Design. In the following sections, we provide a brief description of each one of the research articles accepted in this special issue. To achieve the best performance possible, algorithms must be carefully parallelized and optimized for multi-core or many-core processors. Current multi-core and many-core architectures may feature different technologies, such as distributed memory banks, vector instructions, and specialized cores. Oftentimes, different classes of multi-core and many-core processors are combined to construct a heterogeneous multiprocessing system. Today's technologies enable heterogeneous multiprocessing systems on a chip containing multi-cores, many-cores (eg, GPUs), and FPGAs. The following research articles present parallel solutions and optimizations to different classes of algorithms on multi-core and many-core processors. The paper “Optimized implementation of QC-MDPC code-based cryptography” presents a new enhanced version of the QcBits key encapsulation mechanism (KEM), which is a constant time implementation of the Niederreiter cryptosystem using QC-MDPC codes.1 The parallel solution uses vector instructions (AVX 512) and applies several other techniques to achieve a competitive performance level. The enhanced version is 1.9x faster when decrypting messages when compared with BIKE, which was the state-of-the-art implementation for QC-MDPC codes. The paper “On the Parallelization of Hirschberg's Algorithm for Multi-core and Many-core Systems” focuses on improving the execution efficiency of Hirschberg's algorithm, which aims at finding the longest common subsequence between two strings, on multi-core and many-core systems.2 The proposed solution exploits vector instructions and different parallelization strategies to achieve the best performance possible. Results showed that the parallel solution can achieve speedups of up to 15.5x on a 18-core Xeon processor and of up to 105x on a 68-core Intel Xeon Phi many-core processor. Finally, the paper “A Hybrid CPU-GPU-MIC Algorithm for Minimal Hitting Set Enumeration” proposes a hybrid exact algorithm for the Minimal Hitting Set (MHS) Enumeration Problem for highly heterogeneous platforms.3 The experiments were carried out on heterogeneous platforms composed of Intel Xeon E5-2620v2 CPUs, Intel Xeon Phi 3120A, and a GTX TITAN X GPUs. The results showed that the proposed algorithm was able to distribute parallel tasks among the processing units according to their computational efficiency in processing the task batches, achieving speedups of up to 25.3x in comparison with using two Intel Xeon E5-2620v2 CPUs. To deliver high performance to large-scale engineering and scientific applications, particular intricacies of the application and the underlying platform should be considered, so that tailored techniques can be employed to map one into another. In this context, evenly distributing the workload of an application among its threads, processes, or virtual machines is an NP-Hard minimization problem known as scheduling, and allocating these work abstractions to the underlying infrastructure is called placement. These problems are significant to the academic community and industry, and they are a hot research topic in High Performance Computing (HPC). The following research articles present contributions to these problems for different abstraction levels and domains. The paper “A Comprehensive Performance Evaluation of the BinLPT Workload-Aware Loop Scheduler” focuses on improving a workload-aware scheduling strategy called BinLPT, previously proposed by the same authors.4 Two new contributions are presented to the state of the art. First, a multiloop support feature was introduced to BinLPT, which enables the reuse of workload estimations across loops. Based on this feature, BinLPT was integrated into a real-world elastodynamics application and evaluated running on a supercomputer. Second, BinLPT was evaluated using simulations as well as synthetic and application kernels. This analysis was carried out on a large-scale NUMA machine under a variety of workloads. The results revealed that BinLPT is able to balance the load of irregular OpenMP parallel loops among the application threads, delivering up to 37% and 9% better performance than well-known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively. Finally, the paper “Optimizing the Performance of Multi-tier Applications Using Interference and Affinity-aware Placement Algorithms” proposes a combined approach that considers both resource interference and network affinity to decide the best placement of multi-tier applications in consolidated environments.5 In their previous work, the same authors identified that a combined approach could result in better solutions for this problem and proposed a set of placement policies that explore this tradeoff. The authors propose a new family of placement algorithms based on these policies and evaluated for different workload scenarios using a visual simulation tool called CIAPA. CIAPA introduces a performance degradation model, a cost function, and heuristics to find a placement with the minimum cost for a specific workload of multi-tier applications. The solution generated by CIAPA was compared to other placement strategies from related work, and delivered placement decisions with better cost, and, consequently, improved performance. An average reduction in response time of 10% was observed when compared to interference strategies, and up to 18% when considering only affinity strategies. Dataflow-based FPGA accelerators have become a promising alternative to deliver energy efficient platforms for the HPC domain. However, FPGA programming is still a challenge. Although reconfigurable FPGA technologies have been around since the 1980s, their utilization as a general-purpose processing platform is recent. Historically, both FPGA and ASIC developers have employed Hardware Description Languages (HDLs) to implement their designs, which is usually outside the main expertise area of software developers. The lack of simple and common programming models prevents software developers from easily designing accelerators and delays a broader adoption of this technology. In this context, the paper “ADD: Accelerator Design and Deploy - A Tool for FPGA High Performance Dataflow Computing” presents a high-level framework to specify, to simulate, and to implement dataflow accelerators for streaming applications.6 The Accelerator Design and Deploy (ADD) framework includes an open dataflow operator library, and templates are provided to easily design new operators. The framework also provides a high-level and an accurate simulation at circuit level with short execution times. Moreover, ADD provides software and hardware APIs to simplify the integration process, extending the benefits of portability from low-cost FPGA boards to high performance datacenter FPGA platforms. The framework supports coupling with high-level programming languages, and it has been validated on two FPGA platforms: the Intel high-performance CPU-FPGA heterogeneous computing platform and an educational FPGA kit. The authors show that the proposed approach presents competitive performance, both in time and energy, when compared to multi-core and GPU accelerators. Concerning energy, it is 18.8x and 193.2x more efficient than the GPU and multi-core evaluated platforms, respectively. The research articles presented in this special issue provide insights in fields related to High Performance Computing, including Parallel Algorithms and Optimizations, Scheduling and Placement, and Parallel Architecture Design. We believe that the main contributions presented in the research articles are timely and important. We hope that readers can benefit from insights of these research articles and contribute to these rapidly growing areas. Dr. César A. F. De Rose has a B.Sc. degree in Computer Science from the Pontifical Catholic University of Rio Grande do Sul (PUCRS, Porto Alegre, Brazil, 1990), an M.Sc. in Computer Science from the Federal University of Rio Grande do Sul (PGCC/UFRGS, Porto Alegre, Brazil, 1993), and a Doctoral degree from Karlsruhe Institute of technology (KIT - Karlsruhe, Germany, 1998). In 1998, he joined the Faculty of Informatics at PUCRS as an associate professor and member of the Resource Management and Virtualization Group (full professor since 2012). His research interests include resource management, dynamic provisioning and allocation, monitoring techniques (resource and application), application modeling, scheduling and optimization in parallel and distributed environments (Cluster, Grid, Cloud), and virtualization. In 2009, he founded PUCRS High Performance Computing Laboratory (LAD-PUCRS) being nowadays senior researcher. Dr. Márcio Castro received a B.Sc. in Computer Science with honors (Summa Cum Laude) from Pontifical Catholic University of Rio Grande do Sul (PUCRS, Brazil) in 2006 and an M.Sc. degree in Computer Science from the same university in 2009. He received a Ph.D. in Computer Science in 2012 from the University of Grenoble Alpes, France. Then, he worked as a postdoctoral fellow at the Federal University of Rio Grande do Sul (UFRGS), Brazil. Since 2014, he is an associate professor at the Federal University of Santa Catarina (UFSC), Brazil. His main research area is High Performance Computing, with focus on parallel programming models, load balancing, high performance parallel applications, and parallel and distributed computing on multi-core and many-core architectures. We would like to thank all the authors who provided valuable contributions to this special issue. We are also grateful to the reviewers for their feedback to the authors. Indeed, their advices were essential to further improve the quality of the papers. Finally, we would like to express our sincere gratitude to Professor Geoffrey Fox, the Editor in Chief, for providing us with this unique opportunity to present the selected papers from WSCAD 2017 in the International Journal of Concurrency and Computation: Practice and Experience.
César A. F. De Rose, Márcio Castro 0001
Concurr. Comput. Pract. Exp.1
2017 Cloud Storage Cost Modeling for Cryptographic File Systems
abstract
Nowadays, data security is a demand for companies when adopting storage services on public clouds. From long term persistence services, such as Amazon Glacier, to online block storage systems for virtual machines disks, security principles can be part of the cloud context, especially for customer's sensitive data. The confidentiality of storage services considers aspects such as data life-cycle, location, and size, besides that, this principle is often provided by a cryptography mechanism applied in one of the persistence layers, such as in the File-System (FS). However, to add cryptography for data security demands extra CPU cycles for ciphering the data during its persistence. Although these extra CPU cycles are not considered on current cloud costs estimations, it should be part of the total application execution's costs. This paper presents the architectures for Cryptography File Systems (CFS) adoption for data storing in cloud computing. Furthermore, a mathematical model is presented and discussed as an estimation tool of cryptography overhead when using CFSs in the cloud storage stack. The model is verified in a real scenario for estimating the total cost when adding security for storage in a cloud environment. As main result, the model could estimate the overhead within 90% to 92% of accuracy for the AES algorithm, according to real cases traces, considering available memory, {I/O} throughput and workload size.
Mauro Storch, César A. F. De Rose
PDP2
2017 Mobile Application Testing on Clouds: Challenges, Opportunities and Architectural Elements
abstract
With mobility increasing more each year, mobile devices and operating system (OS) fragmentation are increasing at an even faster pace. A multitude of screen sizes, network connection types, and OS versions have emerged in the market and led mobile developers to rethink testing practices to ensure quality and a good experience for increasingly demanding users who crave highly reliable and stable applications. Such best practices come at the high cost for testing infrastructure and maintenance, making most development teams bypass test cycles and deliver applications before they are thoroughly validated. And when teams pursue automated test cycles on clouds, expenses due to high-cost services are not always worth the investment in application development phases. As a result, test cycles which are essential to validate application reliability and stability, such as regression, functional, and leakage tests are left out, primarily affecting user experience. This paper shows the challenges intrinsic to mobile application testing in consonance with the opportunities provided by clouds. We explored the state of the art to synthesize current cloud-based mobile application testing architectures to convey the need for a new concept and platform to minimize maintenance and make testing infrastructures more cost-effective. Hence, we proposed an alternative architecture using emulated devices for testing automation, which aims for massive test cycles at a lower cost.
Miguel G. Xavier, Kassiano J. Matteussi, Gabriel R. Franca, Wagner P. Pereira, César A. F. De Rose
PDP5
2017 Modeling and simulation of global and sleep states in ACPI-compliant energy-efficient cloud environments
abstract
Summary The more large‐scale data centers infrastructure costs increase, the more simulation‐based evaluations are needed to understand better the trade‐off between energy and performance and support the development of new energy‐aware resource allocation policies. Specifically, in the cloud computing field, various simulators are able to predict and measure the behavior of applications on different architectures using different resource allocation policies. Yet, only a few of them have the ability to simulate energy‐saving strategies, and none of them support the complete advanced configuration and power interface (ACPI) specification. ACPI defines a terminology for all possible power states of a machine and their associated power rate. The hardware industry has relied on ACPI to provide up‐to‐date standard interfaces for hardware discovery, configuration, power management, and monitoring, enabling a better understanding of the energy consumption level of different hardware states, referred to as ACPI G‐states, S‐states, and P‐states. In this paper, we improve the modeling and simulation of the ACPI G/S‐states and show not only that these states offer different energy‐saving levels but also that state transitions consume energy. In addition, we model the latency to transit between two states and the effects on the turnaround time when the transitions are not performed conservatively. Furthermore, the equations provide essential information to quantify the trade‐off between energy consumption and performance and assist in the analysis/decision on which strategy fits better in the environment and how it could be refined. Our expanded energy model was implemented in CloudSim and validated with simulation‐based experiments with a very high level of accuracy, with a standard deviation of at most 6%. Copyright © 2016 John Wiley & Sons, Ltd.
Miguel G. Xavier, Fábio D. Rossi, César A. F. De Rose, Rodrigo N. Calheiros, Danielo Goncalves Gomes
Concurr. Comput. Pract. Exp.3
2017 E-eco: Performance-aware energy-efficient cloud data center orchestration
Fábio D. Rossi, Miguel G. Xavier, César A. F. De Rose, Rodrigo N. Calheiros, Rajkumar Buyya
J. Netw. Comput. Appl.3
2016 Understanding performance interference in multi-tenant cloud databases and web applications
abstract
The number of e-commerce customers and database services in cloud computing platforms has grown increasingly, leading providers to adopt resource-sharing solutions to meet growing demand for infrastructure resources, such as processing and storage. Consolidating database applications has become arguably a de-facto solution to support a large number of customers/tenants at low infrastructure costs. However, the friction generated in shared hardwares (resource contention) is converted to performance interference, which is felt by tenants' database applications running on upper layers (VMs). Hence, there is a real concern on how to manage and prevent multi-tenant cloud databases from performance interferences sourced by either resource contention or isolation flaws. In this paper, we analyzed the performance interference tolerated by multi-tenant e-commerce cloud databases in resource-sharing infrastructures. We claimed that multiple-different workloads (e.g. memory-/CPU-intensive, and e-commerce applications) might be consolidated with database systems to minimize performance interference and increase resource-efficiency.
Miguel G. Xavier, Kassiano J. Matteussi, Fabian Lorenzo, César A. F. De Rose
IEEE BigData4
2015 Modeling power consumption for DVFS policies
abstract
Power-aware management strategies are a trend towards achieving energy-efficient computing environments. One of the approaches behind those strategies is dynamic frequency and voltage scaling (DVFS). Since frequency adjustments may have a negative impact on system performance, users often have to experiment with these policies to find the optimal configuration for their application and energy reduction goals. While the performance impact can be easily measured by the total execution time of an application, power consumption measurements require additional logging and frequently external equipment. The following paper presents a mathematical model to help users estimate the power consumption of their application when using different DVFS policies. A preliminary evaluation shows that the model has 94% accuracy when compared against real-time measurements.
Fábio D. Rossi, Mauro Storch, Israel C. De Oliveira, César A. F. De Rose
ISCAS4
2015 MRemu: An Emulation-Based Framework for Datacenter Network Experimentation Using Realistic MapReduce Traffic
abstract
As data volumes and the need for timely analysis grow, Big Data analytics frameworks have to scale out to hundred or even thousands of commodity servers. While such a scale-out is crucial to sustain desired computational throughput/latency and storage capacity, it comes at the cost of increased network traffic volumes and multiplicity of traffic patterns. Despite the sheer reality of the dependency between datacenter network (DCN) and time-to-insight through big data analysis, our experience as active networking researchers conveys that a large fraction of DCN research experimentation is conducted on network traces and/or synthetic flow traces. And while the respective results are often valuable as standalone contributions, in practice it turns out extremely difficult to quantitatively assess how the reported network optimization results translate to performance or fault-tolerance improvement for actual analytics runtimes, e.g., due to the ability of these runtimes to overlap communication with computation. This paper presents MRemu, an emulation-based framework for conducting reproducible datacenter network research using accurate MapReduce workloads and at system scales that are relevant to the size of target deployments, albeit without requiring access to a hardware infrastructure of such scale. We choose the MapReduce (MR) framework as a design point, for it is a common representative of the most widely deployed frameworks for analysis of large volumes of - structured and unstructured - data and is reported to be highly sensitive to network performance. With MRemu, it is possible to quantify the impact of various network design parameters and software-defined control techniques to key performance indicators of a given MR application. We show through targeted experimental validation that MRemu exhibits high fidelity, when compared to the performance of MR applications on a real scale-out cluster of 16 high-end servers.
Marcelo Veiga Neves, César A. F. De Rose, Kostas Katrinis
MASCOTS2
2015 On the Impact of Energy-Efficient Strategies in HPC Clusters
abstract
Energy-aware management strategies are a recent trend towards achieving energy-efficient computing in HPC clusters. One of the approaches behind those strategies is to apply energy-saving states on idle nodes, alternating them among different sleep states that reflect on many power consumption levels. This paper investigated the way such energy-efficient strategies affected the job turnaround time - the elapsed time between when the job is submitted and when the job is completed, including the wait time as well as the job's actual execution time - in these clusters. Based on the results we proposed a Best-Fit Energy-Aware Strategy that switches the nodes to a sleep state, depending on the throughput of the resource manager's job queue. We simulated the proposed strategy using the SimGrid simulator. Our preliminary results showed a reduction of up to 19% in the overall energy consumption and give us a better understanding of the trade-offs involved in using energy-efficient strategies.
Fábio D. Rossi, Miguel G. Xavier, Yuri J. Monti, César A. F. De Rose
PDP4
2015 A Performance Isolation Analysis of Disk-Intensive Workloads on Container-Based Clouds
abstract
The popularity of Cloud computing due to the increasing number of customers has led Cloud providers to adopt resource-sharing solutions to meet growing demand for infrastructure resources. As the adoption of resource-sharing/consolidation in Cloud computing became arguably a well-established solution, the ability the underlying virtualization systems of preventing performance interferences from customers must also be understood. Virtualization systems based on containers, such as LXC, are the basis of the next-generation of Cloud computing and have become the most popular solution under PaaS/IaaS Cloud platforms with the rise of Docker -- an open platform for developers and sysadmins to build, ship, and run distributed applications. Such platforms have enticed many attentions globally, since they leverage container-based virtualization systems to offer high scalability while low performance overheads, the performance might be solely aggravated if the customers' workloads are consolidated onto the same hardware and the isolation layer does not properly isolate the shared resources. Performance isolation is an inherent concern of such systems due to the nature as they are conceived and is still an unexplored and open research topic, the consequences might influence in the adoption under shared Cloud computing platforms where Quality-of-Service is a crucial factor that cannot be disregarded. In this paper we analyze the performance interference suffered by disk-intensive workloads within very noisy-perturbed containers (different hardware components stressed). Our results show workload combinations whose performance degradation goes up to 38%, but in contrast we expose a workload-balanced scenario wherein the performance does not suffer any interference.
Miguel G. Xavier, Israel C. De Oliveira, Fábio D. Rossi, Robson D. Dos Passos, Kassiano J. Matteussi, César A. F. De Rose
PDP6
2014 Pythia: Faster Big Data in Motion through Predictive Software-Defined Network Optimization at Runtime
abstract
The rise of Internet of Things sensors, social networking and mobile devices has led to an explosion of available data. Gaining insights into this data has led to the area of Big Data analytics. The MapReduce framework, as implemented in Hadoop, is one of the most popular frameworks for Big Data analysis. To handle the ever-increasing data size, Hadoop is a scalable framework that allows dedicated, seemingly unbound numbers of servers to participate in the analytics process. Response time of an analytics request is an important factor for time to value/insights. While the compute and disk I/O requirements can be scaled with the number of servers, scaling the system leads to increased network traffic. Arguably, the communication-heavy phase of MapReduce contributes significantly to the overall response time, the problem is further aggravated, if communication patterns are heavily skewed, as is not uncommon in many MapReduce workloads. In this paper we present a system that reduces the skew impact by transparently predicting data communication volume at runtime and mapping the many end-to-end flows among the various processes to the underlying network, using emerging software-defined networking technologies to avoid hotspots in the network. Dependent on the network oversubscription ratio, we demonstrate reduction in job completion time between 3% and 46% for popular MapReduce benchmarks like Sort and Nutch.
Marcelo Veiga Neves, César A. F. De Rose, Kostas Katrinis, Hubertus Franke
IPDPS2
2014 Green software development for multi-core architectures
abstract
Advances in computer architecture to provide higher parallelism (e.g. hyper threading and multi-core) usually incur in higher complexity in software development. Applications should be designed to use efficiently the additional resources in order to improve its performance. However, the popularity of mobile devices and recent studies in IT-related energy consumption have driven software developers to focus also on energy efficiency. Besides improving applications' performance, software developers should aim at minimizing the amount of energy consumed by the applications. Energy saving becomes an important non-functional requirement for new applications. This paper evaluates the behavior of applications on multi-core architectures and proposes energy-saving alternatives for software development.
Fábio D. Rossi, Miguel G. Xavier, Endrigo D'Agostini Conte, Tiago Ferreto, César A. F. De Rose
ISCC5
2014 A Performance Comparison of Container-Based Virtualization Systems for MapReduce Clusters
abstract
Virtualization as a platform for resource-intensive applications, such as MapReduce (MR), has been the subject of many studies in the last years, as it has brought benefits such as better manageability, overall resource utilization, security and scalability. Nevertheless, because of the performance overheads, virtualization has traditionally been avoided in computing environments where performance is a critical factor. In this context, container-based virtualization can be considered a lightweight alternative to the traditional hypervisor-based virtualization systems. In fact, there is a trend towards using containers in MR clusters in order to provide resource sharing and performance isolation (e.g., Mesos and YARN). However, there are still no studies evaluating the performance overhead of the current container-based systems and their ability to provide performance isolation when running MR applications. In this work, we conducted experiments to effectively compare and contrast the current container-based systems (Linux VServer, OpenVZ and Linux Containers (LXC)) in terms of performance and manageability when running on MR clusters. Our results showed that although all container-based systems reach a near-native performance for MapReduce workloads, LXC is the one that offers the best relationship between performance and management capabilities (specially regarding to performance isolation).
Miguel G. Xavier, Marcelo Veiga Neves, César A. F. De Rose
PDP3
2013 Optimizing the management of a database in a virtual environment
abstract
Recent studies have demonstrated advantages in using Data Base Management System (DBMS) in virtual environments, like the consolidation of several DBMS isolated by virtual machines on a single physical machine to reduce maintenance costs and energy consumption. Furthermore, live migration can improve database availability, allowing transparent maintenance operations on host machines. However, there are issues that still need to be addressed, like overall performance degradation of the DBMS when running in virtual environments and connections instabilities during a live migration. In this context, new virtualization techniques are emerging, like the virtual database, which is considered a less intrusive alternative for the traditional database virtualization over virtual machines. This paper analyzes aspects of this new virtualization approach, like performance and connection stability during a database migration process and its isolation capabilities. Our evaluation shows very promising results compared to the traditional approach over virtual machines, including a more efficient and stable live migration, maintaining the required isolation characteristics for a virtualized DBMS.
Timoteo Lange, Paolo Cemim, Miguel G. Xavier, César A. F. De Rose
ISCC4
2013 Performance Evaluation of Container-Based Virtualization for High Performance Computing Environments
abstract
The use of virtualization technologies in high performance computing (HPC) environments has traditionally been avoided due to their inherent performance overhead. However, with the rise of container-based virtualization implementations, such as Linux VServer, OpenVZ and Linux Containers (LXC), it is possible to obtain a very low overhead leading to near-native performance. In this work, we conducted a number of experiments in order to perform an in-depth performance evaluation of container-based virtualization for HPC. We also evaluated the trade-off between performance and isolation in container-based virtualization systems and compared them with Xen, which is a representative of the traditional hypervisor-based virtualization systems used today.
Miguel G. Xavier, Marcelo Veiga Neves, Fábio D. Rossi, Tiago Ferreto, Timoteo Lange, César A. F. De Rose
PDP6
2013 EMUSIM: an integrated emulation and simulation environment for modeling, evaluation, and validation of performance of Cloud computing applications
abstract
SUMMARY Cloud computing allows the deployment and delivery of application services for users worldwide. Software as a Service providers with limited upfront budget can take advantage of Cloud computing and lease the required capacity in a pay‐as‐you‐go basis, which also enables flexible and dynamic resource allocation according to service demand. One key challenge potential Cloud customers have before renting resources is to know how their services will behave in a set of resources and the costs involved when growing and shrinking their resource pool. Most of the studies in this area rely on simulation‐based experiments, which consider simplified modeling of applications and computing environment. In order to better predict service's behavior on Cloud platforms, we developed an integrated architecture that is based on both simulation and emulation. The proposed architecture, named EMUSIM, automatically extracts information from application behavior via emulation and then uses this information to generate the corresponding simulation model. We performed experiments using an image processing application as a case study and found that EMUSIM was able to accurately model such application via emulation and use the model to supply information about its potential performance in a Cloud provider. We also discuss our experience using EMUSIM for deploying applications in a real public Cloud provider. EMUSIM is based on an open source software stack and therefore it can be extended for analysis behavior of several other applications. Copyright © 2012 John Wiley & Sons, Ltd.
Rodrigo N. Calheiros, Marco Aurélio Stelmar Netto, César A. F. De Rose, Rajkumar Buyya
Softw. Pract. Exp.3
2012 CASViD: Application Level Monitoring for SLA Violation Detection in Clouds
abstract
Cloud resources and services are offered based on Service Level Agreements (SLAs) that state usage terms and penalties in case of violations. Although, there is a large body of work in the area of SLA provisioning and monitoring at infrastructure and platform layers, SLAs are usually assumed to be guaranteed at the application layer. However, application monitoring is a challenging task due to monitored metrics of the platform or infrastructure layer that cannot be easily mapped to the required metrics at the application layer. Sophisticated SLA monitoring among those layers to avoid costly SLA penalties and maximize the provider profit is still an open research challenge. This paper proposes an application monitoring architecture named CASViD, which stands for Cloud Application SLA Violation Detection architecture. CASViD architecture monitors and detects SLA violations at the application layer, and includes tools for resource allocation, scheduling, and deployment. Different from most of the existing monitoring architectures, CASViD focuses on application level monitoring, which is relevant when multiple customers share the same resources in a Cloud environment. We evaluate our architecture in a real Cloud testbed using applications that exhibit heterogeneous behaviors in order to investigate the effective measurement intervals for efficient monitoring of different application types. The achieved results show that our architecture, with low intrusion level, is able to monitor, detect SLA violations, and suggest effective measurement intervals for various workloads.
Vincent C. Emeakaroha, Tiago Ferreto, Marco Aurélio Stelmar Netto, Ivona Brandic, César A. F. De Rose
COMPSAC5
2012 Scheduling MapReduce Jobs in HPC Clusters
Marcelo Veiga Neves, Tiago Ferreto, César A. F. De Rose
Euro-Par3
2012 Towards autonomic detection of SLA violations in Cloud infrastructures
Vincent C. Emeakaroha, Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, Ivona Brandic, Rajkumar Buyya, César A. F. De Rose
Future Gener. Comput. Syst.6
2012 Performance evaluation of OpenMP-based algorithms for handling Kronecker descriptors
Antonio M. de Lima, Marco Aurélio Stelmar Netto, Thais Webber, Ricardo M. Czekster, César A. F. De Rose, Paulo Fernandes 0001
J. Parallel Distributed Comput.5
2011 Maximum Migration Time Guarantees in Dynamic Server Consolidation for Virtualized Data Centers
Tiago Ferreto, César A. F. De Rose, Hans-Ulrich Heiß
Euro-Par (1)2
2011 Server consolidation with migration control for virtualized data centers
Tiago Ferreto, Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, César A. F. De Rose
Future Gener. Comput. Syst.4
2011 CloudSim: a toolkit for modeling and simulation of cloud computing environments and evaluation of resource provisioning algorithms
abstract
Abstract Cloud computing is a recent advancement wherein IT infrastructure and applications are provided as ‘services’ to end‐users under a usage‐based payment model. It can leverage virtualized services even on the fly based on requirements (workload patterns and QoS) varying with time. The application services hosted under Cloud computing model have complex provisioning, composition, configuration, and deployment requirements. Evaluating the performance of Cloud provisioning policies, application workload models, and resources performance models in a repeatable manner under varying system and user configurations and requirements is difficult to achieve. To overcome this challenge, we propose CloudSim: an extensible simulation toolkit that enables modeling and simulation of Cloud computing systems and application provisioning environments. The CloudSim toolkit supports both system and behavior modeling of Cloud system components such as data centers, virtual machines (VMs) and resource provisioning policies. It implements generic application provisioning techniques that can be extended with ease and limited effort. Currently, it supports modeling and simulation of Cloud computing environments consisting of both single and inter‐networked clouds (federation of clouds). Moreover, it exposes custom interfaces for implementing policies and provisioning techniques for allocation of VMs under inter‐networked Cloud computing scenarios. Several researchers from organizations, such as HP Labs in U.S.A., are using CloudSim in their investigation on Cloud resource provisioning and energy‐efficient management of data center resources. The usefulness of CloudSim is demonstrated by a case study involving dynamic provisioning of application services in the hybrid federated clouds environment. The result of this case study proves that the federated Cloud computing model significantly improves the application QoS requirements under fluctuating resource and service demand patterns. Copyright © 2010 John Wiley & Sons, Ltd.
Rodrigo N. Calheiros, Rajiv Ranjan 0001, Anton Beloglazov, César A. F. De Rose, Rajkumar Buyya
Softw. Pract. Exp.4
2010 Building an automated and self-configurable emulation testbed for grid applications
abstract
Abstract Distributed systems, such as grids, are composed of geographically distributed computing elements that belong to multiple administrative domains and are controlled by multiple entities. It is unlikely that testers are able to acquire repeatedly the same resources, for the same amount of time, and under the same network conditions, which are paramount requirements for enabling reproducible and controlled tests in software under development. An alternative to experiments in real testbeds is the use of emulation tools, which allow the software to run in an environment that behaves like a distributed system. Although advances in virtualization technology allowed the development of efficient emulators, few efforts were put in making operation of such emulators easier. This paper presents the design and the development of the Automated Emulation Framework that allows automatic mapping of virtual machines to hosts, virtual machine deployment, network configuration, and proactive management and reconfiguration of the virtual infrastructure. Copyright © 2010 John Wiley & Sons, Ltd.
Rodrigo N. Calheiros, Rajkumar Buyya, César A. F. De Rose
Softw. Pract. Exp.3
2009 A Heuristic for Mapping Virtual Machines and Links in Emulation Testbeds
abstract
Distributed system emulators provide a paramount platform for testing of network protocols and distributed applications in clusters and networks of workstations. However, to allow testers to benefit from these systems, it is necessary an efficient and automatic mapping of hundreds, or even thousands, of virtual nodes to physical hosts-and the mapping of the virtual links between guests to physical paths in the physical environment. In this paper we present a heuristic to map both virtual machines to hosts and virtual links between virtual machines to paths in the real system. We define the problem we are addressing, present the solution for it and evaluate it in different usage scenarios.
Rodrigo N. Calheiros, Rajkumar Buyya, César A. F. De Rose
ICPP3
2009 Towards self-managed adaptive emulation of grid environments
abstract
Distributed systems emulators built with the aid of virtualization tools allow testing of systems in a testbed whose number of real elements are orders of magnitude smaller than the number of virtual elements being tested. However, to allow testers to benefit from these systems, operation of the virtual environment should be hidden from them and performed automatically by the emulator. Moreover, testers may be unsure on the exact needs of their environment, and thus can request an environment that does not fit the experiment. In this paper we present our achievements in providing an emulation framework able to provide environment reconfiguration if the requested one does not comply with experiment's demands. Also, it supplies services such as execution log, environment monitoring, and automatic management of applications running in the virtual environment.
Rodrigo N. Calheiros, Everton Alexandre, Andriele B. do Carmo, César A. F. De Rose, Rajkumar Buyya
ISCC4
2009 Design of a Grid workflow for a climate application
abstract
Grid applications can be modeled as a composition of rather independent tasks. There are two approaches to define such a workflow either by combining multiple applications to build a more complex functionality or by splitting up an existing application. In this paper we analyze the latter process. We present a compute intensive application for climatology simulation and the options available to split it up. Using the simulation mode of our grid broker, we were able to compare the different workflow specifications before actually executing the workflows. This case study showed, using finer grained workflows-which usually need more adjustments to the software-allows better performance in the grid.
Jörg Schneider 0001, Julius Gehr, Hans-Ulrich Heiß, Tiago Ferreto, César A. F. De Rose, Rodrigo da Rosa Righi, Eduardo Rocha Rodrigues, Nicolas Maillard, Philippe Olivier Alexandre Navaux
ISCC5
2008 Applying Virtualization and System Management in a Cluster to Implement an Automated Emulation Testbed for Grid Applications
abstract
Although grid systems have evolved in such a way that they are largely used both in industry and academy, techniques to test and evaluate them, such as simulation and emulation, have limitations on both their applicability and their reliability. We are investigating the utilization of paravirtualization techniques merged with systems management tools to build an automated emulation framework for grid experiments. This framework accesses standard network resources to manage communication among virtual nodes, allowing virtual machines to behave like a real grid environment. The development of this framework involves the mapping of virtual machines to physical hosts, automatic deployment and management of virtual machines, automatic configuration of virtual network and experiment control. In this paper, we address these issues and present results demonstrating the feasibility and advantages of our approach.
Rodrigo N. Calheiros, Mauro Storch, Everton Alexandre, César A. F. De Rose, Marcus Breda
SBAC-PAD4
2008 Allocation strategies for utilization of space-shared resources in Bag of Tasks grids
César A. F. De Rose, Tiago Ferreto, Rodrigo N. Calheiros, Walfredo Cirne, Lauro Beltrão Costa, Daniel Fireman
Future Gener. Comput. Syst.1
2007 Distributed dynamic processor allocation for multicomputers
César A. F. De Rose, Hans-Ulrich Heiß, Barry Linnert
Parallel Comput.1
2006 GerpavGrid: using the Grid to maintain the city road system
abstract
This paper presents and evaluates a governmental application that has been ported to run in a Grid. The Gerpav application is used in the city of Porto Alegre, located in the south of Brazil, to maintain and plan the investments in the city road system, and Grid technology brought substantial performance gains to its distributed version called GerpavGrid. We describe several optimizations strategies used in GerpavGrid to move the bottleneck of the sequential application from database to memory to facilitate distribution of tasks in the Grid. Results of its execution in a real Grid are also presented, and show that a Grid can be also an interesting execution platform for non-classical HPC applications that are database intensive.
César A. F. De Rose, Tiago Ferreto, Marcelo B. de Farias, Vladimir G. Dias, Walfredo Cirne, Milena P. M. Oliveira, Katia Barbosa Saikoski, Maria Luiza Danieleski
SBAC-PAD1
2005 Transparent Resource Allocation to Exploit Idle Cluster Nodes in Computational Grids
abstract
Clusters of workstations are one of the most suitable resources to assist e-scientists in the execution of large-scale experiments that demand processing power. The utilization rate of these machines is usually far from 100%, and hence this should motivate administrators to share their clusters to grid communities. However, exploiting these resources in computational grids is challenging and brings several problems. This paper presents a transparent resource allocation strategy to harness idle cluster resources aimed at executing grid applications. This novel approach does not make use of a formal allocation request to cluster resource managers. Moreover, it does not interfere with local cluster users, being non-intrusive, and hence motivating cluster administrators to publish their resources to grid communities. We present experimental results regarding the effects of the proposed strategy on the attendance time of both cluster and grid requests and we also analyze its effectiveness in clusters with different utilization rates.
Marco Aurélio Stelmar Netto, Rodrigo N. Calheiros, Rafael K. S. Silva, César A. F. De Rose, Caio Northfleet, Walfredo Cirne
e-Science4
2005 Scheduling Divisible Workloads Using the Adaptive Time Factoring Algorithm
Tiago Ferreto, César A. F. De Rose
ICA3PP2
2005 VRM: A Failure-Aware Grid Resource Management System
abstract
For resource management in grid environments, advance reservations turned out to be very useful and hence are supported by a variety of grid toolkits. However, failure recovery for such systems has not yet received the attention it deserves. In this paper, we address the problem of remapping reservations to other resources, when the originally selected resource fails. Instead of dealing with jobs already running, which usually means checkpointing and migration, our focus is on jobs that are scheduled on the failed resource for a specific future period of time but not started yet. The most critical factor when solving this problem is the estimation of the downtime. We avoid the drawbacks of under- or overestimating the downtime by a dynamic load-based approach that is evaluated by extensive simulations in a grid environment and shows superior performance compared to estimation-based approaches.
Lars-Olof Burchard, César A. F. De Rose, Hans-Ulrich Heiß, Barry Linnert, Jörg Schneider 0001
SBAC-PAD2
2005 The I-Cluster Cloud: distributed management of idle resources for intense computing
Bruno Richard 0002, Nicolas Maillard, César A. F. De Rose, Reynaldo Novaes
Parallel Comput.3
2004 Scheduling BoT Applications in Grids Using a Slave Oriented Adaptive Algorithm
Tiago Ferreto, César A. F. De Rose, Caio Northfleet
ISPA2
2004 Scheduling in Bag-of-Task Grids: The PAUÁ Case
abstract
In this paper we discuss the difficulties involved in the scheduling of applications on computational grids. We highlight two main sources of difficulties: 1) the size of the grid rules out the possibility of using a centralized scheduler; 2) since resources are managed by different parties, the scheduler must consider several different policies. Thus, we argue that scheduling applications on a grid require the orchestration of several schedulers, with possibly conflicting goals. We discuss how we have addressed this issue in the context of PAUA, a grid for Bag-of-Tasks applications (i.e. parallel applications whose tasks are independent) that we are currently deploying throughout Brazil.
Walfredo Cirne, Francisco Vilar Brasileiro, Lauro Beltrão Costa, Daniel Paranhos da Silva, Elizeu Santos-Neto, Nazareno Andrade, César A. F. De Rose, Tiago Ferreto, Miranda Mowbray, Roque Scheer, João Jornada
SBAC-PAD7
2003 A Hardware Counters Based Tool for System Monitoring
Tiago Ferreto, Luiz De Rose, César A. F. De Rose
Euro-Par3
2003 Improving Performance Analysis Using Resource Management Information
Tiago Ferreto, César A. F. De Rose
HiPC2
2003 CRONO: A Configurable and Easy to Maintain Resource Manager Optimized for Small and Mid-Size GNU/Linux Cluster
abstract
We present the design and implementation of a new management system called CRONO aimed at small and mid-size GNU/Linux cluster installations owned by nonspecialized users. CRONO implements only the basic management services needed to share a cluster among several users and is optimized for machines with up to 64 nodes, being therefore easy to install, maintain and use, while still being highly configurable. We also show how to configure CRONO for an environment with one and other with three clusters, as well as some maintenance procedures to give an idea of the simplicity of both tasks
Marco Aurélio Stelmar Netto, César A. F. De Rose
ICPP2
2003 Performance Issues of Bandwidth Reservations for Grid Computing
abstract
In general, two types of resource reservations in computer networks can be distinguished: immediate reservations which are made in a just-in-time manner and advance reservations which allow to reserve resources a long time before they are actually used. Advance reservations are especially useful for grid computing but also for a variety of other applications that require network quality-of-service, such as content distribution networks or even mobile clients, which need advance reservation to support handovers for streaming video. With the emerged MPLS standard, explicit routing can be implemented also in IP networks, thus overcoming the unpredictable routing behavior which so far prevented the implementation of advance reservation services. The impact of such advance reservation mechanisms on the performance of the network with respect to the amount of admitted requests and the allocated bandwidth has so far not been examined in detail. We show that advance reservations can lead to a reduced performance of the network with respect to both metrics. The analysis of the reasons shows a fragmentation of the network resources. In advance reservation environments, additional new services can be defined such as malleable reservations and can lead to an increased performance of the network. Four strategies for scheduling malleable reservations are presented and compared. The results of the comparisons show that some strategies increase the resource fragmentation and are therefore unsuitable in the considered environment while others lead to a significantly better performance of the network. Besides discussing the performance issue, the software architecture of a management system for advance reservations is presented.
Lars-Olof Burchard, Hans-Ulrich Heiß, César A. F. De Rose
SBAC-PAD3
2002 RVision: An Open and High Configurable Tool for Cluster Monitoring
abstract
In this paper we present the design and implementation of RVision (Remote Vision), an open architecture, high configurable tool for cluster monitoring. We focus on the description of it's modular architecture, emphasizing the new concepts we are introducing for cluster monitoring, such as monitoring sessions and the support for dynamic linking of monitoring libraries. In addition, we measure intrusion with several benchmarks and applications under different scenarios. RVision distinguishes itself from other available tools for cluster monitoring because of its open architecture, high configurability, and low intrusion. It is being used in production mode, in our research center, and has proven itself as a powerful alternative for cluster monitoring, especially in heterogeneous clusters and cluster of clusters.
Tiago Ferreto, César A. F. De Rose, Luiz De Rose
CCGRID2
2002 The Virtual Cluster: A Dynamic Environment for Exploitation of Idle Network Resources
abstract
Standard environments for exploiting idle time of workstations are based on some kind of spying process that detects low CPU usage and informs to a scheduler so that work can be dispatched. This approach generates local interference and, since the same local environment is used, could lead to security problems. We are investigating the exploitation of idle times in network resources based on a complete mode change in a candidate node. After the detection that some node is idle, a mode-switcher boots a new operating system that will work over a separate disk partition. After the boot phase the node is linked to a logical network topology and is available to receive jobs. Users can allocate nodes from this virtual cluster through a standard frontend as they would do in a "conventional" cluster. Because nodes may leave and join this virtual machine we use a distributed processor management to allow user applications to cope with this dynamic resource behavior. In this paper we describe the architecture of the virtual cluster and present the results obtained with a mode-switcher and a prototype application under real use conditions.
César A. F. De Rose, Franco Blanco, Nicolas Maillard, Katia Barbosa Saikoski, Reynaldo Novaes, Olivier Richard, Bruno Richard 0002
SBAC-PAD1
2001 DECK-SCI: High-Performance Communication and Multithreading for SCI Clusters
abstract
This paper presents the design and implementation of DECK-SCI, a multithreaded communication library that fully exploits the high-performance capabilities of the SCI technology. We compare DECK-SCI, in terms of performance, to a commercially distributed MPI implementation and to a freely available MPICH distribution, both specifically designed for SCI clusters.
Fabio A. D. de Oliveira, Rafael Bohrer Ávila, Marcos E. Barreto, Philippe Olivier Alexandre Navaux, César A. F. De Rose
CLUSTER5
2001 Dynamic Processor Allocation in Large Mesh-Connected Multicomputers
César A. F. De Rose, Hans-Ulrich Heiß
Euro-Par1
2000 Distributed Processor Allocation in Large PC Clusters
abstract
Current processor allocation techniques for highly parallel systems are based on centralized front-end based algorithms. As a result, the applied strategies are restricted to static allocation, low parallelism and weak fault tolerance. To lift these restrictions, we are investigating a distributed approach to the processor allocation problem in large distributed memory machines. A contiguous and a noncontiguous version of a distributed dynamic processor allocation strategy are proposed and studied. Simulations compare the performance of the proposed strategies with that of well-known centralized algorithms. We also present the results of experiments on a Simens hpcline Primergy Server with 96 nodes that show distributed allocation is feasible with current technologies.
Hans-Ulrich Heiß, César A. F. De Rose, Philippe Olivier Alexandre Navaux
HPDC2
1995 Performance evaluation in image processing with GAPP array processor
Philippe Olivier Alexandre Navaux, César A. F. De Rose, Gerson G. H. Cavalheiro
Microprocess. Microprogramming2