Marcelo Amaral

dblp:95/10031 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2024
0000-0003-3212-2312ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 Process-Based Efficient Power Level Exporter
abstract
In this paper, we present the Kepler framework, designed to address the critical need for precise power and energy measurement in on-prem cloud-native, containerized environments, with a specific focus on processes, containers, and Kubernetes pods. The framework aims to support other tools in making informed decisions regarding provisioning, scheduling, and energy-optimization in cloud environments. Our approach involves leveraging the Kepler framework to create power models using Hardware Counters (HC), and real- time system power metrics from hardware sensors like x86 Running Average Power Limit (RAPL). Unlike previous methods that create and validate power models using aggregated system metrics, we propose a versatile process-level power model trained with per-process metrics. Those metrics are collected via a series of experiments in a controlled environment, measuring the incremental power consumption of processes under different scenarios. The collected data is then utilized to create a power model to be used in a shared cloud environment, and to validate the created power models using different set of input metrics. Our results show a significant improvement in the model accuracy compared to prior works, when incorporating per- process metrics and real-time system power metrics into the power estimation process. For instance, using the simplest power model, which is based on CPU utilization ratio, resulted in a Sum of Squared Error (SSE) of 75. In contrast, a power model created using aggregated system metrics, as the related works, had an SSE of 175 without real-time power metrics, and 5.6 with our proposed model refinement by normalizing the model results with the real-time system power metrics. On the other hand, training the power model with per-process metrics from controlled experiments yielded an SSE as low as 1.68 using real- time system power metrics, representing a 70% improvement in model accuracy compared to using aggregated system metrics, and an SSE 8.7 without power metrics, representing a 95% improvement in model accuracy. Furthermore, the results show that Kepler has a notable lower overhead by utilizing extended Berkeley Packet Filter (eBPF) for HC collection than alternative methods.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD1
2024 Best-Effort Power Model Serving for Energy Quantification of Cloud Instances
abstract
Quantifying energy consumption is a fundamental element of green computing. Power models trained by resource utilization allow quantifying the energy number and enable energy-efficient resource management systems without raising the concerns of complexity, cost, and security. However, energy consuming behavior on different machines varies by several factors. In this paper, we address the challenges of power modeling for cloud instances where information about these factors is obscured or unseen in the training set, and propose a best-effort method to train and serve a power model as precise as possible by leveraging a large, industry-standard power database. The proposed method prioritizes the modeling precision, and offers similarity and uncertainty indicators to elucidate the confidence level when serving an unseen instance. The results have demonstrated feasibility and precision of the proposed method against comparable approaches.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Marcelo Amaral, Rina Nakazawa, Scott Trent, UmaMaheswari Devi, Tamar Eilam
MASCOTS3
2023 Kepler: A Framework to Calculate the Energy Consumption of Containerized Applications
abstract
Energy accounting is crucial in data centers for optimizing power provisioning, capping, and tuning. This paper introduces the Kepler framework, which estimates power consumption at the process, container, and Kubernetes pod levels. Kepler offers a set of power models applicable to various architectures and metrics. In this study, we propose a generic power model that utilizes hardware counters (HC) and realtime system power metrics (e.g., running average power limit (RAPL)) as independent variables in a regression model. Unlike previous approaches that rely on aggregate power consumption, our methodology measures individual process power consumption to train the power model. We provide step-by-step instructions to measure process power consumption in a controlled environment, considering the activation constant and load-dependent dynamic power consumption in different executions. By following the Greenhouse Gas (GHG) Protocol, our approach ensures the fair distribution of constant power among the user's processes. The results demonstrate significantly improved accuracy with a mean squared error (MSE) as low as 0.010 for the proposed method, compared with an MSE of 0.16 for a simple ratio approach and 0.92 when training the model using aggregated workload power.
Marcelo Amaral, Tatsuhiro Chiba, Rina Nakazawa, Sunyanan Choochotkaew, Tamar Eilam
CLOUD1
2023 Advancing Cloud Sustainability: A Versatile Framework for Container Power Model Training
abstract
Estimating power consumption in modern Cloud is important to account for the power consumed by each container. The challenge is that multiple customers are sharing the same hardware platform, where physical information is mostly obscured. In addition, there is the overhead in power consumption that the Cloud control plane induces. This paper addresses these challenges and introduces a pipeline framework for container power model training on the basis of available performance counters and other metrics. The proposed model utilizes machine learning techniques to predict the power consumed by the control plane and associated processes when running together with the user containers, and uses it for isolating the dynamic power consumed by the user-inducing workload. Applying the proposed power model does not require online power measurements, nor does it need machine information, or information on other tenants sharing the same machine. The results of cross-workload, cross-platform experiments demonstrated the higher accuracy of the model when predicting power consumption of unseen containers on unknown platforms, including on virtual machines.
Sunyanan Choochotkaew, Chen Wang 0039, Tatsuhiro Chiba, Marcelo Amaral, Tamar Eilam
MASCOTS5
2022 MicroLens: A Performance Analysis Framework for Microservices Using Hidden Metrics With BPF
abstract
Determining the root cause of performance regression for microservices is challenging. The topological cascading performance implications among microservices hide the source of the problem. Additionally, the lack of knowledge about application phases can potentially lead to false-positive critical service detection. Service resource utilization is an imperfect proxy for application performance, potentially leading to false positives. Therefore, in this work, we propose a new performance testing framework that leverages hidden Berkeley Packet Filter (BPF) kernel metrics to locate root causes of performance regression. The framework applies a systematic multi-level approach to analyze microservice performance without intrusive code instrumentation. First, the framework constructs an attributed graph with microservice requests, scores the services to identify the critical paths, and ranks the low-level metrics to highlight the root cause of performance regression. Through judiciously designed experiments, we evaluated the metric collection overhead, showing less than 18% more latency when the application is running across hosts and 9% within the same host. In addition, depending on the application, no overhead is experienced, while the state-of-the-art approach presented up to 1060% more latency. The microservice benchmark evaluation shows that MicroLens can successfully identify the set of root causes and that the causes vary when the application is running in different infrastructures.
Marcelo Amaral, Tatsuhiro Chiba, Scott Trent, Takeshi Yoshimura, Sunyanan Choochotkaew
CLOUD1
2022 Bypass Container Overlay Networks with Transparent BPF-driven Socket Replacement
abstract
Containerization on the cloud offers several crucial benefits. However, these benefits are negated by the effects of virtual network stack and address encapsulation, especially for workloads that require intense communication. Socket replacement is a promising approach to breach this wall without changing the underlay infrastructure by replacing a nested network stack with a simple host network stack. Current state-of-the-art approaches perform this replacement by preloading the overridden socket library in a containerized process. However, the preloading approach requires user effort to modify the deploying manifests and a compromised security policy configuration of privileged containers to access the host namespace. This paper introduces a new replacement framework where a secured control plane agent performs the replacement by utilizing low-overhead BPF kernel tracing technology. As a result, containers can obtain host-native network performance and neither modification nor escalated privileges are required for user containers. Experiments on multiple benchmarks including iPerf, MPI, memslap, and GROMACS have been conducted to confirm efficacy.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Marcelo Amaral
CLOUD4
2022 AutoDECK: Automated Declarative Performance Evaluation and Tuning Framework on Kubernetes
abstract
Containerization and application variety bring many challenges in automating evaluations for performance tuning and comparison among infrastructure choices. Due to the tightly-coupled design of benchmarks and evaluation tools, the present automated tools on Kubernetes are limited to trivial microbenchmarks and cannot be extended to complex cloudnative architectures such as microservices and serverless, which are usually managed by customized operators for setting up workload dependencies. In this paper, we propose AutoDECK, a performance evaluation framework with a fully declarative manner. The proposed framework automates configuring, deploying, evaluating, summarizing, and visualizing the benchmarking workload. It seamlessly integrates mature Kubernetes-native systems and extends multiple functionalities such as tracking the image-build pipeline, and auto-tuning. We present five use cases of evaluations and analysis through various kinds of bench-marks including microbenchmarks and HPC/AI benchmarks. The evaluation results can also differentiate characteristics such as resource usage behavior and parallelism effectiveness between different clusters. Furthermore, the results demonstrate the benefit of integrating an auto-tuning feature in the proposed framework, as shown by the 10% transferred memory bytes in the Sysbench benchmark.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Takeshi Yoshimura, Marcelo Amaral
CLOUD5
2022 Detecting Layered Bottlenecks in Microservices
abstract
We propose a method to detect both software and hardware bottlenecks in a web service consisting of microservices. A bottleneck is a resource that limits the maximum performance of the entire web service. Bottlenecks often include both software resources such as threads, locks, and channels, and hardware resources such as processors, memories, and disks. Bottlenecks form a layered structure since a single request can utilize multiple software resources and a hardware resource simultaneously. The microservice architecture makes the detection of layered bottlenecks challenging due to the lack of a uniform analysis perspective across languages, libraries, frameworks, and middle-ware.We detect layered bottlenecks in microservices by profiling numbers and status of working threads in each microservice and dependency among microservices via network connections. Our approach can be applied to various programming languages since it relies only on standard debugging tools. Nevertheless, our approach not only detects which microservice is a bottleneck but also enables us to understand why it becomes a bottleneck. This is enabled by a novel visualization method to show layered bottlenecks in microservices at a glance. We demonstrate that our approach successfully detects and visualizes layered bottlenecks in the state-of-the-art microservice benchmarks, DeathStarBench and Acme Air microservices. This enables us to optimize the microservices themselves to achieve a higher throughput per re-source utilization rate compared with simply scaling the number of replicas of microservices.
Tatsushi Inagaki, Yohei Ueda, Moriyoshi Ohara, Sunyanan Choochotkaew, Marcelo Amaral, Scott Trent, Tatsuhiro Chiba, Qi Zhang 0009
CLOUD5
2021 Run Wild: Resource Management System with Generalized Modeling for Microservices on Cloud
abstract
Microservice architecture competes with the traditional monolithic design by offering benefits of agility, flexibility, reusability resilience, and ease of use. Nevertheless, due to the increase in internal communication complexity, care must be taken for resource-usage scaling in harmony with placement scheduling, and request balancing to prevent cascading performance degradation across microservices. We prototype Run Wild, a resource management system that controls all mechanisms in the microservice-deployment process covering scaling, scheduling, and balancing to optimize for desirable performance on the dynamic cloud driven by an automatic, united, and consistent deployment plan. In this paper, we also highlight the significance of co-location aware metrics on predicting the resource usage and computing the deployment plan. We conducted experiments with an actual cluster on the IBM Cloud platform. RunWild reduced the 90th percentile response time by 11% and increased average throughput by 10% with more than 30% lower resource usage for widely used autoscaling benchmarks on Kubernetes clusters.
Sunyanan Choochotkaew, Tatsuhiro Chiba, Scott Trent, Marcelo Amaral
CLOUD4
2017 Topology-aware GPU scheduling for learning workloads in cloud environments
abstract
Recent advances in hardware, such as systems with multiple GPUs and their availability in the cloud, are enabling deep learning in various domains including health care, autonomous vehicles, and Internet of Things. Multi-GPU systems exhibit complex connectivity among GPUs and between GPUs and CPUs. Workload schedulers must consider hardware topology and workload communication requirements in order to allocate CPU and GPU resources for optimal execution time and improved utilization in shared cloud environments.
Marcelo Amaral, Jorda Polo, David Carrera 0001, Seetharami R. Seelam, Malgorzata Steinder
SC1
2015 Performance Evaluation of Microservices Architectures Using Containers
abstract
Micro services architecture has started a new trend for application development for a number of reasons: (1) to reduce complexity by using tiny services, (2) to scale, remove and deploy parts of the system easily, (3) to improve flexibility to use different frameworks and tools, (4) to increase the overall scalability, and (5) to improve the resilience of the system. Containers have empowered the usage of micro services architectures by being lightweight, providing fast start-up times, and having a low overhead. Containers can be used to develop applications based on monolithic architectures where the whole system runs inside a single container or inside a micro services architecture where one or few processes run inside the containers. Two models can be used to implement a micro services architecture using containers: master-slave, or nested-container. The goal of this work is to compare the performance of CPU and network running benchmarks in the two aforementioned models of micro services architecture hence provide a benchmark analysis guidance for system designers.
Marcelo Amaral, Jorda Polo, David Carrera 0001, Iqbal Mohomed, Merve Unuvar, Malgorzata Steinder
NCA1
2011 EbitSim: An Enhanced BitTorrent Simulation Using OMNeT++ 4
abstract
The BitTorrent protocol is one of the most successful P2P applications, being largely studied by the research community. Nevertheless, studying the dynamics of a large BitTorrent network presents several challenges, such as difficulty in acquiring network traces or building measurement experiments. Evaluation through simulation is usually utilized for studying BitTorrent networks, yet only a few BitTorrent simulation models are available for the research community. In this article, we present an extensible framework for developing BitTorrent simulations, focusing on realism and without losing on scalability. We developed an accurate version of the protocol by analyzing the source code of mainstream BitTorrent clients and by discussing directly with their developers. The simulation model was developed with the OMNeT++ Framework, inheriting its high extensibility, and with the INET Framework, for accuracy of the underlying network model. We also took into account the effects of multitasking in our model, since BitTorrent applications acquires content from several sources simultaneously, and utilized real world traces for modeling the processing times. We present an analysis of our simulator regarding performance aspects and BitTorrent-related results.
Pedro Evangelista, Marcelo Amaral, Charles Miers, Walter Akio Goya, Marcos A. Simplício Jr., Tereza Cristina M. B. Carvalho, Victor Souza
MASCOTS2