Márcio Castro 0001

dblp:39/6633 · also Márcio Bastos Castro · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-9992-8540ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Performance and Cost Evaluation of StarPU on AWS: Case Studies With Dense Linear Algebra Kernels and N-Body Simulations
abstract
ABSTRACT Task‐based programming interfaces introduce a paradigm in which computations are decomposed into fine‐grained units of work known as “tasks”. StarPU is a runtime system originally developed to support task‐based parallelism on on‐premise heterogeneous architectures by abstracting low‐level hardware details and efficiently managing resource scheduling. It enables developers to express applications as task graphs with explicit data dependencies, which are then dynamically scheduled across available processing units, such as CPUs and GPUs. In recent years, major cloud providers have begun offering virtual machines equipped with both CPUs and GPUs, allowing researchers to deploy and execute parallel workloads in virtual heterogeneous clusters. However, the performance and cost effectiveness of executing StarPU‐based applications in public cloud environments remain unclear, particularly due to variability in hardware configurations, network performance, ever‐changing pricing models, and computing performance due to virtualization and multi‐tenancy. In this paper, we evaluate the performance and cost‐efficiency of StarPU on Amazon Elastic Compute Cloud (EC2) using dense linear algebra kernels and N‐Body simulations as case studies. Our experiments consider different cluster configurations, including powerful and more expensive instances with four NVIDIA GPUs per node (which we refer to as “fat nodes”), and less powerful and lower‐cost instances with a single NVIDIA GPU per node (which we refer to as “thin nodes”). Our results show that arithmetic precision affects the performance–cost trade‐off for dense linear algebra applications, whereas N‐Body simulations consistently achieve better cost‐efficiency on thin‐node clusters. These findings underscore the challenges of optimizing HPC workloads for performance and cost in cloud environments.
Vanderlei Munhoz, Vinícius Garcia Pinto, João V. F. Lima, Márcio Castro 0001, Daniel Cordeiro, Emilio Francesquini
Concurr. Comput. Pract. Exp.4
2025 Multicore Environment State Representation for Agent-Directed Test Generation
abstract
A crucial step in the design of multicore systems is to validate the interaction between cores. This involves test program generation and runtime analysis. We propose a novel reinforcement learning approach to directed test generation, where an agent induces a suite of programs, which are executed in a simulation environment for a multicore. It focuses on how to recover state information from raw observations of the environment such that the agent can learn from interaction how to improve coverage for any verification task. We evaluated our state representation for different verification tasks involving 16 and 32-core ARMv8 2-level MOESI designs.
Bruno D. Miranda, Luiz M. V. Pereira, Márcio Castro 0001, Luiz Cláudio Villar dos Santos
DAC3
2025 Task-Based HPC in the Cloud: Price-Performance Analysis of N-Body Simulations with StarPU
abstract
Public cloud environments present significant challenges for traditional High Performance Computing (HPC) applications due to infrastructure limitations that differ substantially from dedicated HPC systems. Unlike traditional HPC clusters optimized for tightly coupled parallel workloads, cloud platforms were designed primarily for web services and data processing applications. Key obstacles include high-latency networks, hardware virtualization overhead, and limited availability of specialized accelerators, all of which can severely impact the performance of compute-intensive applications such as physics simulations. This study investigates the feasibility of running HPC workloads on public cloud infrastructure using standard and cost-effective instance configurations rather than expensive specialized "HPC" offerings. We deploy heterogeneous clusters on Amazon Web Services using the HPC@Cloud Toolkit, incorporating various instance types, including GPU-accelerated nodes with different computational capabilities. Our evaluation focuses on N-body simulations implemented using a task-based parallel programming model, leveraging the StarPU runtime system to dynamically schedule computational tasks across various processing units. Our experimental results demonstrate three key findings: (1) smaller GPU-equipped instances (g6.2xlarge) achieve performance comparable to larger instances while costing approximately one-sixth the price, challenging conventional scaling assumptions for cloud-based HPC; (2) strategic GPU utilization yields up to 8.2× performance improvements over CPU-only configurations while reducing total execution costs by 24.4×; and (3) while task-based programming models effectively address network limitations through dynamic scheduling, complex tree-based algorithms like TBFMM face significant optimization challenges in cloud environments due to load balancing issues and expensive parameter tuning requirements. These findings provide practical guidance for researchers and practitioners seeking cost-effective cloud HPC deployments, demonstrating that commodity cloud infrastructures can be viable for regular computational workloads but require careful algorithmic-resource matching for optimal efficiency.
Nicolas Vanz, Vanderlei Munhoz, Márcio Castro 0001, Laércio Lima Pilla, Olivier Aumage
IC2E3
2025 Dynamic Load Balancing in Kubernetes Environments With Kubernetes Scheduling Extension (KSE)
abstract
ABSTRACT Kubernetes is a flexible and reliable container orchestrator that has been employed to maintain massive cloud infrastructures worldwide. The task of allocating “Pods” (deployable units of computing that have one or more containers) to cluster nodes in Kubernetes is done by the Kube‐Scheduler module, which determines which nodes are valid placements for each Pod according to constraints and available resources. Since it only acts at Pod creation, it does not take any action when the load of the system becomes uneven. In imbalanced scenarios, overloaded nodes can compromise the performance and availability of services hosted by them, whereas underloaded nodes may be a waste of financial resources, especially when using public clouds. In this paper, we propose an extension to the Kubernetes scheduler, called Kubernetes Scheduling Extension (KSE), which allows users to implement dynamic load‐balancing algorithms that can migrate Pods between nodes at runtime. We also provide the implementation of two well‐known load‐balancing algorithms (KSE‐GreedyLB and KSE‐RefineLB) in KSE, which can balance the load of system nodes using CPU and memory consumption metrics. We carried out several experiments to assess the effectiveness of KSE‐GreedyLB and KSE‐RefineLB and compared their results with Kube‐Scheduler. Overall, we evaluated 32 different scenarios using synthetic and realistic applications. Our results showed that KSE‐RefineLB achieves better results than Kube‐Scheduler when the workload is highly imbalanced while keeping similar performance when the load imbalance is low.
Pedro Moritz de Carvalho Neto, Márcio Castro 0001, Frank Siqueira
Concurr. Comput. Pract. Exp.2
2025 A Canonical Test Representation for Verification of Shared-Memory Behavior in Multiprocessor Systems
abstract
The scope of this article is the design verification of a multicore chip or multichip multiprocessor by running concurrent test programs until coverage goals are reached. Interactions between multiple processors through shared memory must obey a memory consistency model, which specifies valid behaviors. We propose a canonical test-program representation that encodes primal shared-memory behaviors to be induced at runtime. It is intended as one of the main keys to the design of new test generators. We prove that our representation does not limit the search space, because it induces equivalence classes that can be completely and uniquely encoded. In particular, we show experimental evidence that our representation is also suitable to learning-based test generators, because it enables the design of effective actions. We have built a generator directed by a Reinforcement Learning agent, designed its actions based on our encoding, and compared it with three generators when targeting 32-core designs. For a given time limit, our generator reached the largest coverage and led to the fastest error diagnosis in 3/4 of the verification scenarios, despite our choice of a minimalist agent. The theoretical guarantees and the experimental evidence indicate that our representation provides proper grounds for defining effective actions, and it prevents them from either inducing redundant tests or limiting the test suite.
Bruno D. Miranda, Márcio Castro 0001, Luiz Cláudio Villar dos Santos
ACM Trans. Design Autom. Electr. Syst.2
2024 Enabling the execution of HPC applications on public clouds with HPC@Cloud toolkit
abstract
Abstract The advent of cloud computing has made access to computing infrastructure available to millions of users that face resource constraints. In the context of high performance computing (HPC), public cloud resources have emerged as a cost‐effective alternative to expensive on‐premises clusters. However, there are several challenges and limitations in adopting this approach. This paper proposes HPC@Cloud , a provider‐agnostic open‐source software toolkit that facilitates the migration, testing, and execution of HPC applications in public clouds. The toolkit takes advantage of various fault tolerance technologies to enable the use of inexpensive transient cloud infrastructure, commonly known as “spot” instances. Also, it features integration with singularity containers, allowing users to run complex applications on virtual HPC clusters in a portable and reproducible way. Finally, it provides a data‐based empirical approach to estimating cloud infrastructure costs for HPC workloads. The results obtained on two public cloud providers (AWS and Vultr) show that: (i) HPC@Cloud can efficiently build virtual HPC clusters on the cloud; (ii) the new adaptive fault tolerance strategy outperforms other existing strategies based on blocking restoration; (iii) the integration of singularity containers into HPC@Cloud improves the portability of HPC applications to public clouds with negligible performance penalty to the applications; (iv) the proposed cost prediction approach can estimate the cost of running the applications on AWS and Vultr with up to 93% accuracy on average.
Vanderlei Munhoz, Márcio Castro 0001
Concurr. Comput. Pract. Exp.2
2023 LWMPI: An MPI library for NoC-based lightweight manycore processors with on-chip memory constraints
abstract
Abstract Lightweight manycore processors deliver high performance and energy efficiency by bundling hundreds of low‐power cores, a distributed memory architecture with small local memories and Networks‐on‐Chip in a single die. However, the lack of rich and portable programming models for these processors makes software development a challenging task. Currently, two approaches are employed to address programmability in lightweight manycores: Operating Systems (OSes) and baremetal runtime libraries. The former provides portability but exposes complex Operating System (OS)‐level programming interfaces to developers. The latter focuses on providing rich and high performance interfaces, which are vendor‐specific and yield to non‐portable software. In this work, we address these programmability and portability challenges by combining a rich OS with a well‐known standard for parallel programming. We propose a portable and lightweight Message Passing Interface (MPI) library (LWMPI) designed from scratch to cope with restrictions and intricacies of lightweight manycores. We integrated LWMPI into Nanvix, an open‐source distributed OS that runs on silicon lightweight manycores. The results obtained with a synthetic benchmark and a subset of the CAP Bench applications running on Kalray MPPA‐256 unveil that LWMPI not only delivers a lightweight and richer programming interface but also presents good performance and scalability results.
João Fellipe Uller, João Vicente Souto, Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
Concurr. Comput. Pract. Exp.4
2023 Improving concurrency and memory usage in distributed operating systems for lightweight manycores via cooperative time-sharing lightweight tasks
João Vicente Souto, Márcio Castro 0001
J. Parallel Distributed Comput.2
2022 Strategies for Fault-Tolerant Tightly-Coupled HPC Workloads Running on Low-Budget Spot Cloud Infrastructures
abstract
Cloud providers can rent their spare computing capacity at substantial discounts, reclaiming it whenever there is a more profitable higher-priority request - a business model well known as spot infrastructure market. Users can attain significant cloud investment savings using spot machines, however with the caveat of increasing software complexity, given the fault tolerance requirements of this environment. Improvements in virtualization and network technology, combined with the development of key new software tools, may allow the HPC community to effectively take advantage of cheap cloud resources, cutting expensive maintenance costs. This study aims to evaluate the viability of budget-constrained cloud environments for tightly-coupled MPI applications, exploring both spot and traditional low-budget infrastructures from real public cloud platforms. We propose and evaluate two different fault tolerance strategies tailored for unreliable spot cloud environments: system-level rollback restart with Berkeley Labs Checkpoint/Restart (BLCR) and in-memory rollback restart with User-Level Failure Mitigation (ULFM). We also propose a provider-agnostic empirical method for testing and predicting MPI workloads execution times and cloud infrastructure costs. A detailed cost analysis and performance benchmark of a case-study application is provided, with data gathered from experiments with both spot and persistent machines from AWS and Vultr Cloud, respectively. Our results show that: (i) adequate cluster sizing plays an important role in the overall job execution performance and cost-effectiveness, regardless of the type of selected instances; (ii) fault tolerance strategies based on BLCR may have worse performance than ULFM, but still be costeffective considering software migration costs; (iii) the use of spot infrastructure does not guarantee costs savings depending on the chosen machine flavors and discounts, as experiments with persistent low-budget options attained better cost-effectiveness in some conditions.
Vanderlei Munhoz, Márcio Castro 0001, Odorico Machado Mendizabal
SBAC-PAD2
2021 A Task-based Execution Engine for Distributed Operating Systems Tailored to Lightweight Manycores with Limited On-Chip Memory
abstract
Operating Systems (OSes) for lightweight manycore processors feature a distributed design, where isolated OS instances cooperate to improve programmability and portability issues that come from their architectural intricacies. However, OS services often resort to processes or threads to implement kernel-level functionalities, consuming much of the already limited on-chip memory available in these processors. In this context, we propose a complementary OS-level execution engine that supports cooperative time-sharing lightweight tasks that share a unique execution stack and features task synchronization via task dependency graphs. This solution provides numerous OS-level execution flows with reduced memory consumption, leaving more room for user-level applications. We implemented our solution in an open-source distributed OS (Nanvix) and we compared it with the standard implementation that uses threads. The results obtained on Kalray MPPA-256 show that our engine: (i) provides 63.2x more execution flows per MB of memory; (ii) features 15.8x and 1.87x less overhead to spawn and destroy an execution flow, respectively, while consuming 6.65x less energy; (iii) runs remote system calls 3x faster; and (iv) reduces the memory footprint in 0.58x when executing OS services requests without impacting the overall system performance.
João Vicente Souto, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna
SBAC-PAD2
2021 PackStealLB: A scalable distributed load balancer based on work stealing and workload discretization
Vinicius Freitas, Laércio Lima Pilla, Alexandre de Limas Santana, Márcio Castro 0001, Johanne Cohen
J. Parallel Distributed Comput.4
2021 Inter-kernel communication facility of a distributed operating system for NoC-based lightweight manycores
Pedro Henrique de Mello Morado Penna, João Vicente Souto, João Fellipe Uller, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
J. Parallel Distributed Comput.4
2021 Dynamic power management under the RUN scheduling algorithm: a slack filling approach
Lais Borin, George Lima 0001, Márcio Castro 0001, Patricia Della Méa Plentz
Real Time Syst.3
2021 ARTful: A model for user-defined schedulers targeting multiple high-performance computing runtime systems
abstract
Abstract Global schedulers are components in parallel runtime libraries that distribute the application's workload across physical resources. More often than not, applications showcase dynamic load imbalance and require customized scheduling solutions to avoid wasting resources. Some libraries lack support for user‐defined schedulers and developers resort to unofficial extensions that are harder to reuse and maintain. We propose a global scheduler software design, entitled ARTful model, to create user‐defined solutions with minimal alterations in the runtime library. Our model uses a component‐based design to separate components from the runtime library and the scheduling policy implementation. The ARTful modeldescribes the interface of a portable scheduler library, allowing policies to operate on different runtime libraries. We study the overhead induced by our design through our ARTful library implementation metaprogramming‐oriented global scheduling library using workload‐aware scheduling policies. We experiment with two different policies from OpenMP and Charm++ runtime systems, also presenting evaluations of the policies outside of their original library context. We observe that our portable schedulers can sometimes perform decisions faster than their native counterparts with negligible overhead in the execution times of synthetic applications and molecular dynamics kernels.
Alexandre de Limas Santana, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla, Jean-François Méhaut
Softw. Pract. Exp.3
2020 Adaptive Load Balancing based on Machine Learning for Iterative Parallel Applications
abstract
The performance of irregular scientific applications can be easily affected by an uneven distribution of work among the computing resources. In this context, Load Balancing (LB) stands as one of the most important solutions to improve resource utilization. However, choosing the best-performing load balancing algorithm for a given application is not a trivial task. For instance, manually and statically choosing an LB algorithm does not work in situations where applications have a dynamic or unknown behavior. In this context, we propose a Machine Learning-based Adaptive Load Balancer (ADAPTIVELB) to automate the load balancing algorithm decision at run time. This approach monitors and collects information about the application dynamically, and according to the analyzed data, it makes a decision of invoking the most suitable LB algorithm. Our experiments show that ADAPTIVELB can select a good load balancing algorithm in most of the cases, leading to performance improvements over statically chosen LB algorithms and over the absence of a load balancer.
C. R. Anna Victoria Oikawa, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla
PDP3
2019 A comprehensive performance evaluation of the BinLPT workload-aware loop scheduler
abstract
Summary Workload‐aware loop schedulers were introduced to deliver better performance than classical loop scheduling strategies. However, they presented limitations such as inflexible built‐in workload estimators and suboptimal chunk scheduling. Targeting these challenges, we proposed previously a workload‐aware scheduling strategy called BinLPT, which relies on three features: (i) user‐supplied estimations of the workload of the loop; (ii) a greedy heuristic that adaptively partitions the iteration space in several chunks; and (iii) a scheduling scheme based on the Longest Processing Time (LPT) rule and on‐demand technique. In this paper, we present two new contributions to the state‐of‐the‐art. First, we introduce a multiloop support feature to BinLPT, which enables the reuse of estimations across loops. Based on this feature, we integrated BinLPT into a real‐world elastodynamics application, and we evaluated it running on a supercomputer. Second, we present an evaluation of BinLPT using simulations as well as synthetic and application kernels. We carried out this analysis on a large‐scale NUMA machine under a variety of workloads. Our results revealed that BinLPT is better at balancing the workloads of the loop iterations and this behavior improves as the algorithmic complexity of the loop increases. Overall, BinLPT delivers up to 37.15% and 9.11% better performance than well‐known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively.
Pedro Henrique de Mello Morado Penna, Antônio Tadeu A. Gomes, Márcio Castro 0001, Patricia Della Méa Plentz, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.3
2019 Foreword to the special issue of the workshop on high performance computing systems (XVIII Simpósio em Sistemas Computacionais de Alto Desempenho, WSCAD 2017)
abstract
This special issue of Concurrency and Computation Practice and Experience gathers extended versions of six selected research articles that were previously presented at the Brazilian Workshop on High Performance Computing Systems (“XVIII Simpósio em Sistemas Computacionais de Alto Desempenho”, WSCAD 2017), held in conjunction with the 29th International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2017, in Campinas, SP, Brazil, from the 17th to the 20th of October 2017. Since 2000, this workshop has presented important and interesting research in the fields of Computer Architecture, High Performance Computing and Distributed Systems. The scope of the current special issue is broad and representative of the multidisciplinary nature of High Performance Computing and Computer Architecture research domains. The set of accepted research articles was organized under three key themes: Parallel Algorithms and Optimizations, Scheduling and Placement, and Parallel Architecture Design. In the following sections, we provide a brief description of each one of the research articles accepted in this special issue. To achieve the best performance possible, algorithms must be carefully parallelized and optimized for multi-core or many-core processors. Current multi-core and many-core architectures may feature different technologies, such as distributed memory banks, vector instructions, and specialized cores. Oftentimes, different classes of multi-core and many-core processors are combined to construct a heterogeneous multiprocessing system. Today's technologies enable heterogeneous multiprocessing systems on a chip containing multi-cores, many-cores (eg, GPUs), and FPGAs. The following research articles present parallel solutions and optimizations to different classes of algorithms on multi-core and many-core processors. The paper “Optimized implementation of QC-MDPC code-based cryptography” presents a new enhanced version of the QcBits key encapsulation mechanism (KEM), which is a constant time implementation of the Niederreiter cryptosystem using QC-MDPC codes.1 The parallel solution uses vector instructions (AVX 512) and applies several other techniques to achieve a competitive performance level. The enhanced version is 1.9x faster when decrypting messages when compared with BIKE, which was the state-of-the-art implementation for QC-MDPC codes. The paper “On the Parallelization of Hirschberg's Algorithm for Multi-core and Many-core Systems” focuses on improving the execution efficiency of Hirschberg's algorithm, which aims at finding the longest common subsequence between two strings, on multi-core and many-core systems.2 The proposed solution exploits vector instructions and different parallelization strategies to achieve the best performance possible. Results showed that the parallel solution can achieve speedups of up to 15.5x on a 18-core Xeon processor and of up to 105x on a 68-core Intel Xeon Phi many-core processor. Finally, the paper “A Hybrid CPU-GPU-MIC Algorithm for Minimal Hitting Set Enumeration” proposes a hybrid exact algorithm for the Minimal Hitting Set (MHS) Enumeration Problem for highly heterogeneous platforms.3 The experiments were carried out on heterogeneous platforms composed of Intel Xeon E5-2620v2 CPUs, Intel Xeon Phi 3120A, and a GTX TITAN X GPUs. The results showed that the proposed algorithm was able to distribute parallel tasks among the processing units according to their computational efficiency in processing the task batches, achieving speedups of up to 25.3x in comparison with using two Intel Xeon E5-2620v2 CPUs. To deliver high performance to large-scale engineering and scientific applications, particular intricacies of the application and the underlying platform should be considered, so that tailored techniques can be employed to map one into another. In this context, evenly distributing the workload of an application among its threads, processes, or virtual machines is an NP-Hard minimization problem known as scheduling, and allocating these work abstractions to the underlying infrastructure is called placement. These problems are significant to the academic community and industry, and they are a hot research topic in High Performance Computing (HPC). The following research articles present contributions to these problems for different abstraction levels and domains. The paper “A Comprehensive Performance Evaluation of the BinLPT Workload-Aware Loop Scheduler” focuses on improving a workload-aware scheduling strategy called BinLPT, previously proposed by the same authors.4 Two new contributions are presented to the state of the art. First, a multiloop support feature was introduced to BinLPT, which enables the reuse of workload estimations across loops. Based on this feature, BinLPT was integrated into a real-world elastodynamics application and evaluated running on a supercomputer. Second, BinLPT was evaluated using simulations as well as synthetic and application kernels. This analysis was carried out on a large-scale NUMA machine under a variety of workloads. The results revealed that BinLPT is able to balance the load of irregular OpenMP parallel loops among the application threads, delivering up to 37% and 9% better performance than well-known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively. Finally, the paper “Optimizing the Performance of Multi-tier Applications Using Interference and Affinity-aware Placement Algorithms” proposes a combined approach that considers both resource interference and network affinity to decide the best placement of multi-tier applications in consolidated environments.5 In their previous work, the same authors identified that a combined approach could result in better solutions for this problem and proposed a set of placement policies that explore this tradeoff. The authors propose a new family of placement algorithms based on these policies and evaluated for different workload scenarios using a visual simulation tool called CIAPA. CIAPA introduces a performance degradation model, a cost function, and heuristics to find a placement with the minimum cost for a specific workload of multi-tier applications. The solution generated by CIAPA was compared to other placement strategies from related work, and delivered placement decisions with better cost, and, consequently, improved performance. An average reduction in response time of 10% was observed when compared to interference strategies, and up to 18% when considering only affinity strategies. Dataflow-based FPGA accelerators have become a promising alternative to deliver energy efficient platforms for the HPC domain. However, FPGA programming is still a challenge. Although reconfigurable FPGA technologies have been around since the 1980s, their utilization as a general-purpose processing platform is recent. Historically, both FPGA and ASIC developers have employed Hardware Description Languages (HDLs) to implement their designs, which is usually outside the main expertise area of software developers. The lack of simple and common programming models prevents software developers from easily designing accelerators and delays a broader adoption of this technology. In this context, the paper “ADD: Accelerator Design and Deploy - A Tool for FPGA High Performance Dataflow Computing” presents a high-level framework to specify, to simulate, and to implement dataflow accelerators for streaming applications.6 The Accelerator Design and Deploy (ADD) framework includes an open dataflow operator library, and templates are provided to easily design new operators. The framework also provides a high-level and an accurate simulation at circuit level with short execution times. Moreover, ADD provides software and hardware APIs to simplify the integration process, extending the benefits of portability from low-cost FPGA boards to high performance datacenter FPGA platforms. The framework supports coupling with high-level programming languages, and it has been validated on two FPGA platforms: the Intel high-performance CPU-FPGA heterogeneous computing platform and an educational FPGA kit. The authors show that the proposed approach presents competitive performance, both in time and energy, when compared to multi-core and GPU accelerators. Concerning energy, it is 18.8x and 193.2x more efficient than the GPU and multi-core evaluated platforms, respectively. The research articles presented in this special issue provide insights in fields related to High Performance Computing, including Parallel Algorithms and Optimizations, Scheduling and Placement, and Parallel Architecture Design. We believe that the main contributions presented in the research articles are timely and important. We hope that readers can benefit from insights of these research articles and contribute to these rapidly growing areas. Dr. César A. F. De Rose has a B.Sc. degree in Computer Science from the Pontifical Catholic University of Rio Grande do Sul (PUCRS, Porto Alegre, Brazil, 1990), an M.Sc. in Computer Science from the Federal University of Rio Grande do Sul (PGCC/UFRGS, Porto Alegre, Brazil, 1993), and a Doctoral degree from Karlsruhe Institute of technology (KIT - Karlsruhe, Germany, 1998). In 1998, he joined the Faculty of Informatics at PUCRS as an associate professor and member of the Resource Management and Virtualization Group (full professor since 2012). His research interests include resource management, dynamic provisioning and allocation, monitoring techniques (resource and application), application modeling, scheduling and optimization in parallel and distributed environments (Cluster, Grid, Cloud), and virtualization. In 2009, he founded PUCRS High Performance Computing Laboratory (LAD-PUCRS) being nowadays senior researcher. Dr. Márcio Castro received a B.Sc. in Computer Science with honors (Summa Cum Laude) from Pontifical Catholic University of Rio Grande do Sul (PUCRS, Brazil) in 2006 and an M.Sc. degree in Computer Science from the same university in 2009. He received a Ph.D. in Computer Science in 2012 from the University of Grenoble Alpes, France. Then, he worked as a postdoctoral fellow at the Federal University of Rio Grande do Sul (UFRGS), Brazil. Since 2014, he is an associate professor at the Federal University of Santa Catarina (UFSC), Brazil. His main research area is High Performance Computing, with focus on parallel programming models, load balancing, high performance parallel applications, and parallel and distributed computing on multi-core and many-core architectures. We would like to thank all the authors who provided valuable contributions to this special issue. We are also grateful to the reviewers for their feedback to the authors. Indeed, their advices were essential to further improve the quality of the papers. Finally, we would like to express our sincere gratitude to Professor Geoffrey Fox, the Editor in Chief, for providing us with this unique opportunity to present the selected papers from WSCAD 2017 in the International Journal of Concurrency and Computation: Practice and Experience.
César A. F. De Rose, Márcio Castro 0001
Concurr. Comput. Pract. Exp.2
2018 Energy Efficient Stencil Computations on the Low-Power Manycore MPPA-256 Processor
Emmanuel Podestá Jr., Bruno Marques do Nascimento, Márcio Castro 0001
Euro-Par3
2018 A Batch Task Migration Approach for Decentralized Global Rescheduling
abstract
Effectively mapping tasks of High Performance Computing (HPC) applications on parallel systems is crucial to assure substantial performance gains. As platforms and applications grow, load imbalance becomes a priority issue. Even though centralized rescheduling has been a viable solution to mitigate this problem, its efficiency is not able to keep up with the increasing size of shared memory platforms. To efficiently solve load imbalance today, and in the years to come, we should prioritize decentralized strategies developed for large scale platforms. In this paper, we propose our Batch Task Migration approach to improve decentralized global rescheduling, ultimately reducing communication costs and preserving task locality. We implemented and evaluated our approach in two different parallel platforms, using both synthetic workloads and a molecular dynamics (MD) benchmark. Our solution was able to achieve speedups of up to 3.75 and 1.15 on rescheduling time, when compared to other centralized and distributed approaches, respectively. Moreover, it improved the execution time of MD by factors up to 1.34 and 1.22 when compared to a scenario without load balancing on two different platforms.
Vinicius Freitas, Alexandre de Limas Santana, Márcio Castro 0001, Laércio Lima Pilla
SBAC-PAD3
2017 Provisioning and Delivering Sepsis Data Supported by an Enhanced SDN Environment
abstract
Medical applications, along with Information and Communication Technology (ICT), have contributed with many solutions to support the treatment of SEPSIS. However, there are few solutions for the transport of sepsis data with Quality of Service (QoS). In this paper we propose a self-manageable architecture for the provision and delivery of sepsis data using Software-Defined Networking (SDN). To evaluate our proposal, we conducted our experiments in the laboratory. The results are promising because our solution with SDN is able to improve performance, ensure QoS and prioritize sepsis flows, when compared to TCP/IP architectures.
Felipe Volpato, Madalena Pereira da Silva, Alexandre L. Gonçalves, Márcio Castro 0001, Mario A. R. Dantas
CBMS4
2017 Design methodology for workload-aware loop scheduling strategies based on genetic algorithm and simulation
abstract
Summary In high‐performance computing, the application's workload must be evenly balanced among threads to deliver cutting‐edge performance and scalability. In OpenMP, the load balancing problem arises when scheduling loop iterations to threads. In this context, several scheduling strategies have been proposed, but they do not take into account the input workload of the application and thus turn out to be suboptimal. In this work, we introduce a design methodology to propose, study, and assess the performance of workload‐aware loop scheduling strategies. In this methodology, a genetic algorithm is employed to explore the state space solution of the problem itself and to guide the design of new loop scheduling strategies, and a simulator is used to evaluate their performance. As a proof of concept, we show how the proposed methodology was used to propose and study a new workload‐aware loop scheduling strategy named smart round‐robin (SRR). We implemented this strategy into GNU Compiler Collection's OpenMP runtime. We carry out several experiments to validate the simulator and to evaluate the performance of SRR. Our experimental results show that SRR may deliver up to 37.89%and 14.10%better performance than OpenMP's dynamic loop scheduling strategy in the simulated environment and in a real‐world application kernel, respectively. Copyright © 2016 John Wiley & Sons, Ltd.
Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.2
2017 CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processors
abstract
Summary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd.
Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.7
2016 Seismic wave propagation simulations on low-power and performance-centric manycores
Márcio Castro 0001, Emilio Francesquini, Fabrice Dupros, Hideo Aochi, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Parallel Comput.1
2015 On the energy efficiency and performance of irregular application executions on multicore, NUMA and manycore platforms
Emilio Francesquini, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna, Fabrice Dupros, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
J. Parallel Distributed Comput.2
2014 Saving energy by exploiting residual imbalances on iterative applications
abstract
The power consumption of High Performance Computing (HPC) systems is an increasing concern as large-scale systems grow in size and, consequently, consume more energy. In response to this challenge, we propose two variants of a new energy-aware load balancer that aim at reducing the energy consumption of parallel platforms running imbalanced scientific applications without degrading their performance. Our research combines dynamic load balancing with DVFS techniques in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experimental results with benchmarks and a real-world application presented energy savings of up to 32% with our fine-grained variant that performs per-core DVFS, and of up to 34% with our coarsegrained variant that performs per-chip DVFS.
Edson L. Padoin, Márcio Castro 0001, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
HiPC2
2014 Evaluating the Impact of Transactional Characteristics on the Performance of Transactional Memory Applications
abstract
Transactional Memory (TM) is reputed by many researchers to be a promising solution to ease parallel programming on multicore processors. This model provides the scalability of fine-grained locking while avoiding common issues of traditional mechanisms, such as deadlocks. During these almost twenty years of research, several TM systems and benchmarks have been proposed. However, TM is not yet widely adopted by the scientific community to develop parallel applications due to unanswered questions in the literature, such as "how to identify if a parallel application can exploit TM to achieve better performance?" or "what are the reasons of poor performances of some TM applications?". In this work, we contribute to answer those questions through a comparative evaluation of a set of TM applications on four different state- of-the-art TM systems. Moreover, we identify some of the most important TM characteristics that impact directly the performance of TM applications. Our results can be useful to identify opportunities for optimizations.
Fernando Rui, Márcio Castro 0001, Dalvan Griebler, Luiz Gustavo Fernandes
PDP2
2014 Energy Efficient Seismic Wave Propagation Simulation on a Low-Power Manycore Processor
abstract
Large-scale simulation of seismic wave propagation is an active research topic. Its high demand for processing power makes it a good match for High Performance Computing (HPC). Although we have observed a steady increase on the processing capabilities of HPC platforms, their energy efficiency is still lacking behind. In this paper, we analyze the use of a low-power manycore processor, the MPPA-256, for seismic wave propagation simulations. First we look at its peculiar characteristics such as limited amount of on-chip memory and describe the intricate solution we brought forth to deal with this processor's idiosyncrasies. Next, we compare the performance and energy efficiency of seismic wave propagation on MPPA-256 to other commonplace platforms such as general-purpose processors and a GPU. Finally, we wrap up with the conclusion that, even if MPPA-256 presents an increased software development complexity, it can indeed be used as an energy efficient alternative to current HPC platforms, resulting in up to 71% and 81% less energy than a GPU and a general-purpose processor, respectively.
Márcio Castro 0001, Fabrice Dupros, Emilio Francesquini, Jean-François Méhaut, Philippe Olivier Alexandre Navaux
SBAC-PAD1
2014 Adaptive thread mapping strategies for transactional memory applications
Márcio Castro 0001, Fabrício Góes, Jean-François Méhaut
J. Parallel Distributed Comput.1
2012 Dynamic Thread Mapping Based on Machine Learning for Transactional Memory Applications
Márcio Castro 0001, Fabrício Góes, Luiz Gustavo Fernandes, Jean-François Méhaut
Euro-Par1
2011 A machine learning-based approach for thread mapping on transactional memory applications
abstract
Thread mapping has been extensively used as a technique to efficiently exploit memory hierarchy on modern chip-multiprocessors. It places threads on cores in order to amortize memory latency and/or to reduce memory contention. However, efficient thread mapping relies upon matching application behavior with system characteristics. Particularly, Software Transactional Memory (STM) applications introduce another dimension due to its runtime system support. Existing STM systems implement several conflict detection and resolution mechanisms, which leads STM applications to behave differently for each combination of these mechanisms. In this paper we propose a machine learning-based approach to automatically infer a suitable thread mapping strategy for transactional memory applications. First, we profile several STM applications from the STAMP benchmark suite considering application, STM system and platform features to build a set of input instances. Then, such data feeds a machine learning algorithm, which produces a decision tree able to predict the most suitable thread mapping strategy for new unobserved instances. Results show that our approach improves performance up to 18.46% compared to the worst case and up to 6.37% over the Linux default thread mapping strategy.
Márcio Castro 0001, Fabrício Góes, Christiane Pousa Ribeiro, Murray Cole, Marcelo Cintra, Jean-François Méhaut
HiPC1
2011 Analysis and Tracing of Applications Based on Software Transactional Memory on Multicore Architectures
abstract
Transactional Memory (TM) is a new programming paradigm that offers an alternative to traditional lock-based concurrency mechanisms. It offers a higher-level programming interface and promises to greatly simplify the development of correct concurrent applications on multicore architectures. However, simplicity often comes with an important performance deterioration and given the variety of TM implementations it is still a challenge to know what kind of applications can really take advantage of TM. In order to gain some insight on these issues, helping developers to understand and improve the performance of TM applications, we propose a generic approach for collecting and tracing relevant information about transactions. Our solution can be applied to different Software Transactional Memory (STM) libraries and applications as it does not modify neither the target application nor the STM library source codes. We show that the collected information can be helpful in order to comprehend the performance of TM applications.
Márcio Castro 0001, Kiril Georgiev, Vania Marangozova-Martin, Jean-François Méhaut, Luiz Gustavo Fernandes, Miguel Santana
PDP1
2009 NUMA-ICTM: A parallel version of ICTM exploiting memory placement strategies for NUMA machines
abstract
In geophysics, the appropriate subdivision of a region into segments is extremely important. ICTM (interval categorizer tesselation model) is an application that categorizes geographic regions using information extracted from satellite images. The categorization of large regions is a computational intensive problem, what justifies the proposal and development of parallel solutions in order to improve its applicability. Recent advances in multiprocessor architectures lead to the emergence of NUMA (non-uniform memory access) machines. In this work, we present NUMA-ICTM: a parallel solution of ICTM for NUMA machines. First, we parallelize ICTM using OpenMP. After, we improve the OpenMP solution using the MAI (memory affinity interface) library, which allows a control of memory allocation in NUMA machines. The results show that the optimization of memory allocation leads to significant performance gains over the pure OpenMP parallel solution.
Márcio Castro 0001, Luiz Gustavo Fernandes, Christiane Pousa Ribeiro, Jean-François Méhaut, Marilton S. de Aguiar
IPDPS1
2009 Memory Affinity for Hierarchical Shared Memory Multiprocessors
abstract
Currently, parallel platforms based on large scale hierarchical shared memory multiprocessors with Non-Uniform Memory Access (NUMA) are becoming a trend in scientific High Performance Computing (HPC). Due to their memory access constraints, these platforms require a very careful data distribution. Many solutions were proposed to resolve this issue. However, most of these solutions did not include optimizations for numerical scientific data (array data structures) and portability issues. Besides, these solutions provide a restrict set of memory policies to deal with data placement. In this paper, we describe an user-level interface named Memory Affinity interface (MAi), which allows memory affinity control on Linux based cache-coherent NUMA (ccNUMA) platforms. Its main goals are, fine data control, flexibility and portability. The performance of MAi is evaluated on three ccNUMA platforms using numerical scientific HPC applications, the NAS Parallel Benchmarks and a Geophysics application. The results show important gains (up to 31\%) when compared to Linux default solution.
Christiane Pousa Ribeiro, Jean-François Méhaut, Alexandre Carissimi, Márcio Castro 0001, Luiz Gustavo Fernandes
SBAC-PAD4