Pedro Henrique de Mello Morado Penna

dblp:210/3783 · also Pedro Henrique Penna · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0003-3617-2915ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 Cornflakes: Zero-Copy Serialization for Microsecond-Scale Networking
abstract
Data serialization is critical for many datacenter applications, but the memory copies required to move application data into packets are costly. Recent zero-copy APIs expose NIC scatter-gather capabilities, raising the possibility of offloading this data movement to the NIC. However, as the memory coordination required for scatter-gather adds bookkeeping overhead, scatter-gather is not always useful. We describe Cornflakes, a hybrid serialization library stack that uses scatter-gather for serialization when it improves performance and falls back to memory copies otherwise. We have implemented Cornflakes within a UDP and TCP networking stack, across Mellanox and Intel NICs. On a Twitter cache trace, Cornflakes achieves 15.4% higher throughput than prior software approaches on a custom key-value store and 8.8% higher throughput than Redis serialization within Redis.
Deepti Raghavan, Shreya Ravi, Gina Yuan, Pratiksha Thaker, Sanjari Srivastava, Micah Murray, Pedro Henrique de Mello Morado Penna, Amy Ousterhout, Philip Alexander Levis, Matei Zaharia, Irene Zhang
SOSP7
2023 LWMPI: An MPI library for NoC-based lightweight manycore processors with on-chip memory constraints
abstract
Abstract Lightweight manycore processors deliver high performance and energy efficiency by bundling hundreds of low‐power cores, a distributed memory architecture with small local memories and Networks‐on‐Chip in a single die. However, the lack of rich and portable programming models for these processors makes software development a challenging task. Currently, two approaches are employed to address programmability in lightweight manycores: Operating Systems (OSes) and baremetal runtime libraries. The former provides portability but exposes complex Operating System (OS)‐level programming interfaces to developers. The latter focuses on providing rich and high performance interfaces, which are vendor‐specific and yield to non‐portable software. In this work, we address these programmability and portability challenges by combining a rich OS with a well‐known standard for parallel programming. We propose a portable and lightweight Message Passing Interface (MPI) library (LWMPI) designed from scratch to cope with restrictions and intricacies of lightweight manycores. We integrated LWMPI into Nanvix, an open‐source distributed OS that runs on silicon lightweight manycores. The results obtained with a synthetic benchmark and a subset of the CAP Bench applications running on Kalray MPPA‐256 unveil that LWMPI not only delivers a lightweight and richer programming interface but also presents good performance and scalability results.
João Fellipe Uller, João Vicente Souto, Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
Concurr. Comput. Pract. Exp.3
2021 A Task-based Execution Engine for Distributed Operating Systems Tailored to Lightweight Manycores with Limited On-Chip Memory
abstract
Operating Systems (OSes) for lightweight manycore processors feature a distributed design, where isolated OS instances cooperate to improve programmability and portability issues that come from their architectural intricacies. However, OS services often resort to processes or threads to implement kernel-level functionalities, consuming much of the already limited on-chip memory available in these processors. In this context, we propose a complementary OS-level execution engine that supports cooperative time-sharing lightweight tasks that share a unique execution stack and features task synchronization via task dependency graphs. This solution provides numerous OS-level execution flows with reduced memory consumption, leaving more room for user-level applications. We implemented our solution in an open-source distributed OS (Nanvix) and we compared it with the standard implementation that uses threads. The results obtained on Kalray MPPA-256 show that our engine: (i) provides 63.2x more execution flows per MB of memory; (ii) features 15.8x and 1.87x less overhead to spawn and destroy an execution flow, respectively, while consuming 6.65x less energy; (iii) runs remote system calls 3x faster; and (iv) reduces the memory footprint in 0.58x when executing OS services requests without impacting the overall system performance.
João Vicente Souto, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna
SBAC-PAD3
2021 The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter Systems
abstract
Datacenter systems and I/O devices now run at single-digit microsecond latencies, requiring ns-scale operating systems. Traditional kernel-based operating systems impose an unaffordable overhead, so recent kernel-bypass OSes [73] and libraries [23] eliminate the OS kernel from the I/O datapath. However, none of these systems offer a general-purpose datapath OS replacement that meet the needs of μs-scale systems.' [email protected] paper proposes Demikernel, a flexible datapath OS and architecture designed for heterogenous kernel-bypass devices and μs-scale datacenter systems. We build two prototype Demikernel OSes and show that minimal effort is needed to port existing μs-scale systems. Once ported, Demikernel lets applications run across heterogenous kernel-bypass devices with ns-scale overheads and no code changes.
Irene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk, Jacob Nelson 0001, Omar S. Navarro Leija, Ashlie Martinez, Anna Kornfeld Simpson, Sujay Jayakar, Pedro Henrique de Mello Morado Penna, Max Demoulin, Piali Choudhury, Anirudh Badam
SOSP11
2021 Inter-kernel communication facility of a distributed operating system for NoC-based lightweight manycores
Pedro Henrique de Mello Morado Penna, João Vicente Souto, João Fellipe Uller, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut
J. Parallel Distributed Comput.1
2019 Teaching Parallel Programming to Freshmen in an Undergraduate Computer Science Program
abstract
This Research to Practice Full Paper proposes a teaching approach that introduces parallel programming early in the undergraduate Computer Science curriculum. Experiments were conducted to freshmen in the second course of algorithms and data structures. The strategy for the evaluation of the early education of parallel programming includes the use of OpenMP Application Programming Interface and sorting algorithms. The results indicate that students improved their skills by participating in parallel programing activities introduced at early stages or even at the very beginning of the undergraduate program. Freshmen could hit about 92%, 63% and 44% of easy, medium and hard questions after theoretical and practice activities. This represents an improvement about 19%, 14% and 39% for each respective difficulty level in comparison to the beginning of the study when all freshmen had no knowledge relative to parallel programming. These results aid to demystify parallel programming and to show that freshmen can learn it.
Leonardo B. A. Vasconcelos, Felipe A. L. Soares, Pedro Henrique de Mello Morado Penna, Max V. Machado, Fabrício Góes, Carlos Augusto Paiva da Silva Martins, Henrique Cota de Freitas
FIE3
2019 A comprehensive performance evaluation of the BinLPT workload-aware loop scheduler
abstract
Summary Workload‐aware loop schedulers were introduced to deliver better performance than classical loop scheduling strategies. However, they presented limitations such as inflexible built‐in workload estimators and suboptimal chunk scheduling. Targeting these challenges, we proposed previously a workload‐aware scheduling strategy called BinLPT, which relies on three features: (i) user‐supplied estimations of the workload of the loop; (ii) a greedy heuristic that adaptively partitions the iteration space in several chunks; and (iii) a scheduling scheme based on the Longest Processing Time (LPT) rule and on‐demand technique. In this paper, we present two new contributions to the state‐of‐the‐art. First, we introduce a multiloop support feature to BinLPT, which enables the reuse of estimations across loops. Based on this feature, we integrated BinLPT into a real‐world elastodynamics application, and we evaluated it running on a supercomputer. Second, we present an evaluation of BinLPT using simulations as well as synthetic and application kernels. We carried out this analysis on a large‐scale NUMA machine under a variety of workloads. Our results revealed that BinLPT is better at balancing the workloads of the loop iterations and this behavior improves as the algorithmic complexity of the loop increases. Overall, BinLPT delivers up to 37.15% and 9.11% better performance than well‐known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively.
Pedro Henrique de Mello Morado Penna, Antônio Tadeu A. Gomes, Márcio Castro 0001, Patricia Della Méa Plentz, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.1
2018 Energy Efficient Parallel K-Means Clustering for an Intel® Hybrid Multi-Chip Package
abstract
FPGA devices have been proving to be good candidates to accelerate applications from different research topics. For instance, machine learning applications such as K-Means clustering usually relies on large amount of data to be processed, and, despite the performance offered by other architectures, FPGAs can offer better energy efficiency. With that in mind, Intel has launched a platform that integrates a multicore and an FPGA in the same package, enabling low latency and coherent fine-grained data offload. In this paper, we present a parallel implementation of the K-Means clustering algorithm, for this novel platform, using OpenCL language, and compared it against other platforms. We found that the CPU+FPGA platform was more energy efficient than the CPU-only approach from 70.71% to 85.92%, with Standard and Tiny input sizes respectively, and up to 68.21% of performance improvement was obtained with Tiny input size. Furthermore, it was up to 7.2×more energy efficient than an Intel® Xeon Phi ™, 21.5×than a cluster of Raspberry Pi boards, and 3.8×than the low-power MPPA-256 architecture, when the Standard input size was used.
Matheus Alcântara Souza, Lucas Andrade Maciel, Pedro Henrique de Mello Morado Penna, Henrique Cota de Freitas
SBAC-PAD3
2017 Design methodology for workload-aware loop scheduling strategies based on genetic algorithm and simulation
abstract
Summary In high‐performance computing, the application's workload must be evenly balanced among threads to deliver cutting‐edge performance and scalability. In OpenMP, the load balancing problem arises when scheduling loop iterations to threads. In this context, several scheduling strategies have been proposed, but they do not take into account the input workload of the application and thus turn out to be suboptimal. In this work, we introduce a design methodology to propose, study, and assess the performance of workload‐aware loop scheduling strategies. In this methodology, a genetic algorithm is employed to explore the state space solution of the problem itself and to guide the design of new loop scheduling strategies, and a simulator is used to evaluate their performance. As a proof of concept, we show how the proposed methodology was used to propose and study a new workload‐aware loop scheduling strategy named smart round‐robin (SRR). We implemented this strategy into GNU Compiler Collection's OpenMP runtime. We carry out several experiments to validate the simulator and to evaluate the performance of SRR. Our experimental results show that SRR may deliver up to 37.89%and 14.10%better performance than OpenMP's dynamic loop scheduling strategy in the simulated environment and in a real‐world application kernel, respectively. Copyright © 2016 John Wiley & Sons, Ltd.
Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut
Concurr. Comput. Pract. Exp.1
2017 CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processors
abstract
Summary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd.
Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.2
2015 On the energy efficiency and performance of irregular application executions on multicore, NUMA and manycore platforms
Emilio Francesquini, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna, Fabrice Dupros, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
J. Parallel Distributed Comput.3