Philippe Olivier Alexandre Navaux

dblp:74/1262 · also Philippe O. A. Navaux · DBLP profile ↗
← Back
168ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0002-9957-5861ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 82 · 2 first-author · 14 since 2021Computer networks · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Artificial intelligence and machine learning · 3Security and privacy · 3 · 1 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Improving Cost-Performance Efficiency of Scientific Cloud Workflows through GPU Sharing
Matheus M. Costa, Tiago Ferreto, César A. F. De Rose, Odej Kao, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
CLOSER5
2026 When Faster Kernels Do Not Mean Faster Applications in Exascale GPU Systems
Mariana Toledo Costa, Antigoni Georgiadou, James B. White, Woong Shin, Bruno Villasenor Alvarez, Jorda Polo, Karl W. Schulz, Philippe Olivier Alexandre Navaux, O. E. Bronson Messer, Arthur Francisco Lorenzon
Euro-Par (1)8
2026 Characterizing Lossless GPU Data Compression Across AMD CDNA and RDNA Architectures
Cristiano A. Künas, Gabriel Freytag, Jean Luca Bez, Thiago da Silva Araújo, Philippe Olivier Alexandre Navaux
ICCSA (1)5
2026 Evaluating Accuracy-Performance Trade-Offs in Deep Learning Frameworks for Diabetic Retinopathy Detection
Bruno Morales, Cristiano A. Künas, Rodrigo C. Machado, Pedro H. C. Rigon, Philippe Olivier Alexandre Navaux
ICCSA (1)5
2025 Energy-Aware Node Selection for Cloud-Based Parallel Workloads with Machine Learning and Infrastructure as Code
Denis B. Citadin, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
CLOSER4
2025 Towards Optimizing Cost and Performance for Parallel Workloads in Cloud Computing
William Maas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
CLOSER4
2025 One GPU, Many Ranks: Enabling Performance and Energy-Efficient In-Transit Visualization via Resource Sharing
abstract
In-transit visualization has become essential in high-performance computing (HPC) to reduce I/O overheads and enable real-time data analysis. However, as simulations grow in scale and complexity, these visualization tasks increasingly demand substantial computational resources, exacerbating energy consumption and limiting system scalability. As we show in this paper, a key bottleneck is the conventional one-rank-per-GPU allocation model, which leads to irregular GPU utilization and waste of hardware resources. To tackle this challenge, we propose using GPU-sharing strategies to improve energy efficiency in in-transit visualization without compromising performance. We evaluate six distinct configurations built upon three NVIDIA GPU-sharing mechanisms: the default CUDA model with context switching between processes, Multi-Process Service (MPS), which enables dynamic context sharing, and Multi-Instance GPU (MIG), which provides hardware-level partitioning. Using the WarpX simulation code and Ascent visualization framework, our experiments on the Polaris supercomputer span multiple rendering techniques, node counts, and data sizes. Results show that GPU-sharing strategies can improve the trade-off between performance and energy, represented by the energy-delay product (EDP) metric, by up to 81.7%. We also show that workload-aware strategy selection is essential to improve performance-energy efficiency: MIG-based configurations are more effective for lightweight and regular workloads, offering up to 64.5% energy savings, while MPS better handles GPU-intensive workloads, achieving up to 71.1% EDP improvement. Finally, we demonstrate that optimized sharing strategies can reduce the required compute nodes by up to 75%, freeing system resources for concurrent workloads.
Matheus M. Costa, Philippe Olivier Alexandre Navaux, Silvio Rizzi 0001, Arthur Francisco Lorenzon
ICPP2
2025 Scalable and Efficient Deep Learning for Diabetic Retinopathy Classification on ARM
abstract
Diabetic retinopathy (DR) diagnosis delays pose a critical challenge for public healthcare systems such as the Brazilian Unified Health System (SUS), where long referral queues often prevent timely treatment and increase the risk of vision loss. Deep learning (DL) models offer an effective solution by automating retinal image analysis, but choosing an appropriate model requires balancing diagnostic accuracy with computational and energy efficiency. This study evaluates 38 convolutional neural networks (CNN) across four key dimensions: Area under the curve (AUC), energy consumption, model size, and training time. Our analysis identifies MobileNet as the superior architecture, demonstrating 77% lower energy use, 83% faster training, and 85% smaller model size than the InceptionV3 baseline, while achieving 3% higher AUC. We further optimize MobileNet through systematic hyperparameter tuning and evaluate its scalability on the ARM-based NVIDIA Grace Superchip, revealing peak efficiency at 36-thread configurations where energy use, CPU utilization, and memory access patterns reach optimal balance. All implementation scripts are publicly available to foster reproducible, sustainable AI development for clinical applications.
Thiago da Silva Araújo, Beatriz Schaan, Carla M. D. S. Freitas, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2025 Efficient Multi-Workload Execution for Sustainable GPU Performance
abstract
Modern scientific research often relies on powerful computing systems that use graphics processing units (GPUs) to run complex applications. However, running these systems requires a large amount of energy, which contributes to carbon emissions and raises concerns about environmental impact. Given this scenario, we explore how sharing a single GPU between multiple applications can improve both performance and sustainability when running scientific workflows. We consider three execution strategies: running applications one after another, running two at the same time, and replacing finished tasks with new ones right away, using eighteen widely used scientific applications on three different GPUs (AMD MI250X, AMD RX 7900XT, and NVIDIA RTX 4090). To demonstrate that finding the best co-execution combination of applications improves resource efficiency, we use a mathematical approach based on linear programming to schedule which applications run together. Our results show that this approach can reduce total execution time by up to 47% and lower carbon emissions by as much as 34%, with minimal impact on the performance of individual applications. Additionally, when optimal combinations of parallel applications are used, the overall performance of a complete scientific workflow can improve by 36%, while carbon emissions are reduced by 25%.
Matheus M. Costa, Philippe Olivier Alexandre Navaux, Silvio Rizzi 0001, O. E. Bronson Messer, Arthur Francisco Lorenzon
SBAC-PAD2
2025 Evaluating Code Portability for Carbon-Efficient RTM Computing
abstract
As GPU architectures continue to diversify across high-performance computing (HPC) systems, ensuring code portability and minimizing environmental impact have become critical challenges. This paper investigates how different programming models affect the carbon-performance efficiency of a Reverse Time Migration (RTM) application, a key workload in geophysics. We provide twelve implementations of the RTM code using CUDA, HIP, Kokkos, RAJA, and OpenMP Target, and evaluate their behavior on eleven GPUs from NVIDIA and AMD. Our analysis covers execution time, energy consumption, and carbon footprint. Results show that HIP, when tuned per architectures and RAJA, achieve the highest code portability in terms of carbon-efficiency, reaching up to 93.7% and 93.4% efficiency across all platforms, respectively. Overall, while HIP and CUDA deliver peak performance and the lowest emissions when properly optimized, the gap between these native models and high-level abstractions such as RAJA and Kokkos is steadily narrowing, indicating growing potential for portable and sustainable HPC development.
Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux, Alexandre Sardinha, O. E. Bronson Messer
SBAC-PAD2
2024 An ANN-Guided Multi-Objective Framework for Power-Performance Balancing in HPC Systems
abstract
Power-performance efficiency has become one of the most critical issues in evolving High-Performance Computing systems (HPC) towards Exaflops. Thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are methods widely applied to better balance the power consumption and performance improvements of parallel applications. However, selecting ideal combinations of these knobs for every application is challenging due to the massive number of possible solutions, as there is no unique combination that delivers at the same time the best performance and the lowest power consumption. Given that, we propose HPC-PPO (power-performance optimizer), a multi-objective optimization strategy driven by an artificial neural network that leverages hardware and software features of parallel applications to predict Pareto-efficient configurations of TLP degree, DVFS, and UFS that optimize the balance between power and performance. When validating HPC-PPO on three multicore processors with twenty-five applications, we show that HPC-PPO can predict combinations very close to the best ones found by an exhaustive search. We also show that the Pareto-efficient configurations predicted by HPC-PPO improve parallel applications' performance by 30.7% while spending 23.9% less power when compared to state-of-the-art strategies.
William Maas, Paulo Silas Severo de Souza, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
CF5
2024 Harnessing Data Movement Strategies to Optimize Performance-Energy Efficiency of Oil & Gas Simulations in HPC
Pedro H. C. Rigon, Brenda S. Schussler, Alexandre Sardinha, Pedro M. Silva, Fábio Oliveira, Alexandre Carissimi, Jairo Panetta, Filippo Spiga, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux
Euro-Par (2)10
2024 Interleaved Execution of Approximated CUDA Kernels in Iterative Applications
abstract
Fine-tuning the floating-point precision of arithmetic operations in applications can be extremely challenging and time-consuming, especially in iterative applications where the output of one iteration serves as the input for the subsequent iteration. Consequently, the accuracy loss can be magnified throughout the execution. Therefore, we propose an alternative approach based on the interleaved execution of multiple approximated kernel versions designed with different precision levels. We demonstrate that creating an interleaved execution configuration of multiple CUDA kernel versions based on their accuracy loss profiles enables us to enhance performance, improve energy efficiency, and manage the accuracy loss of scientific simulation applications in various Target Output Quality (TOQ) scenarios. For a TOQ loss of approximately 3 %, we achieve a speedup of up to 1. 7x and reduce energy consumption by nearly 40 %.
Gabriel Freytag, Cristiano A. Künas, Paolo Rech, Philippe Olivier Alexandre Navaux
PDP4
2024 BTO, Block and Thread Optimization of GPU Kernels on Geophysical Exploration
abstract
The pursuit of performance and energy efficiency of geophysical exploration applications on high-performance computing (HPC) servers has been driving the optimization of hardware resource usage in graphic processing units (GPUs). On such architectures, the execution configuration of each kernel (e.g., the number of blocks and threads per block) plays an essential role in the performance and energy consumption of these applications. However, as we show in this paper, due to the massive number of possible configurations of the number of blocks and threads per block, leveraging solely on the software developer to define the configuration execution for every GPU kernel does not lead to an ideal usage of GPU hardware resources, leading to performance loss and an increase on the energy consumption. To tackle this challenge, we propose BTO, a block and thread optimization strategy driven by a genetic algorithm. It cooperatively optimizes the number of blocks and thread per block for every GPU kernel at runtime with minimum convergence overhead regardless of the number of GPUs available on the system. When employing BTO to optimize the Fletcher modeling, a representative geophysical exploration application, on different AMD and NVIDIA GPUs, we show that BTO improves the energy-delay product (EDP - tradeoff between performance and energy) by up to 83.8% and 81.9% over the standard execution of Fletcher and the default execution of GPU applications on the target architectures. Moreover, by comparing it to an exhaustive search, we show that BTO converges to optimal execution configurations in 96.4% of all evaluated scenarios.
Brenda S. Schussler, Pedro H. C. Rigon, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux
PDP4
2024 Towards Performance Portability of an Oil and Gas Application on Heterogeneous Architectures
abstract
With the widespread use of graphic processing units (GPUs) from different vendors in oil and gas companies, it has become increasingly important to achieve code portability. This allows for evaluating performance across diverse GPU vendors, enabling informed decisions. This capability enables companies to assess hardware solutions based on cost, performance, or energy efficiency, promoting competition and potentially driving innovation and cost reduction in GPU technologies. With that in mind, we address the challenge of achieving performance portability for Fletcher, an anisotropic wave propagation modeling application across various GPU architectures from NVIDIA and AMD. Hence, our first contribution is the parallel implementation of Fletcher on GPUs using eleven variations of portable programming models, including HIP, CUDA, Kokkos, RAJA, and OpenMP Target. Our second contribution is a comprehensive evaluation of these implementations across five generations of GPU from NVIDIA and AMD, assessing performance, performance-portability, and power-performance efficiency. Through an extensive set of experiments, we demonstrate that HIP outperforms the assessed programming models in terms of performance portability across all evaluated GPUs. Specifically, it performs 7.9% better than Kokkos, 8.8% better than RAJA, and 67.8% better than OpenMP Target. We also show that while the automatic translation of CUDA to HIP code allows for execution on AMD GPUs, optimizing for high performance is necessary. Additionally, we show that the programming models’ impact on Fletcher’s performance is heavily influenced by the size of the GPU’s L2 cache, especially given that this is a memory-intensive application.
Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux, Alexandre Sardinha, O. E. Bronson Messer
SBAC-PAD2
2024 HBPB, applying reuse distance to improve cache efficiency proactively
Arthur M. Krause, Paulo C. Santos 0001, Arthur Francisco Lorenzon, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.4
2023 Accelerating Deep Learning Model Training on Cloud Tensor Processing Unit
Cristiano A. Künas, Edson L. Padoin, Philippe Olivier Alexandre Navaux
CLOSER3
2023 NeurOPar, A Neural Network-Driven EDP Optimization Strategy for Parallel Workloads
abstract
The pursuit of energy efficiency has been driving the development of techniques to optimize hardware resource usage in high-performance computing (HPC) servers. On multicore architectures, thread-level parallelism (TLP) exploitation, dynamic voltage and frequency scaling (DVFS), and uncore frequency scaling (UFS) are three popular methods applied to improve the trade-off between performance and energy consumption, represented by the energy-delay product (EDP). However, the complexity of selecting the optimal configuration (TLP degree, DVFS, and UFS) for each application poses a challenge to software developers and end-users due to the massive number of possible configurations. To tackle this challenge, we propose NeurOpar, an optimization strategy for parallel workloads driven by an artificial neural network (ANN). It uses representative hardware and software metrics to build and train an ANN model that predicts combinations of thread count and core/uncore frequency levels that provide optimal EDP results. Through experiments on four multicore processors using twenty-five applications, we demonstrate that NeurOPar predicts combinations that yield EDP values close to the best ones achieved by an exhaustive search and improve the overall EDP by 42% compared to the default execution of HPC applications. We also show that NeurOPar can enhance the execution of parallel applications without incurring the performance and energy penalties associated with online methods by comparing it with two state-of-the-art strategies.
Cristiano A. Künas, Fábio D. Rossi, Marcelo Caggiani Luizelli, Rodrigo N. Calheiros, Philippe Olivier Alexandre Navaux, Arthur Francisco Lorenzon
SBAC-PAD5
2023 Mitigating execution unit contention in parallel applications using instruction-aware mapping
abstract
Summary Parallel applications running on simultaneous multithreading (SMT) processors naturally compete for execution units when their threads are mapped to the same core. This issue is further aggravated when such threads execute similar instructions that stress the same execution unit type, making their execution to behave very similarly as if the threads were running sequentially. This, in turn, will lead to performance degradation and underutilization of hardware resources. This work proposes a completely transparent framework (no modifications to the source code are necessary) that automatically maps threads of multiple parallel applications on SMT processors. The framework focuses on improving performance by mitigating the contention on execution units, considering each thread's instruction types, which are detected at runtime by our framework. Results show performance gains of 21% (geometric mean), compared to the native scheduler of the operating system.
Matheus S. Serpa, Eduardo Henrique Molina da Cruz, Matthias Diener, Arthur Francisco Lorenzon, Antonio Carlos Schneider Beck, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.6
2023 Smart resource allocation of concurrent execution of parallel applications
abstract
Abstract Thread‐level parallelism (TLP) has been widely exploited to optimize computational resource usage in high‐performance systems. However, as many applications do not scale as the number of threads increase, resources will be wasted when the application executes with the maximum possible number of threads (i.e., the default execution) rather than fewer threads (thread throttling) that may use the resources more efficiently. Hence, instead of executing only one application with as many threads as possible, one can run more applications simultaneously by applying thread throttling to each one. The primary outcome of this strategy is a significant reduction in the total execution time and energy consumption when the system needs to execute a list of applications. Given that, we propose a smart resource allocation (SRA) for concurrent parallel application execution. It automatically finds the ideal degree of TLP for each application and guides the simultaneous parallel applications execution. When running 25 well‐known benchmarks on three multicore systems and comparing SRA to state‐of‐the‐art strategies (e.g., Batch, Equal policy, and Scalability), SRA improves the EDP by 87.4% over the Batch strategy; 75.5% over the Equal policy; and 38.8% over the scalability strategy.
Vinicius S. da Silva, Angelo Gaspar Diniz Nogueira, Everton Camargo de Lima, Hiago Rocha, Matheus S. Serpa, Marcelo Caggiani Luizelli, Fábio D. Rossi, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon
Concurr. Comput. Pract. Exp.8
2023 Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos Dumont
abstract
High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.5
2022 Avoiding Unnecessary Caching with History-Based Preemptive Bypassing
abstract
Cache memories can account for more than half of the area and energy consumption on modern processors, which will only increase with the current trend of bigger on die memories. Although these components are very effective when the access pattern is cache-friendly, cache memories incur extra and unnecessary latencies when they cannot serve the data, which adds to significant energy wastes when data that is never reused is placed on them. This work introduces HBPB, a mechanism that detects whether a memory access is cache friendly or not, allowing the bypass of the cache for accesses that are not known to be cache-friendly. Our approach allows the processor to quickly detect when caching accesses is inadequate, improving overall access latency and reducing energy waste and cache pollution. The presented solution achieves reductions of up to 28.6% in energy consumption and 19.5% in latency for SPEC applications, and further improvements in power and performance across various workloads.
Arthur M. Krause, Paulo C. Santos 0001, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2022 Optimizing the EDP of OpenMP applications via concurrency throttling and frequency boosting
Sandro Matheus V. N. Marques, Matheus S. Serpa, Antoni Navarro Muñoz, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon
J. Syst. Archit.6
2022 Terminator: A Secure Coprocessor to Accelerate Real-Time AntiViruses Using Inspection Breakpoints
abstract
AntiViruses (AVs) are essential to face the myriad of malware threatening Internet users. AVs operate in two modes: on-demand checks and real-time verification. Software-based real-time AVs intercept system and function calls to execute AV’s inspection routines, resulting in significant performance penalties as the monitoring code runs among the suspicious code. Simultaneously, dark silicon problems push the industry to add more specialized accelerators inside the processor to mitigate these integration problems. In this article, we propose Terminator , an AV-specific coprocessor to assist software AVs by outsourcing their matching procedures to the hardware, thus saving CPU cycles and mitigating performance degradation. We designed Terminator to be flexible and compatible with existing AVs by using YARA and ClamAV rules. Our experiments show that our approach can save up to 70 million CPU cycles per rule when outsourcing on-demand checks for matching typical, unmodified YARA rules against a dataset of 30 thousand in-the-wild malware samples. Our proposal eliminates the AV’s need for blocking the CPU to perform full system checks, which can now occur in parallel. We also designed a new inspection breakpoint mechanism that signals to the coprocessor the beginning of a monitored region, allowing it to scan the regions in parallel with their execution. Overall, our mechanism mitigated up to 44% of the overhead imposed to execute and monitor the SPEC benchmark applications in the most challenging scenario.
Marcus Botacin, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, André Ricardo Abed Grégio, Marco A. Z. Alves
ACM Trans. Priv. Secur.3
2021 Harnessing Cloud Computing to Power Up HPC Applications: The BRICS CloudHPC Project
Jonatas Adilson Marques, Zhongke Wu, Xingce Wang, Ruslan Kuchumov, Vladimir Korkhov, Weverton Luis da Costa Cordeiro, Philippe Olivier Alexandre Navaux, Luciano Paschoal Gaspary
ICCSA (8)7
2021 Arbitration Policies for On-Demand User-Level I/O Forwarding on HPC Platforms
abstract
I/O forwarding is a well-established and widely-adopted technique in HPC to reduce contention in the access to storage servers and transparently improve I/O performance. Rather than having applications directly accessing the shared parallel file system, the forwarding technique defines a set of I/O nodes responsible for receiving application requests and forwarding them to the file system, thus reshaping the flow of requests. The typical approach is to statically assign I/O nodes to applications depending on the number of compute nodes they use, which is not always necessarily related to their I/O requirements. Thus, this approach leads to inefficient usage of these resources. This paper investigates arbitration policies based on the applications I/O demands, represented by their access patterns. We propose a policy based on the Multiple-Choice Knapsack problem that seeks to maximize global bandwidth by giving more I/O nodes to applications that will benefit the most. Furthermore, we propose a user-level I/O forwarding solution as an on-demand service capable of applying different allocation policies at runtime for machines where this layer is not present. We demonstrate our approach's applicability through extensive experimentation and show it can transparently improve global I/O bandwidth by up to 85% in a live setup compared to the default static policy.
Jean Luca Bez, Alberto Miranda, Ramon Nou, Francieli Zanon Boito, Toni Cortes, Philippe Olivier Alexandre Navaux
IPDPS6
2021 Lightweight Deep Learning Applications on AVX-512
abstract
Machine Learning and Deep Learning applications are of paramount importance these days. Different areas of academia and industry use daily workloads based on these applications. Several aspects are relevant regarding their applicability, such as the complexity and accuracy of the models and their performance and energy efficiency. Currently, there is a trend to usually favor the use of GPUs to train and execute Deep Learning models, intensified by specialized hardware. However, this article demonstrates that using a CPU with AVX-512 instructions can achieve comparable performance to current GPUs and, depending on the workload, suppress it by ≈ 1.8x.
Andre Ramos Carneiro, Matheus S. Serpa, Philippe Olivier Alexandre Navaux
ISCC3
2021 Combining Thread Throttling and Mapping to Optimize the EDP of Parallel Applications
abstract
Thread-throttling and mapping strategies have been used together to make better use of hardware resources and improve the energy-delay product (EDP) of high-performance computing (HPC) systems. However, the design space exploration significantly grows with the increasing number of cores in those systems, making the task of finding the ideal number of active threads and allocating strategy a challenging task. On top of that, parallel applications present various patterns, such as irregularity, unbalanced computations, or high rates of communications. Given these considerations, we propose ETTM, an EDPaware thread-throttling and mapping optimization strategy that automatically finds an ideal combination of number of threads and thread mapping strategy. With the execution of eighteen well-known benchmarks on three multicore architectures, we show that EDP can be significantly improved when running applications with the solution found by EETM1.
Gustavo Berned, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon
PDP6
2021 Optimizing Parallel Applications via Dynamic Concurrency Throttling and Turbo Boosting
abstract
With the increasing number of cores in modern systems, dynamic concurrency throttling (DCT) and turbo-boosting techniques are becoming a solution to better use the hardware resources. While DCT techniques tune the number of running threads, boosting techniques speed up sequential phases or unbalanced threads. However, as each region of an application may behave differently, optimizing both knobs is not straightforward. Hence, we propose two strategies that apply DCT and turbo-boosting: DBF, which aims to find an ideal configuration for each parallel/sequential region, and DBC, which considers the combination of parallel/sequential regions during the optimization. We show that DBF and DBC improve the EDP by up to 19% and 27% compared to a DCT-only strategy and by up to 95% and 96% compared to a Boost-only technique. We also show that DBF is more suitable for applications with high variability in the CPU workload, while DBC is better when there is low workload variability.
Sandro Matheus V. N. Marques, Thiarles S. Medeiros, Matheus S. Serpa, Fábio D. Rossi, Marcelo Caggiani Luizelli, Philippe Olivier Alexandre Navaux, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon
PDP6
2021 HPC Data Storage at a Glance: The Santos Dumont Experience
abstract
High-Performance Computing (HPC) platforms are used to solve the most diverse scientific problems in research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific software, which have different requirements. These requirements include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle a mixed workload scenario when storing data from the applications. Knowledge of the application set and its performance running in a supercomputer is needed to understand the storage system's usage, pinpoint possible bottlenecks, and guide optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, where we were able to identify inefficient usage and problematic factors of performance.
Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2021 Investigating memory prefetcher performance over parallel applications: From real to simulated
abstract
Abstract Memory prefetcher algorithms are widely used in processors to mitigate the performance gap between the processors and the memory subsystem. The complexities behind the architectures and prefetcher algorithms, however, not only hinder the development of accurate architecture simulators, but also hinder understanding the prefetcher's contribution to performance, on both a real hardware and in a simulated environment. In this paper, we contribute to shed light on the memory prefetcher's role in the performance of parallel High‐Performance Computing applications, considering the prefetcher algorithms offered by both the real hardware and the simulators. We performed a careful experimental investigation, executing the NAS parallel benchmark (NPB) on a real Skylake machine, and as well in a simulated environment with the ZSim and Sniper simulators, taking into account the prefetcher algorithms offered by both Skylake and the simulators. Our experimental results show that: (i) prefetching from the L3 to L2 cache presents better performance gains, (ii) the memory contention in the parallel execution constrains the prefetcher's effect, (iii) Skylake's parallel memory contention is poorly simulated by ZSim and Sniper, and (iv) Skylake's noninclusive L3 cache hinders the accurate simulation of NPB with the Sniper's prefetchers.
Valéria Soldera Girelli, Francis B. Moreira 0001, Matheus S. Serpa, Danilo Carastan-Santos, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.5
2021 Energy efficiency and portability of oil and gas simulations on multicore and graphics processing unit architectures
abstract
Summary Reverse time migration (RTM) simulation is the basis of the seismic imaging tools used by the oil and gas industry. Developers have been porting their simulations to the new high‐performance computing architectures, providing faster and more accurate results at each new generation. However, several challenges arrive when trying to achieve high performance on these new architectures. The first one is to choose the architecture that best fits the kind of simulation. After that, researchers should choose the API used to implement the simulation code. These two decisions are strongly related to the effort, performance, and energy efficiency of the simulations. In this article, we propose three optimizations for an oil and gas application, which reduce the floating‐point operations by changing the equation derivatives. We evaluate these optimizations in different multicore and GPU architectures, investigating the impact of different APIs on the performance, energy efficiency, and portability of the code. Our experimental results show that the dedicated CUDA implementation running on the NVIDIA Volta architecture has the best performance and energy efficiency for RTM on GPUs, while the OpenMP version is the best for Intel Broadwell in the multicore. Also, the OpenACC version, which has a lower programming effort and executes on both architectures, has an up to 20% better performance and energy efficiency than the nonportable ones.
Matheus S. Serpa, Pablo J. Pavan, Eduardo Henrique Molina da Cruz, Rodrigo L. Machado, Jairo Panetta, Antônio Azambuja, Alexandre Carissimi, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.8
2021 Collaborative execution of fluid flow simulation using non-uniform decomposition on heterogeneous architectures
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.5
2021 Thermal neutrons: a possible threat for supercomputer reliability
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Andrea Favalli, Opale Schappert, Stephen Wender, Carlo Cazzaniga, Christopher Frost 0002, Paolo Rech
J. Supercomput.6
2020 Thermal Neutrons: a Possible Threat for Supercomputers and Safety Critical Applications
abstract
The high performance, high efficiency, and low cost of Commercial Off-The-Shelf (COTS) devices make them attractive for applications with strict reliability constraints. Today, COTS devices are adopted in HPC and safety-critical applications such as autonomous driving. Unfortunately, the cheap natural Boron widely used in COTS chip manufacturing process makes them highly susceptible to thermal (low energy) neutrons. In this paper, we demonstrate that thermal neutrons are a significant threat to COTS device reliability. For our study, we consider an AMD APU, three NVIDIA GPUs, an Intel accelerator, and an FPGA executing a relevant set of algorithms. We consider different scenarios that impact the thermal neutron flux such as weather, concrete walls and floors, and HPC liquid cooling systems. We show that thermal neutrons FIT rate could be comparable to the high energy neutron FIT rate.
Daniel Oliveira 0002, Sean Blanchard, Nathan DeBardeleben, Fernando Santos 0001, Gabriel Piscoya Davila, Philippe Olivier Alexandre Navaux, Carlo Cazzaniga, Christopher Frost 0002, Robert C. Baumann, Paolo Rech
ETS6
2020 The Impact of CPU Frequency Scaling on Power Consumption of Computing Infrastructures
Adriano Marques Garcia, Matheus S. Serpa, Dalvan Griebler, Claudio Schepke, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux
ICCSA (6)6
2020 Performance Impact of IEEE 802.3ad in Container-Based Clouds for HPC Applications
Anderson M. Maliszewski, Eduardo Roloff, Dalvan Griebler, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
ICCSA (6)5
2020 Performance and Cost-aware HPC in Clouds: A Network Interconnection Assessment
abstract
The availability of computing resources has significantly changed due to the growing adoption of the cloud computing paradigm. Aiming at potential advantages such as cost savings through the pay-per-use method and resource allocation in a scalable/elastic way, we witnessed consistent efforts to execute high-performance computing (HPC) applications in the cloud. Performance in this environment depends heavily upon two main system components: processing power and network interconnection. If, on the one hand, allocating more powerful hardware theoretically boosts performance, on the other hand, it increases the allocation cost. In this paper, we evaluated how the network interconnection impacts on performance and cost efficiency. Our experiments were carried out using NAS Parallel Benchmarks and Alya HPC application on Microsoft Azure public cloud provider, with three different cloud instances/network interconnections. The results revealed that through the use of the accelerated networking approach, which allows the instance to have a high-performance interconnect without additional charges, the performance of HPC applications can be significantly improved with a better cost efficiency.
Anderson M. Maliszewski, Eduardo Roloff, Emmanuell D. Carreño, Dalvan Griebler, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
ISCC6
2020 Attesting L-3 General Program Anomaly Detection Efficiency with SPADA
abstract
One of the main challenges for security systems is the detection of general vulnerability exploitation, especially when the exploit uses valid control flow. Thus, the detection of anomalous behavior provides an exciting research direction, as the research in this field tries to describe what is the standard program execution, to then detect as anomalous any behavior that does not fit that description.In this work, we compare two mechanisms that aim to detect general anomalies: SPADA and LAD. SPADA is an L-3 language mechanism that partitions phases and uses simple phase features to detect anomalies. LAD is a constrained L-1 language mechanism that applies complex clustering and machine learning models on specific functions to detect anomalies. In our experimental campaign with several real-world exploits, we show that SPADA’s detection mechanism performs better than LAD while being much simpler and easier to implement. We therefore show experimental evidence that further attests the efficiency of L-3 attack detection mechanisms for real attacks.
Francis B. Moreira 0001, Danilo Carastan-Santos, Philippe Olivier Alexandre Navaux
ISCC3
2020 Task-based parallel strategies for computational fluid dynamic application in heterogeneous CPU/GPU resources
abstract
Summary Parallel applications executing in contemporary heterogeneous clusters are complex to code and optimize. The task‐based programming model is an alternative to handle the coding complexity. This model consists of splitting the problem domain into tasks with dependencies through a directed acyclic graph, and submit the set of tasks to a runtime scheduler that maps each task dynamically to resources. We consider that computational fluid dynamics applications are typical in scientific computing but not enough exploited by designs that employ the task‐based programming model. This article presents task‐based parallel strategies for a simple CFD application that targets heterogeneous multi‐CPU/multi‐GPU computing resources. We design, develop, evaluate, and compare the performance of three parallel strategies (naive, ghost‐cells, and arrow) of a task‐based heterogeneous (CPU and GPU) application that simulates the flow of an incompressible Newtonian fluid with constant viscosity. All implementations rely on the StarPU runtime, and we use the StarVZ toolkit to conduct comprehensive performance analysis. Results indicate that the ghost cell strategy provides the best speedup (77×) considering the simulation time when the GPU resources still have available memory. However, the arrow strategy achieves better results when the simulation data increases.
Lucas Leandro Nesi, Matheus S. Serpa, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.4
2020 Adaptive request scheduling for the I/O forwarding layer using reinforcement learning
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux
Future Gener. Comput. Syst.6
2019 SPADA: a statistical program attack detection analysis
abstract
One of the main challenges in system security is the detection of vulnerability exploitation, especially valid control flow exploitation. The specificity of state-of-the-art methods, such as signature-based detection, becomes a limiting factor when detecting the latest exploits and attacks uncovered.
Francis B. Moreira 0001, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux
CF3
2019 GaruaGeo: Global Scale Data Aggregation in Hybrid Edge and Cloud Computing Environments
Otávio Carvalho, Eduardo Roloff, Philippe Olivier Alexandre Navaux
CLOSER3
2019 Exploring Instance Heterogeneity in Public Cloud Providers for HPC Applications
Eduardo Roloff, Matthias Diener, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
CLOSER4
2019 Identifying the Most Reliable Collaborative Workload Distribution in Heterogeneous Devices
abstract
The constant need of higher performances and reduced power consumption has lead vendors to design heterogeneous devices that embed traditional CPU and an accelerator, like a GPU or FPGA. When the CPU and the accelerator are used collaboratively the device computational performances reach their peak. However, the higher amount of resources employed for computation has, potentially, the side effect of increasing soft error rate. In this paper we evaluate the reliability behavior of AMD Kaveri Accelerated Processing Units executing a set of heterogeneous applications. We distribute the workload between the CPU and GPU and evaluate which configuration provides the lowest error rate or allows the computation of the highest amount of data before experiencing a failure. We show that, in most cases, the most reliable workload distribution is the one that delivers the highest performances. As experimentally proven, by choosing the correct workload distribution the device reliability can increase of up to 9x.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
DATE3
2019 Impact of Reduced Precision in the Reliability of Deep Neural Networks for Object Detection
abstract
Modern Graphics Processing Units (GPUs) have dedicated hardware to execute floating-point operations with different precisions (64-bit double, 32-bit single, and 16-bit half). Using reduced precision for specific applications like Deep Neural Networks (DNNs) has been shown to reduce both the execution time and power consumption with negligible effects on the DNNs' accuracy. As GPUs are playing a critical role in DNN for object detection and get into safety-critical environments, their reliability is becoming a growing concern. In this paper, we evaluate the reliability of a DNN implemented in three different precisions (half, single, and double) on NVIDIA mixed-precision GPUs. We evaluate not only the error rate of the applications but also the effects of the errors on the final detection. We perform extensive fault-injection campaign on the register file of NVIDIA mixed-precision GPUs. We found that reducing data and operation precision increases the probability for the fault to impact the DNN detection and classification. Then, we complement the fault injection study with beam experiments. We exposed YOLOv3 running on Tesla V100s to neutron beams and found that the use of half precision reduces the error rate of up to 2x. The smaller exposed area and improved performances brought by reduced precision is then likely to increase the DNN reliability.
Fernando Santos 0001, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
ETS2
2019 Boosting HPC Applications in the Cloud Through JIT Traffic-Aware Path Provisioning
Guilherme R. Pretto, Bruno Lopes Dalmazo, Jonatas Adilson Marques, Zhongke Wu, Xingce Wang, Vladimir Korkhov, Philippe Olivier Alexandre Navaux, Luciano Paschoal Gaspary
ICCSA (4)7
2019 Minimizing Communication Overheads in Container-based Clouds for HPC Applications
abstract
Although the industry has embraced the cloud computing model, there are still significant challenges to be addressed concerning the quality of cloud services. Network-intensive applications may not scale in the cloud due to the sharing of the network infrastructure. In the literature, performance evaluation studies are showing that the network tends to limit the scalability and performance of HPC applications. Therefore, we proposed the aggregation of Network Interface Cards (NICs) in a ready-to-use integration with the OpenNebula cloud manager using Linux containers. We perform a set of experiments using a network microbenchmark to get specific network performance metrics and NAS parallel benchmarks to analyze the performance impact on HPC applications. Our results highlight that the implementation of NIC aggregation improves network performance in terms of throughput and latency. Moreover, HPC applications have different patterns of behavior when using our approach, which depends on communication and the amount of data transferring. While network-intensive applications increased the performance up to 38%, other applications with aggregated NICs maintained the same performance or presented slightly worse performance.
Anderson M. Maliszewski, Adriano Vogel, Dalvan Griebler, Eduardo Roloff, Luiz Gustavo Fernandes, Philippe Olivier Alexandre Navaux
ISCC6
2019 Multi-phased Task Placement of HPC Applications in the Cloud
abstract
Many high-performance computing applications present different phases during their execution. Nevertheless, thread and process placement techniques usually provide static-only methods to improve the data and thread locality. Similarly, cloud computing datacenters may present variations in terms of latency over the execution time of applications. To overcome these two problems, in this paper we analyze scientific applications that have different communication patterns along with its execution. For such applications, we evaluate the performance variation of traditional static placement techniques to our new approach that uses code annotations to perform the new placement of tasks, matching also the variations on network performance of Virtual Machines (VMs) during the run time. For our experiments, we use applications from the NAS parallel benchmark suite, running them on two VM sizes with 32 and 64 cores respectively, from the same family of instance types at the West US datacenter from Azure. Results show that compared to traditional static process mapping, our multi-phased placement mechanism achieves average performance gains of 13.57% up to 28.32% on the evaluated scenarios. These results show that there is an opportunity to improve performance by correctly identifying the network variations and reacting by generating a new task-to-instance mapping.
Emmanuell D. Carreño, Marco A. Z. Alves, Matthias Diener, Eduardo Roloff, Philippe Olivier Alexandre Navaux
ISPDC5
2019 Impact of Workload Distribution on Energy Consumption, Performance, and Reliability of Heterogeneous Devices
abstract
Devices integrating cores of different nature in the same chip achieve very high computation efficiency by reducing the power consumption and latencies of moving data from a chip to an external device. Heterogeneous devices are commonly used in high performance computing applications and, lately, their market has expanded from portable and gaming to safety-critical applications. In this paper, we evaluate how the collaborative workload distribution impact the energy consumption, performance, and reliability of heterogeneous devices. Then, we use the Energy-Delay-Fit Product (EDFP) to evaluate the trade off between the measured metrics and find how they correlate. To perform the proposed study we consider AMD Accelerated Processing Units (APUs) that embed a CPU and a GPU. We run four applications, each one representing an algorithm class, gradually distributing the workload from the CPU to the GPU and measuring both the energy consumption and execution time. Then, we take advantage of accelerated neutron beams to measure the realistic error rates of the different workload distributions. As we show in the paper, energy consumption and execution time are mold by the same trend while FIT rates highly depend on algorithm class and workload distribution. Additionally, we found that execution time is the most influencing factor for the device EDFP. The application EDFP varies of up tp 6 orders of magnitude depending on the workload distribution. An unwise configuration can then jeopardize the device efficiency and reliability.
Gabriel Piscoya Davila, Daniel Oliveira 0002, Philippe Olivier Alexandre Navaux, Paolo Rech
PDP3
2019 A Dynamic Task-Based D3Q19 Lattice-Boltzmann Method for Heterogeneous Architectures
abstract
Nowadays computing platforms expose a significant number of heterogeneous processing units such as multicore processors and accelerators. The task-based programming model has been a de facto standard model for such architectures since its model simplifies programming by unfolding parallelism at runtime based on data-flow dependencies between tasks. Many studies have proposed parallel strategies over heterogeneous platforms with accelerators. However, to the best of our knowledge, no dynamic task-based strategy of the Lattice-Boltzmann Method (LBM) has been proposed to exploit CPU+GPU computing nodes. In this paper, we present a dynamic task-based D3Q19 LBM implementation using three runtime systems for heterogeneous architectures: OmpSs, StarPU, and XKaapi. We detail our implementations and compare performance over two heterogeneous platforms. Experimental results demonstrate that our task-based approach attained up to 8.8 of speedup over an OpenMP parallel loop version.
João V. F. Lima, Gabriel Freytag, Vinícius Garcia Pinto, Claudio Schepke, Philippe Olivier Alexandre Navaux
PDP5
2019 Memory Performance and Bottlenecks in Multicore and GPU Architectures
abstract
Nowadays, there are several different architectures available not only for the industry, but also for normal consumers. Traditional multicore processors, GPUs, accelerators such as the Sunway SW26010, or even energy efficiency-driven processors such as the ARM family, present very different architectural characteristics. This wide range of characteristics presents a challenge for the developers of applications. Developers must deal with different instruction sets, memory hierarchies, or even different programming paradigms when programming for these architectures. Therefore, the same application can perform well when executing on one architecture, but poorly on another architecture. To optimize an application, it is important to have a deep understanding of how it behaves on different architectures. The related work in this area mostly focuses on a limited analysis encompassing execution time and energy. In this paper, we perform a detailed investigation on the impact of the memory subsystem of different architectures, which is one of the most important aspects to be considered. For this study, we performed experiments in the Broadwell CPU and Pascal GPU, using applications from the Rodinia benchmark suite. In this way, we were able to understand why an application performs well on one architecture and poorly on others.
Matheus S. Serpa, Francis B. Moreira 0001, Philippe Olivier Alexandre Navaux, Eduardo Henrique Molina da Cruz, Matthias Diener, Dalvan Griebler, Luiz Gustavo Fernandes
PDP3
2019 Detecting I/O Access Patterns of HPC Workloads at Runtime
abstract
In this paper, we seek to guide optimization and tuning strategies by identifying the application's I/O access pattern. We evaluate three machine learning techniques to automatically detect the I/O access pattern of HPC applications at runtime: decision trees, random forests, and neural networks. We focus on the detection using metrics from file-level accesses as seen by the clients, I/O nodes, and parallel file system servers. We evaluated these detection strategies in a case study in which the accurate detection of the current access pattern is fundamental to adjust a parameter of an I/O scheduling algorithm. We demonstrate that such approaches correctly classify the access pattern, regarding file layout and spatiality of accesses - into the most common ones used by the community and by I/O benchmarking tools to test new I/O optimization - with up to 99% precision. Furthermore, when applied to our study case, it guides a tuning mechanism to achieve 99% of the performance of an Oracle solution.
Jean Luca Bez, Francieli Zanon Boito, Ramon Nou, Alberto Miranda, Toni Cortes, Philippe Olivier Alexandre Navaux
SBAC-PAD6
2019 Non-uniform Partitioning for Collaborative Execution on Heterogeneous Architectures
abstract
Since the demand for computing power increases, new architectures arise to obtain better performance. An important class of integrated devices is heterogeneous architectures, which join different specialized hardware into a single chip, composing a System on Chip - SoC. Within this context, effectively splitting tasks between the different architectures is primal to obtain efficiency and performance. In this work, we evaluate two heterogeneous architectures: one composed of a general-purpose CPU and a graphics processing unit (GPU) integrated into a single chip (AMD Kaveri SoC), and another composed by a general-purpose CPU and a Field Programmable Gate Array (FPGA) integrated into a single chip (Intel Arria 10 SoC). We investigate how data partitioning affects the performance of each device in a collaborative execution through the decomposition of the data domain. As a case study, we apply the technique in the well-known Lattice Boltzmann Method (LBM), analyzing the performance of five kernels in both architectures. Our experimental results show that non-uniform partitioning improves LBM kernels performance by up to 11.40% and 15.15% on AMD Kaveri and Intel Arria 10, respectively.
Gabriel Freytag, Matheus S. Serpa, João V. F. Lima, Paolo Rech, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2019 Managing Power Demand and Load Imbalance to Save Energy on Systems with Heterogeneous CPU Speeds
abstract
Different simulations of real problems have been executed in High Performance Computing systems. However, the power consumption of these systems is an increasing concern once more energy are consumed to large simulations. In this context, load balancers emerge as a promising alternative for supporting the computational science methods. In response to this challenge, we developed a new heterogeneous energy-aware load balancer called H-ENERGYLB to reduce the average power demand of systems with heterogeneous processors and save energy when scientific applications with imbalanced load are executed. Our new load balancing strategy combines dynamic load balancing with DVFS techniques to mitigate the imbalanced workloads in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experiments with three applications on two different heterogeneous architectures show that H-ENERGYLB results in power reductions of 7.14% in average with the energy saving of 36.6% in average compared to others load balancers.
Edson L. Padoin, Matthias Diener, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
SBAC-PAD3
2019 An Unsupervised Learning Approach for I/O Behavior Characterization
abstract
I/O operations are the bottleneck of several applications due to the difference between processing and data access speeds. Hence, understanding the I/O behavior is vital to find problems and propose solutions. Thus, identifying and characterizing the I/O access pattern is important, since it reflects directly on applications' performance. With this premise, we propose an I/O characterization approach that uses unsupervised learning to cluster jobs with similar I/O behavior, using information from high-level aggregated traces. As a case study, we apply our approach on four months of activity - a total of 28, 938 jobs - from the Intrepid supercomputer located at Argonne Laboratory. Our experimental results show that nine access patterns represent the I/O behavior in 73% of the clusters. From these nine patterns, we learn some aspects about the I/O such as the most accesses patterns are made using POSIX and small requests, also, the most patterns are accessing unique files. Lastly, analyzing the I/O workload over four months, we can notice that it is composed by several applications that spend a short time on I/O activity, but when compared to the others, the total I/O time represents a greater portion of the overall system.
Pablo J. Pavan, Jean Luca Bez, Matheus S. Serpa, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2019 Energy efficiency and I/O performance of low-power architectures
abstract
Summary This paper presents an energy efficiency and I/O performance analysis of low‐power architectures when compared to conventional architectures, with the goal of studying the viability of using them as storage servers. Our results show that despite the fact the power demand of the storage device amounts for a small fraction of the power demand of the whole system, significant increases in power demand are observed when accessing the storage device. We investigate the access pattern impact on power demand, looking at the whole system and at the storage device by itself, and compare all tested configurations regarding energy efficiency. Then we extrapolate the conclusions from this research to provide guidelines for when considering the replacement of traditional storage servers by low‐power alternatives. We show the choice depends on the expected workload, estimates of power demand of the systems, and factors limiting performance. These guidelines can be applied for other architectures than the ones used in this work.
Pablo J. Pavan, Ricardo K. Lorenzoni, Vinícius Machado 0002, Jean Luca Bez, Edson L. Padoin, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.7
2019 Performance modeling of a geophysics application to accelerate over-decomposition parameter tuning through simulation
abstract
Summary Finite‐difference methods are commonplace in High Performance Computing applications. Despite their apparent regularity, they often exhibit load imbalance that damages their efficiency. We characterize the spatial and temporal load imbalance of Ondes3D, a typical finite‐differences application dedicated to earthquake modeling. Our analysis reveals imbalance originating from the structure of the input data, and from low‐level CPU optimizations. Ondes3D was successfully ported to AMPI/CHARM++ using over‐decomposition and MPI process migration techniques to dynamically rebalance the load. However, this approach requires careful selection of the over‐decomposition level, the load balancing algorithm, and its activation frequency. These choices are usually tied to application structure and platform characteristics. In this article, we propose a workflow that leverages the capabilities of SimGrid to conduct such study at low experimental cost. We rely on a combination of emulation, simulation, and application modeling that requires minimal code modification and manages to capture both spatial and temporal load imbalance to faithfully predict the performance of dynamic load balancing. We evaluate the quality of our simulation by comparing simulation results with the outcome of real executions and demonstrate how this approach can be used to quickly find the optimal load balancing configuration for a given application/hardware configuration.
Rafael Keller Tesser, Lucas Mello Schnorr, Arnaud Legrand, Franz C. Heinrich, Fabrice Dupros, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.6
2018 Exploiting Load Imbalance Patterns for Heterogeneous Cloud Computing Platforms
Eduardo Roloff, Matthias Diener, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
CLOSER4
2018 Collective I/O Performance on the Santos Dumont Supercomputer
abstract
The historical gap between processing and data access speeds causes many applications to spend a large portion of their execution on I/O operations. From the point of view of a large-scale, expensive, supercomputer, it is important to ensure applications achieve the best I/O performance to promote an efficient usage of the machine. In this paper, we evaluate the I/O infrastructure of the Santos Dumont supercomputer, the largest one from Latin America. More specifically, we investigate the performance of collective I/O operations. By conducting an analysis of a scientific application that uses the machine, we identify large performance differences between the available MPI implementations. We then further study the observed phenomenon using the BT-IO and IOR benchmarks, in addition to a custom microbenchmark. We conclude that the customized MPI implementation by Bull (used by more than 20% of the jobs) presents the worst performance for small collective write operations. Our results are being used to help the Santos Dumont users to achieve the best performance for their applications. Additionally, by investigating the observed phenomenon, we provide information to help improve future MPI-IO collective write implementations.
Andre Ramos Carneiro, Jean Luca Bez, Francieli Zanon Boito, Bruno Alves Fagundes, Carla Osthoff, Philippe Olivier Alexandre Navaux
PDP6
2018 Improving Communication and Load Balancing with Thread Mapping in Manycore Systems
abstract
Communication and load balancing have a significant impact on the performance of parallel applications and have been the subject of extensive research in multicore architectures. Thread mapping has been one of the solutions adopted in multicore architectures to address both communication and load balancing. However, the impact of such issues on more recently introduced manycore architectures is still unknown. Most related work on manycore architectures focus on execution time and idleness information for scheduling decisions. In this paper, we improve the state of the art by performing a very detailed analysis of the impact of thread mapping on communication and load balancing in two manycore systems from Intel, namely Knights Corner and Knights Landing. We observed that the widely used metric of CPU time provides very inaccurate information for load balancing. We also evaluated the usage of thread mapping based on the communication and load information of the applications to improve the performance of manycore systems.
Eduardo Henrique Molina da Cruz, Matthias Diener, Matheus S. Serpa, Philippe Olivier Alexandre Navaux, Laércio Lima Pilla, Israel Koren
PDP4
2018 Optimizing Machine Learning Algorithms on Multi-Core and Many-Core Architectures Using Thread and Data Mapping
abstract
Driven by the development of new technologies such as personal assistants or autonomous cars, machine learning has rapidly become one of the most active fields in computer science. The algorithms at the core of machine learning are notoriously demanding in terms of resources. It is therefore of paramount importance to optimize their operation on modern processors. Several approaches have been proposed to accelerate machine learning on GPUs and massively parallel computers, as well as dedicated ASICs. In this paper, we focus on Intel's multi-core Xeon and many-core accelerator Xeon Phi Knights Landing, which can host several hundreds of threads on the same CPU. In such architectures, thread and data mapping are keys for performance. We study the impact of mapping strategies, revealing that, with smart mapping policies, one can indeed significantly speed up machine learning applications on many-core architectures. Execution time was reduced by up to 25.2% and 18.5% on Intel Xeon and Xeon Phi KNL, respectively.
Matheus S. Serpa, Arthur M. Krause, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux, Marcelo Pasin, Pascal Felber
PDP4
2018 Predicting the Reliability Behavior of HPC Applications
abstract
The error rate of current High Performance Computing (HPC) systems is already in the order of one per dozens of hours. Understanding the reliability behavior of HPC applications will be required for the next generation of supercomputers. Using the reliability behavior one can select efficient mitigation techniques for the application and fine-tune parameters such as checkpoint frequency. In this paper, we investigate the application of a machine learning model to predict the reliability behavior of HPC applications. We inject faults in more than 30 HPC applications executing in the Intel Xeon Phi Knights Landing (KNL) and use profiling information to build a predictive model with Support Vector Machines (SVM). We show that the model can predict the Program Vulnerability Factor (PVF) with an average relative error of 7% for certain classes of algorithm, such as linear algebra and sorting. The average relative error for all algorithm classes is 22%. Such a fast and straightforward prediction model can be effective as a filter to select the most unreliable applications to perform an in-depth analysis.
Daniel Oliveira 0002, Francis B. Moreira 0001, Paolo Rech, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2018 MigPF: Towards on self-organizing process rescheduling of Bulk-Synchronous Parallel applications
Rodrigo da Rosa Righi, Roberto de Quadros Gomes, Vinicius Facco Rodrigues, Cristiano André da Costa, Antônio Marcos Alberti, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
Future Gener. Comput. Syst.7
2018 A lightweight plug-and-play elasticity service for self-organizing resource provisioning on parallel applications
Rodrigo da Rosa Righi, Vinicius Facco Rodrigues, Gustavo Rostirolla, Cristiano André da Costa, Eduardo Roloff, Philippe Olivier Alexandre Navaux
Future Gener. Comput. Syst.6
2017 Leveraging Cloud Heterogeneity for Cost-Efficient Execution of Parallel Applications
Eduardo Roloff, Matthias Diener, Emmanuell D. Carreño, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
Euro-Par5
2017 Using Simulation to Evaluate and Tune the Performance of Dynamic Load Balancing of an Over-Decomposed Geophysics Application
Rafael Keller Tesser, Lucas Mello Schnorr, Arnaud Legrand, Fabrice Dupros, Philippe Olivier Alexandre Navaux
Euro-Par5
2017 Radiation-Induced Error Criticality in Modern HPC Parallel Accelerators
abstract
In this paper, we evaluate the error criticality of radiation-induced errors on modern High-Performance Computing (HPC) accelerators (Intel Xeon Phi and NVIDIA K40) through a dedicated set of metrics. We show that, as long as imprecise computing is concerned, the simple mismatch detection is not sufficient to evaluate and compare the radiation sensitivity of HPC devices and algorithms. Our analysis quantifies and qualifies radiation effects on applications' output correlating the number of corrupted elements with their spatial locality. Also, we provide the mean relative error (dataset-wise) to evaluate radiation-induced error magnitude. We apply the selected metrics to experimental results obtained in various radiation test campaigns for a total of more than 400 hours of beam time per device. The amount of data we gathered allows us to evaluate the error criticality of a representative set of algorithms from HPC suites. Additionally, based on the characteristics of the tested algorithms, we draw generic reliability conclusions for broader classes of codes. We show that arithmetic operations are less critical for the K40, while Xeon Phi is more reliable when executing particles interactions solved through Finite Difference Methods. Finally, iterative stencil operations seem the most reliable on both architectures.
Daniel Oliveira 0002, Laércio Lima Pilla, Mauricio Hanzich, Vinicius Fratin, Fernando Santos 0001, Caio B. Lunardi, José María Cela, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
HPCA8
2017 TWINS: Server Access Coordination in the I/O Forwarding Layer
abstract
This paper presents a study of I/O scheduling techniques applied to the I/O forwarding layer. In high-performance computing environments, applications rely on parallel file systems (PFS) to obtain good I/O performance even when handling large amounts of data. To alleviate the concurrency caused by thousands of nodes accessing a significantly smaller number of PFS servers, intermediate I/O nodes are typically applied between processing nodes and the file system. Each intermediate node forwards requests from multiple clients to the system, a setup which gives this component the opportunity to perform optimizations like I/O scheduling. We evaluate scheduling techniques that improve spatiality and request size of the access patterns. We show they are only partially effective because the access pattern is not the main factor for read performance in the I/O forwarding layer. A new scheduling algorithm, TWINS, is presented to coordinate the access of intermediate I/O nodes to the data servers. Our proposal decreases concurrency at the data servers, a factor previously proven to negatively affect performance. The proposed algorithm is able to improve read performance from shared files by up to 28% over other scheduling algorithms and by up to 50% over not forwarding I/O.
Jean Luca Bez, Francieli Zanon Boito, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
PDP4
2017 High Performance I/O for Seismic Wave Propagation Simulations
abstract
This paper describes our research to provide high performance I/O for seismic wave propagation simulations. Earthquake early warning systems are designed to provide near real-time prediction of strong ground motion. Such systems are crucial tools for risk mitigation and disaster prevention. The ability to accurately and quickly simulate the propagation of seismic waves in complex media lies at the heart of such systems. Besides the processing requirements, it is important for seismic simulations to leverage a high-performance storage infrastructure to output results as frequently as possible, so they can be used for the decision-making process. We propose and evaluate a series of I/O optimizations to the Ondes3D seismic wave propagation simulation, considering its different types of output files separately. These optimizations are designed while keeping the previous output formats, in order not to compromise the application interaction with the other parts of the earthquake early warning system. The optimization techniques presented in this paper have provided I/O performance improvements of up to 85% and decreased the application execution time up to 70%.
Francieli Zanon Boito, Jean Luca Bez, Fabrice Dupros, Mario A. R. Dantas, Philippe Olivier Alexandre Navaux, Hideo Aochi
PDP5
2017 HPC Application Performance and Cost Efficiency in the Cloud
abstract
Unlike traditional cluster systems, the Cloud Computing paradigm provides access to an execution environment without upfront investments in hardware and facilities. Due to the elasticity and the pay-per-use billing model, it is possible to configure experimental environments with minimal idle costs. In this paper, we perform an extensive evaluation of the major commercial public clouds. Our results show that performance degradation due to virtualization and other cloud overheads is insignificant. However, the network interconnection in the cloud still remains a large bottleneck for HPC application performance.
Eduardo Roloff, Matthias Diener, Luciano Paschoal Gaspary, Philippe Olivier Alexandre Navaux
PDP4
2017 Experimental and analytical study of Xeon Phi reliability
abstract
We present an in-depth analysis of transient faults effects on HPC applications in Intel Xeon Phi processors based on radiation experiments and high-level fault injection. Besides measuring the realistic error rates of Xeon Phi, we quantify Silent Data Corruption (SDCs) by correlating the distribution of corrupted elements in the output to the application's characteristics. We evaluate the benefits of imprecise computing for reducing the programs' error rate. For example, for HotSpot a 0.5% tolerance in the output value reduces the error rate by 85%.
Daniel Oliveira 0002, Laércio Lima Pilla, Nathan DeBardeleben, Sean Blanchard, Heather M. Quinn, Israel Koren, Philippe Olivier Alexandre Navaux, Paolo Rech
SC7
2017 Performance and energy efficiency analysis of HPC physics simulation applications in a cluster of ARM processors
abstract
Summary We analyze the feasibility and energy efficiency of using an unconventional cluster of low‐power Advanced RISC Machines processors to execute two scientific parallel applications. For this purpose, we have selected two applications that present high computational and communication cost: the Ondes3D that simulates geophysical events, and the all‐pairs N‐Body that simulates astrophysical events. We compare and discuss the impact of different compilation directives and processor frequency and how they interfere in Time‐to‐Solution and Energy‐to‐Solution. Our results demonstrate that by correctly tuning the application at compile time, for the Advanced RISC Machines architecture, we can considerably reduce the execution time and the energy spent by computing simulations. Furthermore, we observe reductions of up to 54.14% in Time‐to‐Solution and gains of up to 53.65% in Energy‐to‐Solution with two cores. Additionally, we consider the impact of two processor frequency governors on these metrics. Results indicate that the powersave governor presents a smaller instantaneous power consumption. However, it spends more time executing tasks, increasing the energy needed to achieve the solution. Finally, we correlate the energy consumption with the execution time in the experimental results using Pareto. These findings suggest that it is possible to explore low‐powered clusters for high‐performance computing applications by tuning application and hardware configuration to achieve energy efficiency. Copyright © 2016 John Wiley & Sons, Ltd.
Jean Luca Bez, Eliezer E. Bernart, Fernando Santos 0001, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.5
2017 CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processors
abstract
Summary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd.
Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Concurr. Comput. Pract. Exp.8
2016 Automatic Communication Optimization of Parallel Applications in Public Clouds
abstract
One of the most important aspects that influences the performance of parallel applications is the speed of communication between their tasks. To optimize communication, tasks that exchange lots of data should be mapped to processing units that have a high network performance. This technique is called communication-aware task mapping and requires detailed information about the underlying network topology for an accurate mapping. Previous work on task mapping focuses on network clusters or shared memory architectures, in which the topology can be determined directly from the hardware environment. Cloud computing adds significant challenges to task mapping, since information about network topologies is not available to end users. Furthermore, the communication performance might change due to external factors, such as different usage patterns of other users. In this paper, we present a novel solution to perform communication-aware task mapping in the context of commercial cloud environments with multiple instances. Our proposal consists of a short profiling phase to discover the network topology and speed between cloud instances. The profiling can be executed before each application start as it causes only a negligible overhead. This information is then used together with the communication pattern of the parallel application to group tasks based on the amount of communication and to map groups with a lot of communication between them to cloud instances with a high network performance. In this way, application performance is increased, and data traffic between instances is reduced. We evaluated our proposal in a public cloud with a variety of MPI-based parallel benchmarks from the HPC domain, as well as a large scientific application. In the experiments, we observed substantial performance improvements (up to 11 times faster) compared to the default scheduling policies.
Emmanuell D. Carreño, Matthias Diener, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux
CCGrid4
2016 Fostering Collaboration in Energy Research and Technological Developments Applying New Exascale HPC Techniques
abstract
During the last years, High Performance Computing (HPC) resources have undergone a dramatic transformation, with an explosion on the available parallelism and the use of special purpose processors. There are international initiatives focusing on redesigning hardware and software in order to achieve the Exaflop capability. With this aim, the HPC4E project is applying the new exascale HPC techniques to energy industry simulations, customizing them if necessary, and going beyond the state-of-the-art in the required HPC exascale simulations for different energy sources that are the present and the future of energy: wind energy production and design, efficient combustion systems for biomass-derived fuels (biogas), and exploration geophysics for hydrocarbon reservoirs. HPC4E joins efforts of several institutions settled in Brazil and Europe.
José María Cela, Philippe Olivier Alexandre Navaux, Alvaro L. G. A. Coutinho, Rafael Mayo 0001
CCGrid2
2016 enerGyPU and enerGyPhi Monitor for Power Consumption and Performance Evaluation on Nvidia Tesla GPU and Intel Xeon Phi
abstract
The evaluation of performance and power consumption is a key step in the design of applications for large computational systems as supercomputers and clusters (multicore and accelerator nodes, multicore and coprocessor nodes, manycore and accelerator nodes). In these systems the developers must design several experiments for workload characterization observing the architectural implications when using different combinations of computational resources such as number of GPU, number of cores for processing, number of cores for administration of GPU, number of MPI processes and thread affinity policy. It should also engage factors as the clock frequency and memory usage as well select the combination of computational resources that increases the performance and minimizes the power consumption. This research proposes an integrated energy-aware scheme called efficiently energetic acceleration (EEA) for large-scale scientific applications running on heterogeneous architectures. This paper shows the use of a monitoring tool with two components called enerGyPU and enerGyPhi to recording EEA control factors in runtime on two environments: one cluster with multicore and accelerator nodes (2-CPU/8-GPU) and one server with multiple cores and one coprocessor (2-CPU/1-MIC). These monitors allow to analyze multiple testing results under different parameter combinations to observe the EEA control factors that determine the energy efficiency.
John Anderson García Henao, Esteban Hernandez B., Carlos Enrique Montenegro-Marín, Philippe Olivier Alexandre Navaux, Carlos Jaime Barrios Hernández
CCGrid4
2016 A Sharing-Aware Memory Management Unit for Online Mapping in Multi-core Architectures
Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
Euro-Par4
2016 Towards Weather Forecasting in the Cloud
abstract
Scientific applications process a large amount of data and need huge computing power. Traditionally, they are executed in supercomputers, cluster or grid environments. Recently, the cloud also emerged as a feasible execution environment for this kind of application. The viability of using cloud was already extensively validated. In this work, we migrated a weather forecast application to the Microsoft Azure cloud. The intention of our work is to perform an application modernization using the cloud technology. Several pre-and post-processing procedures were modified to decrease or eliminate repetitive tasks. Additionally, the configuration of the application is no longer necessary in the user perspective. We can conclude that with some modification, scientific applications can benefit from cloud technologies, and it is possible to create modern services based on legacy applications.
Emmanuell D. Carreño, Eduardo Roloff, Philippe Olivier Alexandre Navaux
PDP3
2016 Communication in Shared Memory: Concepts, Definitions, and Efficient Detection
abstract
Optimizing the communication behavior of parallel applications has emerged as an important topic in parallel processing. In shared memory architectures, threads communicate implicitly through memory accesses to shared memory areas. The communication behavior can be improved by mapping threads that communicate a lot to processing units that are close to each other in the memory hierarchy, such that they can benefit from shared caches and faster interconnections. An important aspect of such a communication-aware thread mapping is the accurate and efficient detection of communication in shared memory. Previous work used impromptu definitions, without an evaluation of the complexities of different communication types. In this paper, we perform an in-depth, systematic evaluation of communication in shared memory, focusing on its architectural effects. We present an efficient way to detect communication, which is orders of magnitude faster than a cache simulator, while maintaining a high accuracy.
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
PDP4
2016 Analyzing and Improving Memory Access Patterns of Large Irregular Applications on NUMA Machines
abstract
Improving the memory access behavior of parallel applications is one of the most important challenges in high-performance computing. Non-Uniform Memory Access (NUMA) architectures pose particular challenges in this context: they contain multiple memory controllers and the selection of a controller to serve a page request influences the overall locality and balance of memory accesses, which in turn affect performance. In this paper, we analyze and improve the memory access pattern and overall memory usage of large-scale irregular applications on NUMA machines. We selected HashSieve, a very important algorithm in the context of lattice-based cryptography, as a representative example, due to (1) its extremely irregular memory pattern, (2) large memory requirements and (3) unsuitability to other computer architectures, such as GPUs. We optimize HashSieve with a variety of techniques, focusing both on the algorithm itself as well as the mapping of memory pages to NUMA nodes, achieving a speedup of over 2x.
Artur Mariano, Matthias Diener, Christian H. Bischof, Philippe Olivier Alexandre Navaux
PDP4
2016 Exploring Cache Size and Core Count Tradeoffs in Systems with Reduced Memory Access Latency
abstract
One of the main challenges for computer architects is how to hide the high average memory access latency from the processor. In this context, Hybrid Memory Cubes (HMCs) can provide substantial energy and bandwidth improvements compared to traditional memory organizations. However, it is not clear how this reduced average memory access latency will impact the LLC. For applications with high cache miss ratios, the latency to search for the data inside the cache memory will impact negatively on the performance. The importance of this overhead depends on the memory access latency. In this paper, we present an evaluation of the L3 cache importance on a high performance processor using HMC also exploring chip area tradeoffs between the cache size and number of processor cores. We show that the high bandwidth provided by HMC memories can eliminate the need for L3 caches, removing hardware and making room for more processing power. Our evaluations show that performance increased 37% and the EDP improved 12% while maintaining the same original chip area in a wide range of parallel applications, when compared to DDR3 memories.
Paulo C. Santos 0001, Marco A. Z. Alves, Matthias Diener, Luigi Carro, Philippe Olivier Alexandre Navaux
PDP5
2016 Automatic I/O scheduling algorithm selection for parallel file systems
abstract
Summary This article presents our approach to provide input/output (I/O) scheduling with double adaptivity: to applications and devices. In high‐performance computing environments, parallel file systems provide a shared storage infrastructure to applications. In the situation where multiple applications access this shared infrastructure concurrently, their performance can be impaired because of interference. Our work focuses on I/O scheduling as a tool to improve performance by alleviating interference effects. The role of the I/O scheduler is to decide the order in which applications' requests must be processed by the parallel file system's servers, applying optimizations to adjust the resulting access pattern for improved performance. Our approach to improve I/O scheduling results is based on using information from applications' access patterns and storage devices' sensitivity to access sequentiality. We have applied machine learning to provide the ability to automatically select the best scheduling algorithm for each situation. Our approach improves performance by up to 75% over an approach that uses the same scheduling algorithm to all situations, without adaptability. Our results evidence that both aspects – applications and storage devices – are essential to make good scheduling decisions. Copyright © 2015 John Wiley & Sons, Ltd.
Francieli Zanon Boito, Rodrigo Kassick, Philippe Olivier Alexandre Navaux, Yves Denneulin
Concurr. Comput. Pract. Exp.3
2016 Seismic wave propagation simulations on low-power and performance-centric manycores
Márcio Castro 0001, Emilio Francesquini, Fabrice Dupros, Hideo Aochi, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Parallel Comput.5
2016 LAPT: A locality-aware page table for thread and data mapping
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
Parallel Comput.5
2016 A dynamic block-level execution profiler
Francis B. Moreira 0001, Marco A. Z. Alves, Matthias Diener, Philippe Olivier Alexandre Navaux, Israel Koren
Parallel Comput.4
2016 Hardware-Assisted Thread and Data Mapping in Hierarchical Multicore Architectures
abstract
The performance and energy efficiency of modern architectures depend on memory locality, which can be improved by thread and data mappings considering the memory access behavior of parallel applications. In this article, we propose intense pages mapping, a mechanism that analyzes the memory access behavior using information about the time the entry of each page resides in the translation lookaside buffer. It provides accurate information with a very low overhead. We present experimental results with simulation and real machines, with average performance improvements of 13.7% and energy savings of 4.4%, which come from reductions in cache misses and interconnection traffic.
Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
ACM Trans. Archit. Code Optim.4
2016 Evaluation of Histogram of Oriented Gradients Soft Errors Criticality for Automotive Applications
abstract
Pedestrian detection reliability is a key problem for autonomous or aided driving, and methods that use Histogram of Oriented Gradients (HOG) are very popular. Embedded Graphics Processing Units (GPUs) are exploited to run HOG in a very efficient manner. Unfortunately, GPUs architecture has been shown to be particularly vulnerable to radiation-induced failures. This article presents an experimental evaluation and analytical study of HOG reliability. We aim at quantifying and qualifying the radiation-induced errors on pedestrian detection applications executed in embedded GPUs. We analyze experimental results obtained executing HOG on embedded GPUs from two different vendors, exposed for about 100 hours to a controlled neutron beam at Los Alamos National Laboratory. We consider the number and position of detected objects as well as precision and recall to discriminate critical erroneous computations. The reported analysis shows that, while being intrinsically resilient (65% to 85% of output errors only slightly impact detection), HOG experienced some particularly critical errors that could result in undetected pedestrians or unnecessary vehicle stops. Additionally, we perform a fault-injection campaign to identify HOG critical procedures. We observe that Resize and Normalize are the most sensitive and critical phases, as about 20% of injections generate an output error that significantly impacts HOG detection. With our insights, we are able to find those limited portions of HOG that, if hardened, are more likely to increase reliability without introducing unnecessary overhead.
Fernando Santos 0001, Lucas Weigel, Cláudio R. Jung, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech
ACM Trans. Archit. Code Optim.4
2016 Kernel-Based Thread and Data Mapping for Improved Memory Affinity
abstract
Reducing the cost of memory accesses, both in terms of performance and energy consumption, is a major challenge in shared-memory architectures. Modern systems have deep and complex memory hierarchies with multiple cache levels and memory controllers, leading to a Non-Uniform Memory Access (NUMA) behavior. In such systems, there are two ways to improve the memory affinity: First, by mapping threads that share data to cores with a shared cache, cache usage and communication performance are optimized. Second, by mapping memory pages to memory controllers that perform the most accesses to them and are not overloaded, the average cost of accesses is reduced. We call these two techniques thread mapping and data mapping, respectively. Thread and data mapping should be performed in an integrated way to achieve a compounding effect that results in higher improvements overall. Previous work in this area requires expensive tracing operations to perform the mapping, or require changes to the hardware or to the parallel application. In this paper, we propose kMAF, a mechanism that performs integrated thread and data mapping in the kernel. kMAF uses the page faults of parallel applications to characterize their memory access behavior and performs the mapping during the execution of the application based on the detected behavior. In an evaluation with a large set of parallel benchmarks executing on three NUMA architectures, kMAF achieved substantial performance and energy efficiency improvements, close to an Oracle-based mechanism and significantly higher than previous proposals.
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux, Anselm Busse, Hans-Ulrich Heiß
IEEE Trans. Parallel Distributed Syst.4
2015 Locality and Balance for Communication-Aware Thread Mapping in Multicore Systems
Matthias Diener, Eduardo Henrique Molina da Cruz, Marco A. Z. Alves, Mohammad Shadi Al Hakeem, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
Euro-Par5
2015 Understanding GPU errors on large-scale HPC systems and the implications for system design and operation
abstract
Increase in graphics hardware performance and improvements in programmability has enabled GPUs to evolve from a graphics-specific accelerator to a general-purpose computing device. Titan, the world's second fastest supercomputer for open science in 2014, consists of more dum 18,000 GPUs that scientists from various domains such as astrophysics, fusion, climate, and combustion use routinely to run large-scale simulations. Unfortunately, while the performance efficiency of GPUs is well understood, their resilience characteristics in a large-scale computing system have not been fully evaluated. We present a detailed study to provide a thorough understanding of GPU errors on a large-scale GPU-enabled system. Our data was collected from the Titan supercomputer at the Oak Ridge Leadership Computing Facility and a GPU cluster at the Los Alamos National Laboratory. We also present results from our extensive neutron-beam tests, conducted at Los Alamos Neutron Science Center (LANSCE) and at ISIS (Rutherford Appleron Laboratories, UK), to measure the resilience of different generations of GPUs. We present several findings from our field data and neutron-beam experiments, and discuss the implications of our results for future GPU architects, current and future HPC computing facilities, and researchers focusing on GPU resilience.
Devesh Tiwari, Saurabh Gupta 0002, James H. Rogers, Don E. Maxwell, Paolo Rech, Sudharshan S. Vazhkudai, Daniel Oliveira 0002, Dave Londo, Nathan DeBardeleben, Philippe Olivier Alexandre Navaux, Luigi Carro, Arthur S. Bland
HPCA10
2015 An Efficient Algorithm for Communication-Based Task Mapping
abstract
The communication between tasks of a parallel application is an important characteristic to consider when mapping tasks to computing cores due to possible differences in communication performance. Within a machine, performance differences are introduced by the memory hierarchy, in which cache memories can be shared by groups of cores and intra-chip interconnections are faster than inter-chip interconnections. In cluster and grid systems, the network imposes an additional communication latency. By mapping tasks that communicate to cores nearby on the memory hierarchy, or to the same nodes in clusters or grids, the communication of parallel applications is optimized, leading to increased performance and energy efficiency. In the task mapping context, one of the most important aspects to be considered is the mapping algorithm, as it determines the improvements that can be achieved. Since the problem of finding the best mapping is NP-Hard, heuristics must be employed to find an approximate solution in feasible time. In this paper, we present Eager Map, a new algorithm to perform communication-based mapping that is based on a greedy grouping strategy applied hierarchically. Experimental evaluation indicates that the execution time of our algorithm is 10 times faster than the state-of-the-art, and presents higher performance improvements. Due to its low execution time and high stability, Eager Map is also suitable for online task mapping, where tasks are migrated during execution.
Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
PDP4
2015 Locality vs. Balance: Exploring Data Mapping Policies on NUMA Systems
abstract
In parallel architectures that have a Non-Uniform Memory Access (NUMA) behavior, the mapping of memory pages to NUMA nodes influences the performance of parallel applications. In order to improve traditional data mapping policies, two basic strategies can be employed: optimizing locality or balance of memory accesses. In a locality-based policy, memory pages are mapped to nodes that access the page the most. In a balance-based policy, memory pages are mapped such that the number of memory accesses resolved by each memory controller is similar. In this paper, we perform an in-depth exploration of these data mapping policies on the performance of parallel applications. We introduce metrics that describe their memory access behavior and evaluate their suitability for data mapping. We also present new mapping policies that focus on locality, balance or both. These policies were evaluated on three different NUMA architectures with applications from the NAS-OMP and PARSEC benchmark suites. Results show that the performance improvements of each policy depend on the characteristics of the applications and machines. Choosing the wrong policy can actually hurt the performance compared to the default first-touch mapping. Compared to traditional mapping policies and to policies that only focus on either locality or balance, taking into account both locality and balance results in the highest improvements. Furthermore, it avoids the performance reduction caused by the wrong data mapping.
Matthias Diener, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux
PDP3
2015 Towards Seismic Wave Modeling on Heterogeneous Many-Core Architectures Using Task-Based Runtime System
abstract
Understanding three-dimensional seismic wave propagation in complex media is still one of the main challenges of quantitative seismology. Because of its simplicity and numerical efficiency, the finite-differences method is one of the standard techniques implemented to consider the elastodynamics equation. Additionally, this class of modeling heavily relies on parallel architectures in order to tackle large scale geometries including a detailed description of the physics. Last decade, significant efforts have been devoted towards efficient implementation of the finite-differences methods on emerging architectures. These contributions have demonstrated their efficiency leading to robust industrial applications. The growing representation of heterogeneous architectures combining general purpose multicore platforms and accelerators leads to re-design current parallel application. In this paper, we consider Star PU task-based runtime system in order to harness the power of heterogeneous CPU+GPU computing nodes. We detail our implementation and compare the performance obtained with the classical CPU or GPU only versions. Preliminary results demonstrate significant speedups in comparison with the best implementation suitable for homogeneous cores.
Víctor Martínez, David Michéa, Fabrice Dupros, Olivier Aumage, Samuel Thibault, Hideo Aochi, Philippe Olivier Alexandre Navaux
SBAC-PAD7
2015 Communication-aware thread mapping using the translation lookaside buffer
abstract
Summary Threads of parallel applications need to communicate in order to fulfill their tasks. The communication performance between the cores in modern multi‐core architectures differs because of the memory and interconnection hierarchies. In these architectures, it is important to map the threads of parallel applications by taking into account the communication between them, to improve their performance and energy consumption. In parallel applications based on shared memory, communication is implicit, which makes it difficult to detect the communication pattern between the threads. In this paper, we introduce a new lightweight mechanism to detect the communication pattern between threads of shared memory applications using the translation lookaside buffer. Our mechanism relies on hardware features, which make it transparent to the programmer and allow the detection to be performed by the operating system during the execution of the application. We also developed a heuristic mapping algorithm that uses the detected pattern to dynamically map the threads to cores. Experiments were performed with applications from the NAS‐OMP and PARSEC parallel benchmark suites in a simulated machine as well as a real machine. Results show that our mechanism can substantially improve parallel application performance, as well as processor and DRAM energy consumption. Copyright © 2015 John Wiley & Sons, Ltd.
Eduardo Henrique Molina da Cruz, Matthias Diener, Philippe Olivier Alexandre Navaux
Concurr. Comput. Pract. Exp.3
2015 On the energy efficiency and performance of irregular application executions on multicore, NUMA and manycore platforms
Emilio Francesquini, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna, Fabrice Dupros, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
J. Parallel Distributed Comput.6
2015 Communication-aware process and thread mapping using online communication detection
Matthias Diener, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux, Anselm Busse, Hans-Ulrich Heiß
Parallel Comput.3
2015 Characterizing communication and page usage of parallel applications for thread and data mapping
Matthias Diener, Eduardo Henrique Molina da Cruz, Laércio Lima Pilla, Fabrice Dupros, Philippe Olivier Alexandre Navaux
Perform. Evaluation5
2014 kMAF: automatic kernel-level management of thread and data affinity
abstract
One of the main challenges for parallel architectures is the increasing complexity of the memory hierarchy, which consists of several levels of private and shared caches, as well as interconnections between separate memories in NUMA machines. To make full use of this hierarchy, it is necessary to improve the locality of memory accesses by reducing accesses to remote caches and memories, and using local ones instead. Two techniques can be used to increase the memory access locality: executing threads and processes that access shared data close to each other in the memory hierarchy (thread affinity), and placing the memory pages they access on the NUMA node they are executing on (data affinity). Most related work in this area focuses on either thread or data affinity, but not both, which limits the improvements. Other mechanisms require expensive operations, such as memory access traces or binary analysis, require changes to hardware or work only on specific parallel APIs.
Matthias Diener, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux, Anselm Busse, Hans-Ulrich Heiß
PACT3
2014 Radiation Sensitivity of High Performance Computing Applications on Kepler-Based GPGPUs
abstract
In this paper we assess and discuss the radiation sensitivity of a set of HPC applications executed on NVIDIA K20 GPGPUs. The occurrence of both radiation-induced silent data corruption and functional interruption will be experimentally addressed for Hotspot, LavaMD, and Matrix Transponse. Each of the tested codes requires a proper computational power and elaborates a different amount of data. Both these characteristics play a significant role in the application radiations sensitivity. Additionally, an evaluation of the error rate at sea level will be provided for all the tested codes.
Daniel Oliveira 0002, Caio B. Lunardi, Laércio Lima Pilla, Paolo Rech, Philippe Olivier Alexandre Navaux, Luigi Carro
DSN5
2014 Impact of GPUs Parallelism Management on Safety-Critical and HPC Applications Reliability
abstract
Graphics Processing Units (GPUs) offer high computational power but require high scheduling strain to manage parallel processes, which increases the GPU cross section. The results of extensive neutron radiation experiments performed on NVIDIA GPUs confirm this hypothesis. Reducing the application Degree Of Parallelism (DOP) reduces the scheduling strain but also modifies the GPU parallelism management, including memory latency, thread registers number, and the processors occupancy, which influence the sensitivity of the parallel application. An analysis on the overall GPU radiation sensitivity dependence on the code DOP is provided and the most reliable configuration is experimentally detected. Finally, modifying the parallel management affects the GPU cross section but also the code execution time and, thus, the exposure to radiation required to complete computation. The Mean Workload and Executions Between Failures metrics are introduced to evaluate the workload or the number of executions computed correctly by the GPU on a realistic application.
Paolo Rech, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Luigi Carro
DSN3
2014 Saving energy by exploiting residual imbalances on iterative applications
abstract
The power consumption of High Performance Computing (HPC) systems is an increasing concern as large-scale systems grow in size and, consequently, consume more energy. In response to this challenge, we propose two variants of a new energy-aware load balancer that aim at reducing the energy consumption of parallel platforms running imbalanced scientific applications without degrading their performance. Our research combines dynamic load balancing with DVFS techniques in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experimental results with benchmarks and a real-world application presented energy savings of up to 32% with our fine-grained variant that performs per-core DVFS, and of up to 34% with our coarsegrained variant that performs per-chip DVFS.
Edson L. Padoin, Márcio Castro 0001, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
HiPC4
2014 Improving the Performance of Seismic Wave Simulations with Dynamic Load Balancing
abstract
Seismic wave models provide a way to study the consequences of future earthquakes. When modeling a restricted region, these models require a boundary condition to absorb the energy that goes out of the simulated domain. To parallelize these models, the domain is decomposed into a grid of smaller subdomains which are mapped to different tasks. Due to the boundary condition, this division gives rise to load imbalance between the tasks that simulate border regions and those assigned center subdomains. To deal with this imbalance, and therefore improve the simulation's performance, we propose the use of dynamic load balancing. To evaluate our solution, we ported a seismic wave simulator to Adaptive MPI to profit from its load balancing framework. By using dynamic load balancers, we improved the performance of the application by 23.85% when compared to the original MPI implementation. We also show that load balancers are able to adapt to the variation of load imbalance during the application's execution.
Rafael Keller Tesser, Laércio Lima Pilla, Fabrice Dupros, Philippe Olivier Alexandre Navaux, Jean-François Méhaut, Celso L. Mendes
PDP4
2014 Energy Efficient Seismic Wave Propagation Simulation on a Low-Power Manycore Processor
abstract
Large-scale simulation of seismic wave propagation is an active research topic. Its high demand for processing power makes it a good match for High Performance Computing (HPC). Although we have observed a steady increase on the processing capabilities of HPC platforms, their energy efficiency is still lacking behind. In this paper, we analyze the use of a low-power manycore processor, the MPPA-256, for seismic wave propagation simulations. First we look at its peculiar characteristics such as limited amount of on-chip memory and describe the intricate solution we brought forth to deal with this processor's idiosyncrasies. Next, we compare the performance and energy efficiency of seismic wave propagation on MPPA-256 to other commonplace platforms such as general-purpose processors and a GPU. Finally, we wrap up with the conclusion that, even if MPPA-256 presents an increased software development complexity, it can indeed be used as an energy efficient alternative to current HPC platforms, resulting in up to 71% and 81% less energy than a GPU and a general-purpose processor, respectively.
Márcio Castro 0001, Fabrice Dupros, Emilio Francesquini, Jean-François Méhaut, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2014 Optimizing Memory Locality Using a Locality-Aware Page Table
abstract
One of the main challenges for modern parallel shared-memory architectures are accesses to main memory. In current systems, the performance and energy efficiency of memory accesses depend on their locality: accesses to remote caches and NUMA nodes are more expensive than accesses to local ones. Increasing the locality requires knowledge about how the threads of a parallel application access memory pages. With this information, pages can be migrated to the NUMA nodes that access them (data mapping), as well as threads that access the same pages can be migrated to the same node such that locality can be improved even further (thread mapping). In this paper, we propose LAPT, a mechanism to store the memory access pattern of parallel applications in the page table, which is updated by the hardware during TLB misses. This information is used by the operating system to perform an optimized thread and data mapping during the execution of the parallel application. In contrast to previous work, LAPT does not require any previous information about the behavior of the applications, or changes to the application or runtime libraries. Extensive experiments with the NAS Parallel Benchmarks (NPB) and PARSEC showed performance and energy efficiency improvements of up to 19.2% and 15.7%, respectively, (6.7% and 5.3% on average).
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2014 A topology-aware load balancing algorithm for clustered hierarchical multi-core machines
Laércio Lima Pilla, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Philippe Olivier Alexandre Navaux, Jean-François Méhaut
Future Gener. Comput. Syst.6
2014 Dynamic thread mapping of shared memory applications by exploiting cache coherence protocols
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
J. Parallel Distributed Comput.4
2014 Best of SBAC-PAD 2012
Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
Parallel Comput.2
2013 AGIOS: Application-Guided I/O Scheduling for Parallel File Systems
abstract
In this paper, we improve the performance of server-side I/O scheduling on parallel file systems by transparently including information about the applications' access patterns. Server-side I/O scheduling is a valuable tool on multiapplication scenarios, where the applications' spatial locality suffers from interference caused by concurrent accesses to the file system. We present AGIOS, an I/O scheduling library for parallel file systems. We guide scheduler's decisions by including information about the applications' future requests. This information is obtained from traces generated by the scheduler itself, without changes in application or file system. Our approach shows performance improvements under different workloads of 46.3% on average when compared to a scenario without an I/O scheduler, and of 25.1% when compared to a scheduler which does not use information about future accesses.
Francieli Zanon Boito, Rodrigo Kassick, Philippe Olivier Alexandre Navaux, Yves Denneulin
ICPADS3
2013 Communication-Based Mapping Using Shared Pages
abstract
In current shared memory architectures, the complexity of the cache and memory hierarchies is increasing. Therefore, it is becoming more important to analyze the communication behavior of parallel applications when mapping threads to cores, to improve performance and energy efficiency. However, communication is implicit in most programming models for shared memory, which makes it difficult to detect the communication pattern between the threads in an accurate and low-overhead way. We propose a new mechanism to detect the communication pattern of shared memory applications by monitoring page table accesses. Combining this mechanism with a dynamic migration algorithm allows mapping to be performed dynamically by the operating system. We implemented our mechanism in the Linux kernel and performed experiments with applications from the NAS Parallel Benchmarks. Results show a reduction of up to 16.7% of the execution time and 63% of the cache misses, compared to the original scheduler of the operating system. Furthermore, we decrease total processor and DRAM energy consumption by up to 14.7% and 28.5%, respectively.
Matthias Diener, Eduardo Henrique Molina da Cruz, Philippe Olivier Alexandre Navaux
IPDPS3
2013 Energy Efficient Last Level Caches via Last Read/Write Prediction
abstract
The size of the Last Level Caches (LLC) in multi-core architectures is increasing, and so is their power consumption. However, most of this power is wasted on unused or invalid cache lines. For dirty cache lines, the LLC waits until the line is evicted to be written back to memory. Hence, dirty lines compete for the memory bandwidth with read requests (prefetch and demand), increasing pressure on the memory controller. This paper proposes a Dead Line and Early Write-Back Predictor (DEWP) to improve the energy efficiency of the LLC. DEWP early evicts dead cache lines with an average accuracy of 94%, and only 2% false positives. DEWP also allows scheduling of dirty lines for early eviction, allowing earlier write-backs. Using DEWP over a set of single and multi-threaded benchmarks, we obtain an average of 61% static energy savings, while maintaining the performance, for both inclusive and non-inclusive LLCs.
Marco A. Z. Alves, Carlos Villavieja, Matthias Diener, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2013 Preserving the original MPI semantics in a virtualized processor environment
Eduardo Rocha Rodrigues, Philippe Olivier Alexandre Navaux, Jairo Panetta, Celso L. Mendes
Sci. Comput. Program.2
2012 Evaluating High Performance Computing on the Windows Azure Platform
abstract
Using the Cloud Computing paradigm for High-Performance Computing (HPC) is currently a hot topic in the research community and the industry. The attractiveness of Cloud Computing for HPC is the capability to run large applications on powerful, scalable hardware without needing to actually own or maintain this hardware. Most current research focuses on running HPC applications on the Amazon Cloud Computing platform, which is relatively easy because it supports environments that are similar to existing HPC solutions, such as clusters and supercomputers. In this paper, we evaluate the possibility of using Microsoft Windows Azure as a platform for HPC applications. Since most HPC applications are based on the Unix programming model, their source code has to be ported to the Windows programming model in addition to porting it to the Azure platform. We outline the challenges we encountered during porting applications and their resolutions. Furthermore, we introduce a metric to measure the efficiency of Cloud Computing platforms in terms of performance and price. We compared the performance and efficiency of running these benchmarks on a real machine, an Amazon EC2 instance and a Windows Azure instance. Results show that the performance of Azure is close to the performance of running on real machines, and that it is a viable alternative for running HPC applications when compared to other Cloud Computing solutions.
Eduardo Roloff, Francis B. Moreira 0001, Matthias Diener, Alexandre Carissimi, Philippe Olivier Alexandre Navaux
IEEE CLOUD5
2012 High Performance Computing in the cloud: Deployment, performance and cost efficiency
abstract
High-Performance Computing (HPC) in the cloud has reached the mainstream and is currently a hot topic in the research community and the industry. The attractiveness of cloud for HPC is the capability to run large applications on powerful, scalable hardware without needing to actually own or maintain this hardware. In this paper, we conduct a detailed comparison of HPC applications running on three cloud providers, Amazon EC2, Microsoft Azure and Rackspace. We analyze three important characteristics of HPC, deployment facilities, performance and cost efficiency and compare them to a cluster of machines. For the experiments, we used the well-known NAS parallel benchmarks as an example of general scientific HPC applications to examine the computational and communication performance. Our results show that HPC applications can run efficiently on the cloud. However, care must be taken when choosing the provider, as the differences between them are large. The best cloud provider depends on the type and behavior of the application, as well as the intended usage scenario. Furthermore, our results show that HPC in the cloud can have a higher performance and cost efficiency than a traditional cluster, up to 27% and 41%, respectively.
Eduardo Roloff, Matthias Diener, Alexandre Carissimi, Philippe Olivier Alexandre Navaux
CloudCom4
2012 Asymptotically Optimal Load Balancing for Hierarchical Multi-Core Systems
abstract
Current multi-core machines feature a complex and hierarchical core topology, multiple levels of cache and memory subsystem with NUMA design. Although this design provides high processing power to parallel machines, it comes with the cost of asymmetric memory access latencies. Depending on the parallel application communication patterns, this asymmetry may reduce the overall performance of the system. Therefore, to achieve scalable performance in this environment, it becomes crucial to exploit the machine architecture while taking into account the application communication patterns. In this paper, we introduce a topology-aware load balancing algorithm named HWTOPOLB. It combines the machine topology characteristics with the communication patterns of the application to equalize the application load on the available cores while reducing latencies. We also present the proof that the algorithm is asymptotically optimal (Theorem 1). We have implemented our load balancing algorithm using the CHARM++ Parallel System and analyzed its performance using three different benchmarks. Our experimental results show that the HWTOPOLB can achieve average performance improvements of 24% when compared to existing load balancing strategies on three different multi-core machines.
Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Jean-François Méhaut
ICPADS2
2012 A Hierarchical Approach for Load Balancing on Parallel Multi-core Systems
abstract
Multi-core compute nodes with non-uniform memory access (NUMA) are now a common architecture in the assembly of large-scale parallel machines. On these machines, in addition to the network communication costs, the memory access costs within a compute node are also asymmetric. Ignoring this can lead to an increase in the data movement costs. Therefore, to fully exploit the potential of these nodes and reduce data access costs, it becomes crucial to have a complete view of the machine topology (i.e. the compute node topology and the interconnection network among the nodes). Furthermore, the parallel application behavior has an important role in determining how to utilize the machine efficiently. In this paper, we propose a hierarchical load balancing approach to improve the performance of applications on parallel multi-core systems. We introduce NucoLB, a topology-aware load balancer that focuses on redistributing work while reducing communication costs among and within compute nodes. NucoLB takes the asymmetric memory access costs present on NUMA multi-core compute nodes, the interconnection network overheads, and the application communication patterns into account in its balancing decisions. We have implemented NucoLB using the Charm++ parallel runtime system and evaluated its performance. Results show that our load balancer improves performance up to 20% when compared to state-of-the-art load balancers on three different NUMA parallel machines.
Laércio Lima Pilla, Christiane Pousa Ribeiro, Daniel Cordeiro, Abhinav Bhatele, Philippe Olivier Alexandre Navaux, François Broquedis, Jean-François Méhaut, Laxmikant V. Kalé
ICPP6
2012 Using the Translation Lookaside Buffer to Map Threads in Parallel Applications Based on Shared Memory
abstract
The communication latency between the cores in multiprocessor architectures differs depending on the memory hierarchy and the interconnections. With the increase of the number of cores per chip and the number of threads per core, this difference between the communication latencies is increasing. Therefore, it is important to map the threads of parallel applications taking into account the communication between them. In parallel applications based on the shared memory paradigm, the communication is implicit and occurs through accesses to shared variables. For this reason, it is difficult to detect the communication pattern between the threads. Traditional approaches use simulation to monitor the memory accesses performed by the application, requiring modifications to the source code and drastically increasing the overhead. In this paper, we introduce a new light-weight mechanism to detect the communication pattern of threads using the Translation Look aside Buffer (TLB). Our mechanism relies entirely on hardware features, which makes the thread mapping transparent to the programmer and allows it to be performed dynamically by the operating system. Moreover, no time consuming task, such as simulation, is required. We evaluated our mechanism with the NAS Parallel Benchmarks (NPB) and achieved an accurate representation of the communication patterns. Using the detected communication patterns, we generated thread mappings using a heuristic method based on the Edmonds graph matching algorithm. Running the applications with these mappings resulted in performance improvements of up to 15.3%, reducing the number of cache misses by up to 31.1%.
Eduardo Henrique Molina da Cruz, Matthias Diener, Philippe Olivier Alexandre Navaux
IPDPS3
2012 DIMVHCM: An On-line Distributed Monitoring Data Collection Model
abstract
In this work we present a hierarchical distributed data collection model for monitoring systems, called DIMVHCM. Its goal is to provide data for the on-line analysis of the behavior of distributed systems and applications. In our work we focus on analysis through visualization. We implemented a prototype of this model, which was integrated to the DIM Visual prototype and to the TRIVA visualization tool. This way, we obtained a prototype on-line monitoring tool. We measured the time needed to send monitoring data from its origin (a collector) to a client. We also evaluated the performance of the client. Last, we measured the performance overhead caused by DIMVHCM to the programs that compose the NAS Parallel Benchmarks. In this experiment we measured less than 5% intrusiveness from DIMVHCM. We also show that it can support the visual analysis of distributed systems behavior.
Rafael Keller Tesser, Philippe Olivier Alexandre Navaux
PDP2
2012 Energy Savings via Dead Sub-Block Prediction
abstract
Cache memories have traditionally been designed to exploit spatial locality by fetching entire cache lines from memory upon a miss. However, recent studies have shown that often the number of sub-blocks within a line that are actually used is low. Furthermore, those sub-blocks that are used are accessed only a few times before becoming dead (i.e., never accessed again). This results in considerable energy waste since 1) data not needed by the processor is brought into the cache, and 2) data is kept alive in the cache longer than necessary. We propose the Dead Sub-Block Predictor (DSBP) to predict which sub-blocks of a cache line will be actually used and how many times it will be used in order to bring into the cache only those sub-blocks that are necessary, and power them off after they are touched the predicted number of times. We also use DSBP to identify dead lines (i.e., all sub-blocks off) and augment the existing replacement policy by prioritizing dead lines for eviction. Our results show a 24% energy reduction for the whole cache hierarchy when averaged over the SPEC2000, SPEC2006 and NAS-NPB benchmarks.
Marco A. Z. Alves, Khubaib, Eiman Ebrahimi, Veynu Narasiman, Carlos Villavieja, Philippe Olivier Alexandre Navaux, Yale N. Patt
SBAC-PAD6
2012 A hierarchical aggregation model to achieve visualization scalability in the analysis of parallel applications
Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux
Parallel Comput.3
2011 Combining Multiple Metrics to Control BSP Process Rescheduling in Response to Resource and Application Dynamics
abstract
This article discusses MigBSP: a rescheduling model that acts on Bulk Synchronous Parallel applications running over computational Grids. It combines the metrics Computation, Communication and Memory to make migration decisions. MigBSP also offers efficient adaptations to reduce its overhead. Additionally, MigBSP is infrastructure and application independent and tries to handle dynamicity on both levels. MigBSP's results show application performance improvements of up to 16% on dynamic environments while maintaining a small overhead when migrations do not take place.
Rodrigo da Rosa Righi, Lucas Graebin, Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, Laércio Lima Pilla
ICPADS4
2011 Improving Performance on Atmospheric Models through a Hybrid OpenMP/MPI Implementation
abstract
This work shows how a Hybrid MPI/OpenMP implementation can improve the performance of the Ocean-Land-Atmosphere Model (OLAM) on a multi-core cluster environment, which is a typical HPC many small files workload application. Previous experiments have shown that the scalability of this application on clusters is limited by the performance of the output operations. We show that the Hybrid MPI/OpenMP version of OLAM decreases the number of output files, resulting in better performance for I/O operations. We also observe that the MPI version of OLAM performs better for unbalanced workloads and that further parallel optimizations should be included on the hybrid version in order to improve the parallel execution time of OLAM.
Carla Osthoff, Pablo Javier Grunmann, Francieli Zanon Boito, Rodrigo Kassick, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Claudio Schepke, Jairo Panetta, Nicolas Maillard, Pedro Leite da Silva Dias, Robert L. Walko
ISPA6
2011 Dynamic I/O Reconfiguration for a NFS-Based Parallel File System
abstract
The large gap between the speed in which data can be processed and the performance of I/O devices makes the shared storage infrastructure of a cluster a great bottle-neck. Parallel File Systems try to smooth such difference by distributing data onto several servers, increasing the system's available bandwidth. However, most implementations use a fixed number of I/O servers, defined during the initialization of the system, and can not add new resources without a complete redistribution of the existing data. With the execution of different applications at the same time, the concurrent access to these resources can aggravate the existing bottleneck, making very hard to define an initial number of servers that satisfies the performance requirements of different applications. This paper presents a reconfiguration mechanism for the dNFSp file system that uses on-line monitoring of application's I/O behavior to detect performance contention and dedicate more I/O resources to applications with higher demands. These extra resources are taken from the available nodes of the cluster, using their I/O devices as a temporary storage. We show that this strategy is capable of increasing the I/O performance in up to 200% for access patterns with short I/O phases and 47% for longer I/O phases.
Rodrigo Kassick, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux
PDP3
2010 Optimizing an MPI weather forecasting model via processor virtualization
abstract
Weather forecasting models are computationally intensive applications. These models are typically executed in parallel machines and a major obstacle for their scalability is load imbalance. The causes of such imbalance are either static (e.g. topography) or dynamic (e.g. shortwave radiation, moving thunderstorms). Various techniques, often embedded in the application's source code, have been used to address both sources. However, these techniques are inflexible and hard to use in legacy codes. In this paper, we demonstrate the effectiveness of processor virtualization for dynamically balancing the load in BRAMS, a mesoscale weather forecasting model based on MPI parallelization. We use the Charm++ infrastructure, with its over-decomposition and object-migration capabilities, to move subdomains across processors during execution of the model. Processor virtualization enables better overlap between computation and communication and improved cache efficiency. Furthermore, by employing an appropriate load balancer, we achieve better processor utilization while requiring minimal changes to the model's code.
Eduardo Rocha Rodrigues, Philippe Olivier Alexandre Navaux, Jairo Panetta, Celso L. Mendes, Laxmikant V. Kalé
HiPC2
2010 Evaluating Thread Placement Based on Memory Access Patterns for Multi-core Processors
abstract
Process placement is a technique widely used on parallel machines with heterogeneous interconnects to reduce the overall communication time. For instance, two processes which communicate frequently are mapped close to each other. Finding the optimal mapping between threads and cores in a shared-memory environment (for example, OpenMP and Pthreads) is an even more complex task due to implicit communication. In this work, we examine data sharing patterns between threads in different workloads and use those patterns in a similar way as messages are used to map processes in cluster computers. We evaluated our technique on a state-of-the-art multicore processor and achieved moderate improvements in the common case and considerable improvements in some cases, reducing execution time by up to 45%.
Matthias Diener, Felipe Lopes Madruga, Eduardo Rocha Rodrigues, Marco A. Z. Alves, Jörg Schneider 0001, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
HPCC6
2010 Supporting performance and adaptivity on BSP process rescheduling
abstract
In this paper we will describe a model for BSP (Bulk Synchronous Parallel) process rescheduling called MigBSP. Considering the scope of BSP applications, its differential approach is the combination of three metrics - Memory, Computation and Communication - in order to measure the Potential of Migration of each BSP process. In this context, this paper addresses both the efficiency and the adaptivity perspectives of this model over our multi-cluster machine.
Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
ISCC4
2010 Impact of Parallel Workloads on NoC Architecture Design
abstract
Due to the multi-core processors, the importance of parallel workloads has increased considerably. However, many-core chips demand new interconnection strategies, since traditional crossbars or buses, common for current multi-core processors, have problems related to wires and scalability. For this reason, Networks-on-Chip (NoCs) have been developed in order to support the performance and parallelism focused on several workloads. Although a Network-on-Chip is a good option, most designs consist of a large number of routers. These routers are responsible for forwarding packets, and consequently, for supporting message-passing workloads. In this context, the NoC performance is a problem. Therefore, the main goal of this paper is to evaluate the impact of well-known parallel workloads on NoC architecture design. In order to achieve high performance, the results point out to parallel workloads with small packets and cluster-based NoCs with circuit switching and adaptable topologies.
Henrique Cota de Freitas, Lucas Mello Schnorr, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
PDP4
2010 Parallel Shared-Memory Workloads Performance on Asymmetric Multi-core Architectures
abstract
Putting performance asymmetric cores inside the same processor can be a good alternative to obtain high performance per area, throughput and single-threaded performance. However, the impact of running parallel applications on this type of machine is not clear, since most of previous work focused on multi-programmed and server workloads where there is low or no dependence between threads. In this work, we analyze the impact of running parallel shared-memory programs on heterogeneous multi-core setups using six parallel applications with diverse parallelization schemes. Moreover, we show that, in some cases, with a high number of cores, it is better to put one complex core than several simple ones. The impact of sharing the address space between asymmetric cores with private caches was also investigated and the number of invalidations per write access was not greater than a comparable homogeneous configuration.
Felipe Lopes Madruga, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux
PDP3
2010 Challenges and Issues of Supporting Task Parallelism in MPI
Márcia C. Cera, João V. F. Lima, Nicolas Maillard, Philippe Olivier Alexandre Navaux
EuroMPI4
2010 Impact of I/O Coordination on a NFS-Based Parallel File System with Dynamic Reconfiguration
abstract
The large gap between processing and I/O speed makes the storage infrastructure of a cluster a great bottleneck for UPC applications. Parallel File Systems propose a solution to this issue by distributing data onto several servers, dividing the load of I/O operations and increasing the available bandwidth. However, most parallel file systems use a fixed number of I/O servers defined during initialization and do not support addition of new resources as applications' demands grow. With the execution of different applications at the same time, the concurrent access to these resources can impact the performance and aggravate the existing bottleneck. The dNFSp File System proposes a reconfiguration mechanism that aims to include new I/O resources as application's demands grow. These resources are standard cluster nodes and are dedicated to a single application. This paper presents a study of the I/O performance of this reconfiguration mechanism under two circumstances: the use of several independent processes on a multi-core system or of a single centralized I/O process that coordinates the requests from all instances on a node. We show that the use of coordination can improve performance of applications with regular intervals between I/O phases. For applications with no such intervals, on the other hand, uncoordinated I/O presents better performance.
Rodrigo Kassick, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2010 A Comparative Analysis of Load Balancing Algorithms Applied to a Weather Forecast Model
abstract
Among the many reasons for load imbalance in weather forecasting models, the dynamic imbalance caused by localized variations on the state of the atmosphere is the hardest one to handle. As an example, active thunderstorms may substantially increase load at a certain time step with respect to previous time steps in an unpredictable manner - after all, tracking storms is one of the reasons for running a weather forecasting model. In this paper, we present a comparative analysis of different load balancing algorithms to deal with this kind of load imbalance. We analyze the impact of these strategies on computation and communication and the effects caused by the frequency at which the load balancer is invoked on execution time. This is done without any code modification, employing the concept of processor virtualization, which basically means that the domain is over-decomposed and the unit of rebalance is a sub-domain. With this approach, we were able to reduce the execution time of a full, real-world weather model.
Eduardo Rocha Rodrigues, Philippe Olivier Alexandre Navaux, Jairo Panetta, Alvaro Luiz Fazenda 0001, Celso L. Mendes, Laxmikant V. Kalé
SBAC-PAD2
2010 Triva: Interactive 3D visualization for performance analysis of parallel applications
Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux
Future Gener. Comput. Syst.3
2009 Towards Visualization Scalability through Time Intervals and Hierarchical Organization of Monitoring Data
abstract
Highly distributed systems such as grids are used today to the execution of large-scale parallel applications. The behavior analysis of these applications is not trivial. The complexity appears because of the event correlation among processes, external influences like time-sharing mechanisms and saturation of network links, and also the amount of data that registers the application behavior. Almost all visualization tools to analysis of parallel applications offer a space-time representation of the application behavior. This paper presents a novel technique that combines traces from grid applications with a treemap visualization of the data. With this combination, we dynamically create an annotated hierarchical structure that represents the application behavior for the selected time interval. The experiments in the grid show that we can readily use our technique to the analysis of large-scale parallel applications with thousands of processes.
Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux
CCGRID3
2009 On the design of reconfigurable crossbar switch for adaptable on-chip topologies in programmable NoC routers
abstract
Research works have focused on high-performance on-chip interconnections with low cost and energy consumption for the next generation of many-core processors. In the same way, parallel applications will explore thread level parallelism and message-passing communication through a Network-on-Chip (NoC) to perform a high data throughput. Due to the dynamic changing of communication patterns, the goal of this paper is to present the design of Reconfigurable Crossbar Switch for NoC Routers (RCS-NR) capable of adapting topologies on demand. RCS-NR has an optimized architecture, similar area and lower energy consumption (up to 87.41%) relative to a traditional crossbar switch. Furthermore, RCS-NR-based NoC router has a similar area, higher throughput, and it is up to 98.76% more efficient in energy consumption than a conventional NoC.
Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux
ACM Great Lakes Symposium on VLSI2
2009 MigBSP: A Novel Migration Model for Bulk-Synchronous Parallel Processes Rescheduling
abstract
We have developed a model called MigBSP that controls processes rescheduling in BSP (bulk synchronous parallel)applications. A BSP application is composed by one or more supersteps, each one containing both computation and communication phases followed by a synchronization barrier. Since the barrier waits for the slowest process, MigBSPpsilas final idea is to adjust the processes location in order to reduce the superstepspsila times. Considering the scope of the BSP model, the novel ideas of MigBSPare: (i) combination of three metrics - memory, computation and communication - to measure the potential of migration of each BSP process; (ii) use of both computation and communication patterns to control processespsila regularity;(iii) adaptation regarding the periodicity to launch the processes rescheduling. This paper describes MigBSP and presents some experimental results and related work.
Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß
HPCC4
2009 Design of Interleaved Multithreading for Network Processors on Chip
abstract
Thread level parallelism and multi-core processors are current alternatives to increase performance of general-purpose applications. In the same way, networks-on-ohip (NoCs) are the main alternatives for supporting packet throughput for the next generations of many-core processors. NPoC (network processor on chip) is a proposal to increase the performance of programmable NoC routers and multi-cluster NoC architectures using interleaved multithreading (IMT) technique. Therefore, the main goal of this paper is to present the design impact of interleaved multithreading for network processors on chip focusing on area and performance feasibility. Results show that NPoC-based router has an acceptable and similar area relative to a conventional NoC, and higher performance up to 7.1% than the same NPoC version without IMT.
Henrique Cota de Freitas, Felipe Lopes Madruga, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux
ISCAS4
2009 Multi-core aware process mapping and its impact on communication overhead of parallel applications
abstract
We propose an approach to reduce the execution time of applications with a steady communication pattern on clusters of multi-core processors by leveraging the asymmetry of core communication speeds. In addition to the well known fact that communication link speeds on a fixed cluster vary with processor selection, we consider one effect of multicore processor chips: link speeds vary with core selection within a single processor chip. The approach requires measuring link speeds among cluster cores as well as communication volumes and computational loads of the selected application processes. This data is fed into the dual recursive bipartitioning method to obtain close to optimal application process placement on cluster cores. We apply this approach to a real world application achieving sensible execution time reduction without even recompiling source code.
Eduardo Rocha Rodrigues, Felipe Lopes Madruga, Philippe Olivier Alexandre Navaux, Jairo Panetta
ISCC3
2009 Design of a Grid workflow for a climate application
abstract
Grid applications can be modeled as a composition of rather independent tasks. There are two approaches to define such a workflow either by combining multiple applications to build a more complex functionality or by splitting up an existing application. In this paper we analyze the latter process. We present a compute intensive application for climatology simulation and the options available to split it up. Using the simulation mode of our grid broker, we were able to compare the different workflow specifications before actually executing the workflows. This case study showed, using finer grained workflows-which usually need more adjustments to the software-allows better performance in the grid.
Jörg Schneider 0001, Julius Gehr, Hans-Ulrich Heiß, Tiago Ferreto, César A. F. De Rose, Rodrigo da Rosa Righi, Eduardo Rocha Rodrigues, Nicolas Maillard, Philippe Olivier Alexandre Navaux
ISCC9
2009 Performance Evaluation of NoC Architectures for Parallel Workloads
abstract
Network-on-Chip is the state-of-the-art approach to interconnect many processing cores in the next generation of general-purpose processors. In this context, the problem is to choose NoC architectures capable of achieving high performance for parallel programs. Therefore, the main goal of this paper is to evaluate the performance of three NoC architectures using well-known parallel workloads.
Henrique Cota de Freitas, Marco A. Z. Alves, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux
NOCS4
2009 Visual Mapping of Program Components to Resources Representation: A 3D Analysis of Grid Parallel Applications
abstract
Highly distributed systems such as Grids are usually interconnected by a hierarchical organization of different types of network. This strong network hierarchy influences directly the behavior of parallel applications. In order to obtain a good understanding of parallel application's behavior, the performance analysis must take into account a correspondence between application and network characteristics. This paper presents a novel way to analyze parallel applications, by using a three dimensional visualization and a technique to visually map the application's components to the used resources. This mapping technique is able to handle different types of resources description: different system logical organizations and the description of the network or system interconnection in different levels. The technique, implemented in our prototype Triva, is evaluated through a series of visual representations of the monitoring data obtained through real executions of parallel applications in a grid. The resulting visualizations enable an alternative and powerful way to view and understand various aspects of application behavior together with the network topology.
Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Guillaume Huard
SBAC-PAD2
2008 NOC architecture design for multi-cluster chips
abstract
For the next generation of multi-core processors, the on-chip interconnection networks must be efficient to achieve high data throughput and performance. Moreover, these interconnections must be flexible and scalable in order to provide parallel on-demand computing. For this reason, the goal of this paper is to present design decisions of a multi-cluster NoC (MCNoC) architecture in order to support collective communication patterns through topology reconfiguration on an FPGA-based multi-cluster chip. The MCNoCpsilas results show a small area occupation, low power consumption and high performance.
Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Tatiana Gadelha Serra dos Santos
FPL2
2008 Controlling Processes Reassignment in BSP Applications
abstract
We have developed a model for dynamic process scheduling in heterogeneous and non-dedicated environments. This model acts over a BSP (Bulk Synchronous Parallel) application, applying runtime processes reassignment to new processors. A BSP application is divided in one or more supersteps, each one containing both computation and communication phases followed by a barrier synchronization. In this context, the developed model combines three metrics - Memory, Computation and Communication - in order to measure the potential of migration of each BSP process. The final idea is to offer a mathematical formalism involving these metrics and to decide the following questions about the process migration: When? Where? Which? This paper presents the algorithms of our model, the parallel machine organization, some experimental results and related work.
Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2007 Processing Mesoscale Climatology in a Grid Environment
abstract
Enhancing the quality of weather and climate forecasts are central scientific research objectives worldwide. However, simulations of the atmosphere, usually demand high processing power and large storage resources. In this context, we present the GBRAMS project, that applies grid computing to speed up the generation of a regional model climatology for Brazil. A grid infrastructure was built to perform long-term integrations of a mesoscale numerical model (BRAMS), managing a queue of up to nine independent jobs submitted to three clusters spread over Brazil- Three distinct middlewares, Globus Toolkit, OurGrid and OAR/CIGRI, were compared in their ability to manage these jobs, and results on the usage of each node of the grid are provided. We analyze the impact of the resulted climatology in the accuracy of climate forecast, showing model bias removal which indicates correctness of the generated climatology. Our central contribution are how to use grid computing to speed-up climatology generation and the middleware impact on this enterprise.
Roberto Pinto Souto, Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, M. X. Py, Tiarajú Asmuz Diverio, Haroldo F. de Campos Velho, Stephan Stephany, Airam Jonatas Preto, Jairo Panetta, Eduardo Rocha Rodrigues, E. S. Almeida, Pedro Leite da Silva Dias, A. W. Gandu
CCGRID3
2007 The Use of Artificial Neural Networks in the Speech Understanding Model - SUM
Daniel Nehme Müller, Mozart Lemos de Siqueira, Philippe Olivier Alexandre Navaux
ICANN (2)3
2007 Evaluating Network-on-Chip for Homogeneous Embedded Multiprocessors in FPGAs
abstract
This paper presents performance and area evaluation of a homogeneous multiprocessor communication system based on network-on-chip (NoC) in FPGA platforms. Two homogenous chip multiprocessor proposals were designed and compared for Xilinx FPGAs using MicroBlaze processors: one based on NoC and the other based on shared memory/bus. One of the main findings is the communication performance evaluation of NoC for parallel computing applications. The comparison results show that an efficient implementation of NoC on FPGA can improve communication speed by up to seven times with low area overhead, according to the data size and the number of processors connected to the network.
Henrique Cota de Freitas, Dalton Martini Colombo, Fernanda Lima Kastensmidt, Philippe Olivier Alexandre Navaux
ISCAS4
2007 On-line Scheduling of MPI-2 Programs with Hierarchical Work Stealing
abstract
MPI (Message Passing Interface) is the de facto standard in High Performance Computing. By using some MPI- 2 new features, such as the dynamic creation of processes, it is possible to implement highly efficient parallel programs that can run on dynamic and/or heterogeneous resources, provided a good schedule of the processes can be computed at run-time. A classical solution to schedule parallel programs on-line is Work Stealing. However, its use with MPI- 2 is complicated by a restricted communication scheme between the processes: namely, spawned processes in MPI-2 can only communicate with their direct parents. This work presents an on-line scheduling algorithm, called Hierarchical Work Stealing, to obtain good load-balancing of MPI- 2 programs that follow a Divide & Conquer strategy. Experimental results are provided, based on a synthetic application, the N-Queens computation. The results show that the Hierarchical Work Stealing algorithm enables the use of MPI with high efficiency, even in parallel dynamic HPC platforms that are not as homogeneous as clusters.
Guilherme P. Pezzi, Márcia C. Cera, Elton N. Mathias, Nicolas Maillard, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2006 ICE: A Service Oriented Approach to Uniform the Access and Management of Cluster Environments
Clarissa Cassales Marquezan, Rodrigo da Rosa Righi, Lucas Mello Schnorr, Alexandre Carissimi, Nicolas Maillard, Philippe Olivier Alexandre Navaux
CCGRID6
2006 DIMVisual: Data Integration Model for Visualization of Parallel Programs Behavior
abstract
The development of high performance parallel applications for clusters is considered a complex task. This can happen because the influence of the execution environment and the non-deterministic natural behavior of this kind of applications. In such development, the programmer uses application traces and cluster monitoring tools to register the events of the application and the underlying execution environment. Generally, the analysis of the information from each source is made independently, making the correlation of events from the application with events from the execution environment difficult. This paper presents DIMVisual, a Data Integration Model which addresses this problem by integrating information from different sources and providing a unified visualization. An implementation of this model is also presented, using as data sources traces from MPI and DECK applications, events from Ganglia and Performance Co-Pilot cluster monitoring tools and operating system context switches. The results show the information gathered by these data sources integrated and visualized together in the generic visualization tool Paj´e, allowing the programmer a more complete view of his application behavior.
Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Benhur de Oliveira Stein
CCGRID2
2006 A Connectionist Approach to Speech Understanding
abstract
This paper argues that connectionist systems are a good approach to implement a speech understanding computational model. In this direction, we propose SUM, a speech understanding model, which is a software architecture based on neurocognitive researches. The SUM's computational implementation applies wavelets transforms to speech signal processing and connectionist models to syntactic parsing and prosodic-semantic mapping. This approach enables a system of analysis by compensation, when the syntactic analysis does not offer a good reliable level, it is possible to evaluate prosodic-semantic analysis, such as in human speech understanding.
Daniel Nehme Müller, Mozart Lemos de Siqueira, Philippe Olivier Alexandre Navaux
IJCNN3
2006 Scheduling Dynamically Spawned Processes in MPI-2
Márcia C. Cera, Guilherme P. Pezzi, Maurício L. Pilla, Nicolas Maillard, Philippe Olivier Alexandre Navaux
JSSPP5
2006 A Speculative Trace Reuse Architecture with Reduced Hardware Requirements
abstract
Trace reuse is an effective way of improving the performance of superscalar processors by skipping the execution of a sequence of instructions with known input and output values. However, the extra hardware complexity is of special concern when implementing such mechanisms. In this paper, we describe ways to reduce these requirements for Reuse through Speculation on Traces (RST). RST combines instruction and trace reuse with value prediction in an integrated mechanism to provide missing trace inputs when execution reaches the beginning of a trace. Speculatively reused traces do not consume resources in the execution pipeline, as they are not executed. In this paper, we study the effects of constraining reuse tables to effectively reduce the number of reuse candidates and comparisons. We compare our approach to instruction reuse, trace reuse and value prediction. We show that RST reuses more instructions and has better performance than traditional trace reuse, with an average speedup over a baseline without reuse of 1.21.
Maurício L. Pilla, Bruce R. Childers, Amarildo T. da Costa, Felipe M. G. França, Philippe Olivier Alexandre Navaux
SBAC-PAD5
2005 Evaluating the performance of the dNFSP file system
abstract
Parallel I/O in cluster computing is one of the most important issues to be tackled as clusters grow larger and larger. Many solutions have been proposed for the problem and, while effective in terms of performance, they usually represent a considerable amount of hacking into a "traditional" Beowulf cluster installation. In this paper, we investigate a parallel solution based on NFS, which reduces the level of intrusion in the file server installation, keeps the client side untouched, and still provides an improved level of performance and scalability for parallel applications. We compare our proposal to other existing file systems using known benchmarks, and demonstrate that it is a valid alternative for general-purpose cluster computing.
Rodrigo Kassick, Caciano Machado, Everton Hermann, Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, Yves Denneulin
CCGRID5
2005 Branch Prediction Topologies for SMT Architectures
abstract
The exploitation of instruction level parallelism in superscalar architectures is limited by data and control dependencies. Simultaneous multi-threaded (SMT) architectures can explore another level of parallelism, called thread-level parallelism, to fetch and execute instructions from different tasks at the same time. While a task is blocked by control or data dependencies, other tasks may continue executing, thus masking latencies caused by mispredicted branches and memory accesses, and increasing the occupation of functional units. However, the design of SMT architectures brings new challenges, such as determining the most efficient way to share resources among different threads. In this paper, we present different branch prediction topologies for SMT architectures. We show that the best results are obtained by matching the number of i-cache modules (fetch width) with the number of branch prediction modules (number of lookups and updates), while increasing the number of modules also helps increasing clock rates. Moreover, contention on branch prediction lookup and updates buses cannot be ignored on such architectures.
Guilherme Dal Pizzol, Philippe Olivier Alexandre Navaux
SBAC-PAD2
2005 Asynchronous Communication in Java over Infiniband and DECK
abstract
Java is becoming an attractive and easy to use programming language. It provides two systems for distributed computing, RMI and sockets, which describe a synchronous communication over TCP/IP. These Java core features may not be the best choice for cluster computing environments, since they do not provide high performance, a critical factor in this scenario. In this paper, presented is the development of Aldeia system, a library proposal to provide asynchronous communication in Java for cluster programming. Aldeia has been deployed on SAN hardware through Infiniband and DECK high-speed substrates. This paper shows the rationale for Aldeia's creation, its structure and encouraging results in synchronous and asynchronous approaches.
Rodrigo da Rosa Righi, Philippe Olivier Alexandre Navaux, Márcia C. Cera, Marcelo Pasin
SBAC-PAD2
2005 Reusing Traces in a Dynamic Conditional Execution Architecture
abstract
The cost of control and data dependences in superscalar processors is still an open issue, for what no definitive solution was yet found. Moreover, the cost of branch mispredictions is getting worse due to the increasing number of pipeline stages. The dynamic conditional execution (DCE) is a new approach to address this problem. The basic idea is to fetch and execute all paths produced by a branch that obey certain restrictions regarding complexity and size. As a consequence, a smaller number of predictions is performed, and therefore, a smaller number of branches is mispredicted. Although the execution of multiple paths of certain branches allows for a reduction in branch misprediction penalties, it implies on an increase in the number of executed instructions. Thus, an alternative to reduce the overhead created by DCE pipeline is to reuse previously executed values, freeing up resources for more useful instructions. The goal of this work is to analyze the impact of value reuse in DCE architecture. As it is presented, this effectively reduces the overhead produced by the architecture, increasing the overall performance. This paper shows that, in some cases, the speedup gain exceeds 60% over the original DCE architecture.
Tatiana Gadelha Serra dos Santos, Sergio Bampi, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2004 Performance Evaluation of a Prototype Distributed NFS Server
abstract
A high-performance file system is normally a key point for large cluster installations, where hundreds or even thousands of nodes frequently need to manage large volumes of data. While most solutions usually make use of dedicated hardware and/or specific distribution and replication protocols, the NFSP (NFS Parallel) project aims at improving performance within a standard NFS client/server system. In this paper we investigate the possibilities of a replication model for the NFS server, which is based on Lasy Release Consistency (LRC). A prototype has been built upon the user-level NFSv2 server and a performance evaluation is carried out.
Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, Pierre Lombard, Adrien Lèbre, Yves Denneulin
SBAC-PAD2
2004 Value Predictors for Reuse through Speculation on Traces
abstract
Reusing dynamic sequences of instructions - i.e., traces - improves performance for many benchmarks. However, many traces are not reused because of unavailable inputs in the reuse test. Reuse through speculation on traces (RST) aims to increase the number of reused traces by predicting those inputs when necessary, with minimal additional hardware when compared to nonspeculative trace reuse. In this paper, we compare last n-value and stride-aware prediction for trace inputs. Last n-value prediction uses the last recorded values as predictions, while stride-aware prediction identifies and uses strides to compute new predictions. Stride-aware RST has a higher hardware cost than last n-value RST and has also the shortcoming of not allowing branches inside predicted traces. This paper aims to determine which scheme is the most beneficial for RST. We show that stride values are important for reuse in RST and that last n-value prediction works as well as the more sophisticated stride-aware approach with simpler hardware.
Maurício L. Pilla, Philippe Olivier Alexandre Navaux, Bruce R. Childers, Amarildo T. da Costa, Felipe M. G. França
SBAC-PAD2
2003 An Oscillatory Neural Network for Image Segmentation
Dênis Fernandes, Philippe Olivier Alexandre Navaux
CIARP2
2003 Dynamic Load Balancing in PC Clusters: An Application to a Multi-Physics Model
abstract
We describe the use of dynamic load balancing in a PC cluster, applied to a multiphysics model that combines the parallel solution for three-dimensional (3D) PDEs of shallow water bodies flow and the parallel solution for the three-dimensional PDEs of scalar transportation of substances. The dynamic load balancing is obtained via diffusion algorithms. The numerical mesh is partitioned using RCB algorithm, in order to minimize communication and balance the load. Parallelism is obtained through Schwarz's additive domain decomposition method (DDM), so that the subproblems are solved concurrently. SPMD is the programming model used and the message passing between processes in the PC cluster is done with MPICH library.
Ricardo Vargas Dorneles, Rogério Luís Rizzi, Tiarajú Asmuz Diverio, Philippe Olivier Alexandre Navaux
SBAC-PAD4
2003 Complex Branch Profiling for Dynamic Conditional Execution
abstract
Branch predictors are widely used as an alternative to deal with conditional branches. Despite the high accuracy rates, misprediction penalties are still large in any superscalar pipeline. DCE, or dynamic conditional execution, is an alternative to reduce the number of predicted branches by executing both paths of certain branches, reducing the number of predictions and, therefore, the occurrence of mispredictions. The goal of this work is to analyze the complexity of branch structures and determine the number of branches that can be predicated in DCE and the distribution of mispredictions according to the proposed classification. The complex branch classification proposed extends the classification presented by Klauser [A. Klauser, et al., (1998)]. As result, we show that an average of 35% of all branches can be predicated in DCE and around 32% of all mispredictions fall into these branches.
Rafael R. dos Santos, Tatiana Gadelha Serra dos Santos, Maurício L. Pilla, Philippe Olivier Alexandre Navaux, Sergio Bampi, Mario Nemirovsky
SBAC-PAD4
2003 Performance Analysis of DECK Collective Communication Service
abstract
Collective communication is very useful for parallel applications, especially those in which matrix and vector data structures need to be manipulated by a group of processes. We present a performance analysis of collective communication primitives designed for the DECK parallel programming environment, with the aid of different numerical methods used to solve hydrodynamics and mass transportation models.
Rafael Ennes Silva, Delcino Picinin, Marcos E. Barreto, Rafael Bohrer Ávila, Tiarajú Asmuz Diverio, Philippe Olivier Alexandre Navaux
SBAC-PAD6
2002 Architecture of Oscillatory Neural Network for Image Segmentation
abstract
Oscillatory neural networks are a recent approach for applications in image segmentation. In this context, the LEGION (Locally Excitatory Globally Inhibitory Oscillator Network) is the most consistent proposal. As positive aspects, the network has got a parallel architecture and capacity to separate the segments in time. On the other hand, the structure based on differential equations presents high computational complexity and limited capacity of segmentation, which restricts practical applications. In this paper, a proposal of a parallel architecture for implementation of an oscillatory neural network suitable for image segmentation is presented. The proposed network keeps the positive features of the LEGION network, offering lower complexity for implementation in digital hardware and capacity of segmentation unlimited, as well as a few parameters, with an intuitive setting. Preliminary results confirm the successful operation of the proposed network in applications of image segmentation.
Dênis Fernandes, J. Stedile, Philippe Olivier Alexandre Navaux
SBAC-PAD3
2001 DECK-SCI: High-Performance Communication and Multithreading for SCI Clusters
abstract
This paper presents the design and implementation of DECK-SCI, a multithreaded communication library that fully exploits the high-performance capabilities of the SCI technology. We compare DECK-SCI, in terms of performance, to a commercially distributed MPI implementation and to a freely available MPICH distribution, both specifically designed for SCI clusters.
Fabio A. D. de Oliveira, Rafael Bohrer Ávila, Marcos E. Barreto, Philippe Olivier Alexandre Navaux, César A. F. De Rose
CLUSTER4
2000 Distributed Processor Allocation in Multicomputers
Rose Rose, Hans-Ulrich Heiß, Philippe Olivier Alexandre Navaux
CLUSTER3
2000 Distributed Processor Allocation in Large PC Clusters
abstract
Current processor allocation techniques for highly parallel systems are based on centralized front-end based algorithms. As a result, the applied strategies are restricted to static allocation, low parallelism and weak fault tolerance. To lift these restrictions, we are investigating a distributed approach to the processor allocation problem in large distributed memory machines. A contiguous and a noncontiguous version of a distributed dynamic processor allocation strategy are proposed and studied. Simulations compare the performance of the proposed strategies with that of well-known centralized algorithms. We also present the results of experiments on a Simens hpcline Primergy Server with 96 nodes that show distributed allocation is feasible with current technologies.
Hans-Ulrich Heiß, César A. F. De Rose, Philippe Olivier Alexandre Navaux
HPDC3
1998 Analysing a Multistreamed Superscalar Speculative Fetch Mechanism
Rafael R. dos Santos, Philippe Olivier Alexandre Navaux
Euro-Par2
1995 Performance evaluation in image processing with GAPP array processor
Philippe Olivier Alexandre Navaux, César A. F. De Rose, Gerson G. H. Cavalheiro
Microprocess. Microprogramming1
1988 SARA: A processor interconnection performance analysis tool
Philippe Olivier Alexandre Navaux, Paulo Fernandes 0001, Maurizio Tazza
Microprocess. Microprogramming1