VLDB 2026 Research / reviewers in the wild / expert
Laércio Lima Pilla
dblp:75/2582
· DBLP profile ↗
47ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0003-0997-586XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 13 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TOTO: Transparent I/O Tuning for HPC ApplicationsabstractHigh-performance computing applications rely on parallel file systems, where I/O performance is strongly affected by configuration parameters such as stripe count. However, the ideal stripe count is highly application- and system-dependent, making it difficult to predict and rarely tuned in practice. As a result, substantial I/O performance potential remains unexplored. We present TOTO, a transparent tool that improves I/O performance without requiring application modifications. TOTO intercepts POSIX calls, characterizes application behavior, and uses a machine learning model to select an appropriate stripe count, even for already opened files. We also introduce an allocation algorithm that balances performance and resource occupation, and describe a methodology for training the model once per system using limited data. Our results show that TOTO can improve I/O performance by up to 4.6 × compared to using a default stripe count, while imposing an overhead of at most \(8\%\). Moreover, compared to the state of the art, TOTO can optimize more applications with a lower resource occupation, which is expected to decrease contention in the I/O infrastructure. Francieli Zanon Boito, Luan Teylo, Mihail Popov, Laora Aimi, Alexis Bandet, Laércio Lima Pilla, Guillaume Pallez |
ICS | 6 |
| 2026 | MetaCS-FL: A metaheuristic-based framework for client selection in federated learning systems
Alan L. Nunes, Cristina Boeres, Laércio Lima Pilla, Lúcia M. A. Drummond |
Future Gener. Comput. Syst. | 3 |
| 2026 | Energy-aware scheduling strategies for partially-replicable task chains on heterogeneous processors
Yacine Idouar, Adrien Cassagne, Laércio Lima Pilla, Julien Sopena, Manuel Bouyer, Diane Orhan, Lionel Lacassagne, Dimitri Galayko, Denis Barthou, Christophe Jégo |
Parallel Comput. | 3 |
| 2025 | Near-Optimal Contraction Strategies for the Scalar Product in the Tensor-Train Format
Atte Torri, Przemyslaw Dominikowski, Brice Pointal, Oguz Kaya, Laércio Lima Pilla, Olivier Coulaud |
Euro-Par (3) | 5 |
| 2025 | Task-Based HPC in the Cloud: Price-Performance Analysis of N-Body Simulations with StarPUabstractPublic cloud environments present significant challenges for traditional High Performance Computing (HPC) applications due to infrastructure limitations that differ substantially from dedicated HPC systems. Unlike traditional HPC clusters optimized for tightly coupled parallel workloads, cloud platforms were designed primarily for web services and data processing applications. Key obstacles include high-latency networks, hardware virtualization overhead, and limited availability of specialized accelerators, all of which can severely impact the performance of compute-intensive applications such as physics simulations. This study investigates the feasibility of running HPC workloads on public cloud infrastructure using standard and cost-effective instance configurations rather than expensive specialized "HPC" offerings. We deploy heterogeneous clusters on Amazon Web Services using the HPC@Cloud Toolkit, incorporating various instance types, including GPU-accelerated nodes with different computational capabilities. Our evaluation focuses on N-body simulations implemented using a task-based parallel programming model, leveraging the StarPU runtime system to dynamically schedule computational tasks across various processing units. Our experimental results demonstrate three key findings: (1) smaller GPU-equipped instances (g6.2xlarge) achieve performance comparable to larger instances while costing approximately one-sixth the price, challenging conventional scaling assumptions for cloud-based HPC; (2) strategic GPU utilization yields up to 8.2× performance improvements over CPU-only configurations while reducing total execution costs by 24.4×; and (3) while task-based programming models effectively address network limitations through dynamic scheduling, complex tree-based algorithms like TBFMM face significant optimization challenges in cloud environments due to load balancing issues and expensive parameter tuning requirements. These findings provide practical guidance for researchers and practitioners seeking cost-effective cloud HPC deployments, demonstrating that commodity cloud infrastructures can be viable for regular computational workloads but require careful algorithmic-resource matching for optimal efficiency. Nicolas Vanz, Vanderlei Munhoz, Márcio Castro 0001, Laércio Lima Pilla, Olivier Aumage |
IC2E | 4 |
| 2025 | Optimal scheduling algorithms for software-defined radio pipelined and replicated task chains on multicore architecturesabstractSoftware-Defined Radio (SDR) represents a move from dedicated hardware to software implementations of digital communication standards. This approach offers flexibility, shorter time to market, maintainability , and lower costs, but it requires an optimized distribution tasks in order to meet performance requirements. Thus, we study the problem of scheduling SDR linear task chains of stateless and stateful tasks for streaming processing. We model this problem as a pipelined workflow scheduling problem based on pipelined and replicated parallelism on homogeneous resources. We propose an optimal dynamic programming solution and an optimal greedy algorithm named OTAC for maximizing throughput while also minimizing resource utilization . Moreover, the optimality of the proposed scheduling algorithm is proved. We evaluate our solutions and compare their execution times and schedules to other algorithms using synthetic task chains and an implementation of the DVB-S2 communication standard on the AFF3CT SDR Domain Specific Language . Our results demonstrate how OTAC quickly finds optimal schedules, leading consistently to better results than other algorithms, or equivalent results with fewer resources. Diane Orhan, Laércio Lima Pilla, Denis Barthou, Adrien Cassagne, Olivier Aumage, Romain Tajan, Christophe Jégo, Camille Leroux |
J. Parallel Distributed Comput. | 2 |
| 2025 | Approximation Algorithms for Scheduling With/Without Deadline Constraints Where Rejection Costs are Proportional to Processing TimesabstractWe study two offline job scheduling problems where tasks can be processed on a limited number of energy-efficient edge machines or offloaded to an unlimited supply of energy-inefficient cloud machines (called rejected). The objective is to minimize total energy consumption. First, we consider scheduling without deadlines, formulating it as a scheduling problem with rejection, where rejection costs are proportional to processing times. We propose a novel$\frac{5}{4}(1+\epsilon )$-approximation algorithm,$\mathcal{BEKP}$, by associating it to a Multiple Subset Sum problem, improving upon the existing$(\frac{3}{2} - \frac{1}{2m})$-approximation for arbitrary rejection costs. Next, we address scheduling with deadlines, aiming to minimize the weighted number of rejected jobs. We position this problem within the literature and introduce a new$(1-\frac{(m-1)^{m}}{m^{m}})$-approximation algorithm,$\mathcal{MDP}$, inspired by an interval selection algorithm with a$(1-\frac{m^{m}}{(m+1)^{m}})$-approximation for arbitrary rejection costs. Experimental results demonstrate that$\mathcal{BEKP}$and$\mathcal{MDP}$obtain better results (lower costs or higher profits) than other state-of-the-art algorithms while maintaining a competitive or better time complexity. Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz, Laércio Lima Pilla, Alexandre van Kempen |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | A 1.25(1+ε )-Approximation Algorithm for Scheduling with Rejection Costs Proportional to Processing Times
Olivier Beaumont, Rémi Bouzel, Lionel Eyraud-Dubois, Esragul Korkmaz, Laércio Lima Pilla, Alexandre van Kempen |
Euro-Par (1) | 5 |
| 2024 | Optimal Time and Energy-Aware Client Selection Algorithms for Federated Learning on Heterogeneous ResourcesabstractFederated Learning systems allow training machine learning models distributed across multiple clients, each one using private local data. Iteratively, the clients send their training contributions to a server, which performs a merge to produce an enhanced global model. Due to resource and data heterogeneity, client selection is crucial to optimize the system efficiency and improve the global model generalization. Selecting more clients is likely to increase the overall energy consumption, while a small number of clients may decline the performance of the trained model or require longer training time. We propose two time- and energy-aware client selection algorithms, MEC and ECMTC, which are proven regarding their optimality and evaluated against state-of-the-art algorithms on an extensive series of experiments in both simulation and HPC platform scenarios. The results indicate the benefits of jointly optimizing the time and energy consumption metrics using our proposals. Alan L. Nunes, Cristina Boeres, Lúcia M. A. Drummond, Laércio Lima Pilla |
SBAC-PAD | 4 |
| 2023 | Optimizing performance and energy across problem sizes through a search space exploration and machine learning
Lana Scravaglieri, Mihail Popov, Laércio Lima Pilla, Amina Guermouche, Olivier Aumage, Emmanuelle Saillard |
J. Parallel Distributed Comput. | 3 |
| 2023 | Scheduling Algorithms for Federated Learning With Minimal Energy ConsumptionabstractFederated Learning (FL) has opened the opportunity for collaboratively training machine learning models on heterogeneous mobile or Edge devices while keeping local data private. With an increase in its adoption, a growing concern is related to its economic and environmental cost (as is also the case for other machine learning techniques). Unfortunately, little work has been done to optimize its energy consumption or emissions of carbon dioxide or equivalents, as energy minimization is usually left as a secondary objective. In this paper, we investigate the problem of minimizing the energy consumption of FL training on heterogeneous devices by controlling the workload distribution. We model this as the Minimal Cost FL Schedule problem, a total cost minimization problem with identical, independent, and atomic tasks that have to be assigned to heterogeneous resources with arbitrary cost functions. We propose a pseudo-polynomial optimal solution to the problem based on the previously unexplored Multiple-Choice Minimum-Cost Maximal Knapsack Packing Problem. We also provide four algorithms for scenarios where cost functions are monotonically increasing and follow the same behavior. These solutions are likewise applicable on the minimization of other kinds of costs, and in other one-dimensional data partition problems. Laércio Lima Pilla |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Exploring Scheduling Algorithms for Parallel Task Graphs: A Modern Game Engine Case Study
Mustapha Regragui, Baptiste Coye, Laércio Lima Pilla, Raymond Namyst, Denis Barthou |
Euro-Par | 3 |
| 2022 | Algorithm Selection Framework for Legalization Using Deep Convolutional Neural Networks and Transfer LearningabstractMachine learning (ML) models have been used to improve the quality of different physical design steps, such as timing analysis, clock tree synthesis, and routing. However, so far very few works have addressed the problem of algorithm selection during physical design, which can drastically reduce the computational effort of some steps. This work proposes a legalization algorithm selection framework using deep convolutional neural networks (CNNs). To extract features, we used snapshots of circuit placements and used transfer learning to train the models using pretrained weights of the Squeezenet architecture. By doing so, we can greatly reduce the training time and required data even though the pretrained weights come from a different problem. We performed extensive experimental analysis of ML models, providing details on how we chose the parameters of our model, such as CNN architecture, learning rate, and number of epochs. We evaluated the proposed framework by training a model to select between different legalization algorithms according to cell displacement and wirelength variation. The trained models achieved an average$F$-score of 0.98 when predicting cell displacement and 0.83 when predicting wirelength variation. When integrated into the physical design flow, the cell displacement model achieved the best results on 15 out of 16 designs, while the wirelength variation model achieved that for 10 out of 16 designs, being better than any individual legalization algorithm. Finally, using the proposed ML model for algorithm selection resulted in a speedup of up to$10\times $compared to running all the algorithms separately. Renan Netto, Sheiny Fabre Almeida, Tiago Fontana, Vinicius S. Livramento, Laércio Lima Pilla, Laleh Behjat, José Luís Güntzel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Optimal Task Assignment for Heterogeneous Federated Learning DevicesabstractFederated Learning provides new opportunities for training machine learning models while respecting data privacy. This technique is based on heterogeneous devices that work together to iteratively train a model while never sharing their own data. Given the synchronous nature of this training, the performance of Federated Learning systems is dictated by the slowest devices, also known as stragglers. In this paper, we investigate the problem of minimizing the duration of Federated Learning rounds by controlling how much data each device uses for training. We formulate this as a makespan minimization problem with identical, independent, and atomic tasks that have to be assigned to heterogeneous resources with non-decreasing cost functions, while also respecting lower and upper limits of tasks per resource. Based on this formulation, we propose a polynomial-time algorithm named OLAR and prove that it provides optimal schedules. We evaluate OLAR in an extensive series of experiments using simulation that includes comparisons to other algorithms from the state of the art, and new extensions to them. Our results indicate that OLAR provides optimal solutions with a small execution time. They also show that the presence of lower and upper limits of tasks per resource erase any benefits that suboptimal heuristics could provide in terms of algorithm execution time. Laércio Lima Pilla |
IPDPS | 1 |
| 2021 | PackStealLB: A scalable distributed load balancer based on work stealing and workload discretization
Vinicius Freitas, Laércio Lima Pilla, Alexandre de Limas Santana, Márcio Castro 0001, Johanne Cohen |
J. Parallel Distributed Comput. | 2 |
| 2021 | ARTful: A model for user-defined schedulers targeting multiple high-performance computing runtime systemsabstractAbstract Global schedulers are components in parallel runtime libraries that distribute the application's workload across physical resources. More often than not, applications showcase dynamic load imbalance and require customized scheduling solutions to avoid wasting resources. Some libraries lack support for user‐defined schedulers and developers resort to unofficial extensions that are harder to reuse and maintain. We propose a global scheduler software design, entitled ARTful model, to create user‐defined solutions with minimal alterations in the runtime library. Our model uses a component‐based design to separate components from the runtime library and the scheduling policy implementation. The ARTful modeldescribes the interface of a portable scheduler library, allowing policies to operate on different runtime libraries. We study the overhead induced by our design through our ARTful library implementation metaprogramming‐oriented global scheduling library using workload‐aware scheduling policies. We experiment with two different policies from OpenMP and Charm++ runtime systems, also presenting evaluations of the policies outside of their original library context. We observe that our portable schedulers can sometimes perform decisions faster than their native counterparts with negligible overhead in the execution times of synthetic applications and molecular dynamics kernels. Alexandre de Limas Santana, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla, Jean-François Méhaut |
Softw. Pract. Exp. | 4 |
| 2020 | Adaptive Load Balancing based on Machine Learning for Iterative Parallel ApplicationsabstractThe performance of irregular scientific applications can be easily affected by an uneven distribution of work among the computing resources. In this context, Load Balancing (LB) stands as one of the most important solutions to improve resource utilization. However, choosing the best-performing load balancing algorithm for a given application is not a trivial task. For instance, manually and statically choosing an LB algorithm does not work in situations where applications have a dynamic or unknown behavior. In this context, we propose a Machine Learning-based Adaptive Load Balancer (ADAPTIVELB) to automate the load balancing algorithm decision at run time. This approach monitors and collects information about the application dynamically, and according to the analyzed data, it makes a decision of invoking the most suitable LB algorithm. Our experiments show that ADAPTIVELB can select a good load balancing algorithm in most of the cases, leading to performance improvements over statically chosen LB algorithms and over the absence of a load balancer. C. R. Anna Victoria Oikawa, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla |
PDP | 4 |
| 2019 | How Deep Learning Can Drive Physical Synthesis Towards More Predictable LegalizationabstractMachine learning has been used to improve the predictability of different physical design problems, such as timing, clock tree synthesis and routing, but not for legalization. Predicting the outcome of legalization can be helpful to guide incremental placement and circuit partitioning, speeding up those algorithms. In this work we extract histograms of features and snapshots of the circuit from several regions in a way that the model can be trained independently from region size. Then, we evaluate how traditional and convolutional deep learning models use this set of features to predict the quality of a legalization algorithm without having to executing it. When evaluating the models with holdout cross validation, the best model achieves an accuracy of 80% and an F-score of at least 0.7. Finally, we used the best model to prune partitions with large displacement in a circuit partitioning strategy. Experimental results in circuits (with up to millions of cells) showed that the pruning strategy improved the maximum displacement of the legalized solution by 5% to 94%. In addition, using the machine learning model avoided from 22% to 99% of the calls to the legalization algorithm, which speeds up the pruning process by up to 3x. Renan Netto, Sheiny Fabre Almeida, Tiago Fontana, Vinicius S. Livramento, Laércio Lima Pilla, José Luís Güntzel |
ISPD | 5 |
| 2018 | Improving Communication and Load Balancing with Thread Mapping in Manycore SystemsabstractCommunication and load balancing have a significant impact on the performance of parallel applications and have been the subject of extensive research in multicore architectures. Thread mapping has been one of the solutions adopted in multicore architectures to address both communication and load balancing. However, the impact of such issues on more recently introduced manycore architectures is still unknown. Most related work on manycore architectures focus on execution time and idleness information for scheduling decisions. In this paper, we improve the state of the art by performing a very detailed analysis of the impact of thread mapping on communication and load balancing in two manycore systems from Intel, namely Knights Corner and Knights Landing. We observed that the widely used metric of CPU time provides very inaccurate information for load balancing. We also evaluated the usage of thread mapping based on the communication and load information of the applications to improve the performance of manycore systems. Eduardo Henrique Molina da Cruz, Matthias Diener, Matheus S. Serpa, Philippe Olivier Alexandre Navaux, Laércio Lima Pilla, Israel Koren |
PDP | 5 |
| 2018 | A Batch Task Migration Approach for Decentralized Global ReschedulingabstractEffectively mapping tasks of High Performance Computing (HPC) applications on parallel systems is crucial to assure substantial performance gains. As platforms and applications grow, load imbalance becomes a priority issue. Even though centralized rescheduling has been a viable solution to mitigate this problem, its efficiency is not able to keep up with the increasing size of shared memory platforms. To efficiently solve load imbalance today, and in the years to come, we should prioritize decentralized strategies developed for large scale platforms. In this paper, we propose our Batch Task Migration approach to improve decentralized global rescheduling, ultimately reducing communication costs and preserving task locality. We implemented and evaluated our approach in two different parallel platforms, using both synthetic workloads and a molecular dynamics (MD) benchmark. Our solution was able to achieve speedups of up to 3.75 and 1.15 on rescheduling time, when compared to other centralized and distributed approaches, respectively. Moreover, it improved the execution time of MD by factors up to 1.34 and 1.22 when compared to a scenario without load balancing on two different platforms. Vinicius Freitas, Alexandre de Limas Santana, Márcio Castro 0001, Laércio Lima Pilla |
SBAC-PAD | 4 |
| 2018 | A branch and bound strategy for Fast Trajectory Similarity Measuring
Andre Salvaro Furtado, Laércio Lima Pilla, Vania Bogorny |
Data Knowl. Eng. | 2 |
| 2018 | MigPF: Towards on self-organizing process rescheduling of Bulk-Synchronous Parallel applications
Rodrigo da Rosa Righi, Roberto de Quadros Gomes, Vinicius Facco Rodrigues, Cristiano André da Costa, Antônio Marcos Alberti, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
Future Gener. Comput. Syst. | 6 |
| 2018 | A Cost Model for IaaS Clouds Based on Virtual Machine Energy Consumption
Mauro Hinz, Guilherme P. Koslovski, Charles Miers, Laércio Lima Pilla, Maurício Aronne Pillon |
J. Grid Comput. | 4 |
| 2017 | Radiation-Induced Error Criticality in Modern HPC Parallel AcceleratorsabstractIn this paper, we evaluate the error criticality of radiation-induced errors on modern High-Performance Computing (HPC) accelerators (Intel Xeon Phi and NVIDIA K40) through a dedicated set of metrics. We show that, as long as imprecise computing is concerned, the simple mismatch detection is not sufficient to evaluate and compare the radiation sensitivity of HPC devices and algorithms. Our analysis quantifies and qualifies radiation effects on applications' output correlating the number of corrupted elements with their spatial locality. Also, we provide the mean relative error (dataset-wise) to evaluate radiation-induced error magnitude. We apply the selected metrics to experimental results obtained in various radiation test campaigns for a total of more than 400 hours of beam time per device. The amount of data we gathered allows us to evaluate the error criticality of a representative set of algorithms from HPC suites. Additionally, based on the characteristics of the tested algorithms, we draw generic reliability conclusions for broader classes of codes. We show that arithmetic operations are less critical for the K40, while Xeon Phi is more reliable when executing particles interactions solved through Finite Difference Methods. Finally, iterative stencil operations seem the most reliable on both architectures. Daniel Oliveira 0002, Laércio Lima Pilla, Mauricio Hanzich, Vinicius Fratin, Fernando Santos 0001, Caio B. Lunardi, José María Cela, Philippe Olivier Alexandre Navaux, Luigi Carro, Paolo Rech |
HPCA | 2 |
| 2017 | How Game Engines Can Inspire EDA Tools Development: A use case for an open-source physical design libraryabstractSimilarly to game engines, physical design tools must handle huge amounts of data. Although the game industry has been employing modern software development concepts such as data-oriented design, most physical design tools still relies on object-oriented design. Differently from object-oriented design, data-oriented design focuses on how data is organized in memory and can be used to solve typical object-oriented design problems. However, its adoption is not trivial because most software developers are used to think about objects' relationships rather than data organization. The entity-component design pattern can be used as an efficient alternative. It consists in decomposing a problem into a set of entities and their components (properties). This paper discusses the main data-oriented design concepts, how they improve software quality and how they can be used in the context of physical design problems. In order to evaluate this programming model, we implemented an entity-component system using the open-source library Ophidian. Experimental results for two physical design tasks show that data-oriented design is much faster than object-oriented design for problems with good data locality, while been only sightly slower for other kinds of problems. Tiago Fontana, Renan Netto, Vinicius S. Livramento, Chrystian Guth, Sheiny Fabre Almeida, Laércio Lima Pilla, José Luís Güntzel |
ISPD | 6 |
| 2017 | Experimental and analytical study of Xeon Phi reliabilityabstractWe present an in-depth analysis of transient faults effects on HPC applications in Intel Xeon Phi processors based on radiation experiments and high-level fault injection. Besides measuring the realistic error rates of Xeon Phi, we quantify Silent Data Corruption (SDCs) by correlating the distribution of corrupted elements in the output to the application's characteristics. We evaluate the benefits of imprecise computing for reducing the programs' error rate. For example, for HotSpot a 0.5% tolerance in the output value reduces the error rate by 85%. Daniel Oliveira 0002, Laércio Lima Pilla, Nathan DeBardeleben, Sean Blanchard, Heather M. Quinn, Israel Koren, Philippe Olivier Alexandre Navaux, Paolo Rech |
SC | 2 |
| 2016 | A Sharing-Aware Memory Management Unit for Online Mapping in Multi-core Architectures
Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
Euro-Par | 3 |
| 2016 | Value Reuse Potential in ARM ArchitecturesabstractCode execution in modern superscalar processors is inherently redundant. Many instructions execute repeatedly with the same inputs, producing the same outputs, thus wasting resources in the process. Value reuse techniques memorize previous executions of instructions, blocks or traces which may be reused if they appear again with the same input contexts. Although trace reuse techniques show great potential for both performance and energy consumption improvement, they have not been studied yet in one of the most widely available computer architectures - the ARM architecture. In this paper, the main issues with reusing traces in instruction sets with conditional execution are revisited. Afterwards, the reuse potential in the benchmark suite MiBench is analyzed varying (i) how traces are generated, and (ii) the size of reuse tables. Our results show that a memoization table of 32 KiB allows to reuse 18.36% of the total instructions on average. Rodrigo Costa de Moura, Giovane O. Torres, Maurício L. Pilla, Laércio Lima Pilla, Amarildo T. da Costa, Felipe M. G. França |
SBAC-PAD | 4 |
| 2016 | LAPT: A locality-aware page table for thread and data mapping
Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
Parallel Comput. | 4 |
| 2016 | Hardware-Assisted Thread and Data Mapping in Hierarchical Multicore ArchitecturesabstractThe performance and energy efficiency of modern architectures depend on memory locality, which can be improved by thread and data mappings considering the memory access behavior of parallel applications. In this article, we propose intense pages mapping, a mechanism that analyzes the memory access behavior using information about the time the entry of each page resides in the translation lookaside buffer. It provides accurate information with a very low overhead. We present experimental results with simulation and real machines, with average performance improvements of 13.7% and energy savings of 4.4%, which come from reductions in cache misses and interconnection traffic. Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Evaluation and Mitigation of Radiation-Induced Soft Errors in Graphics Processing UnitsabstractGraphics processing units (GPUs) are increasingly attractive for both safety-critical and High-Performance Computing applications. GPU reliability is a primary concern for both the automotive and aerospace markets and is becoming an issue also for supercomputers. In fact, the high number of devices in large data centers makes the probability of having at least a device corrupted to be very high. In this paper, we aim at giving novel insights on GPU reliability by evaluating the neutron sensitivity of modern GPUs memory structures, highlighting pattern dependence and multiple errors occurrences. Additionally, a wide set of parallel codes are exposed to controlled neutron beams to measure GPUs operative error rates. From experimental data and algorithm analysis we derive general insights on parallel algorithms and programming approaches reliability. Finally, error-correcting code, algorithm-based fault tolerance, and duplication with comparison hardening strategies are presented and evaluated on GPUs through radiation experiments. We present and compare both the reliability improvement and imposed overhead of the selected hardening solutions. Daniel Oliveira 0002, Laércio Lima Pilla, Thiago Santini, Paolo Rech |
IEEE Trans. Computers | 2 |
| 2015 | An Efficient Algorithm for Communication-Based Task MappingabstractThe communication between tasks of a parallel application is an important characteristic to consider when mapping tasks to computing cores due to possible differences in communication performance. Within a machine, performance differences are introduced by the memory hierarchy, in which cache memories can be shared by groups of cores and intra-chip interconnections are faster than inter-chip interconnections. In cluster and grid systems, the network imposes an additional communication latency. By mapping tasks that communicate to cores nearby on the memory hierarchy, or to the same nodes in clusters or grids, the communication of parallel applications is optimized, leading to increased performance and energy efficiency. In the task mapping context, one of the most important aspects to be considered is the mapping algorithm, as it determines the improvements that can be achieved. Since the problem of finding the best mapping is NP-Hard, heuristics must be employed to find an approximate solution in feasible time. In this paper, we present Eager Map, a new algorithm to perform communication-based mapping that is based on a greedy grouping strategy applied hierarchically. Experimental evaluation indicates that the execution time of our algorithm is 10 times faster than the state-of-the-art, and presents higher performance improvements. Due to its low execution time and high stability, Eager Map is also suitable for online task mapping, where tasks are migrated during execution. Eduardo Henrique Molina da Cruz, Matthias Diener, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
PDP | 3 |
| 2015 | Understanding the Effect of Multiple Factors on a Parallel File System's PerformanceabstractThis research work presents an investigation of the impact of a wide range of factors on the performance of parallel file systems (PFSs). It is the result of an extensive test campaign with three distinct computing platforms and value variations for eleven factors that advance the understanding of PFSs' behaviour under different conditions. Our main contributions are the characterization of effects not fully explained or misunderstood in the literature previously. First, we demonstrate that no significant performance variation (≈ 6%) is observed when choosing one among four TCP congestion-avoidance algorithms. Second, we detail the effect of the page cache of I/O nodes on a PFS's write throughput and how it relates to other factors. Eduardo Camilo Inacio, Laércio Lima Pilla, Mario A. R. Dantas |
WETICE | 2 |
| 2015 | Characterizing communication and page usage of parallel applications for thread and data mapping
Matthias Diener, Eduardo Henrique Molina da Cruz, Laércio Lima Pilla, Fabrice Dupros, Philippe Olivier Alexandre Navaux |
Perform. Evaluation | 3 |
| 2014 | Radiation Sensitivity of High Performance Computing Applications on Kepler-Based GPGPUsabstractIn this paper we assess and discuss the radiation sensitivity of a set of HPC applications executed on NVIDIA K20 GPGPUs. The occurrence of both radiation-induced silent data corruption and functional interruption will be experimentally addressed for Hotspot, LavaMD, and Matrix Transponse. Each of the tested codes requires a proper computational power and elaborates a different amount of data. Both these characteristics play a significant role in the application radiations sensitivity. Additionally, an evaluation of the error rate at sea level will be provided for all the tested codes. Daniel Oliveira 0002, Caio B. Lunardi, Laércio Lima Pilla, Paolo Rech, Philippe Olivier Alexandre Navaux, Luigi Carro |
DSN | 3 |
| 2014 | Impact of GPUs Parallelism Management on Safety-Critical and HPC Applications ReliabilityabstractGraphics Processing Units (GPUs) offer high computational power but require high scheduling strain to manage parallel processes, which increases the GPU cross section. The results of extensive neutron radiation experiments performed on NVIDIA GPUs confirm this hypothesis. Reducing the application Degree Of Parallelism (DOP) reduces the scheduling strain but also modifies the GPU parallelism management, including memory latency, thread registers number, and the processors occupancy, which influence the sensitivity of the parallel application. An analysis on the overall GPU radiation sensitivity dependence on the code DOP is provided and the most reliable configuration is experimentally detected. Finally, modifying the parallel management affects the GPU cross section but also the code execution time and, thus, the exposure to radiation required to complete computation. The Mean Workload and Executions Between Failures metrics are introduced to evaluate the workload or the number of executions computed correctly by the GPU on a realistic application. Paolo Rech, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Luigi Carro |
DSN | 2 |
| 2014 | Saving energy by exploiting residual imbalances on iterative applicationsabstractThe power consumption of High Performance Computing (HPC) systems is an increasing concern as large-scale systems grow in size and, consequently, consume more energy. In response to this challenge, we propose two variants of a new energy-aware load balancer that aim at reducing the energy consumption of parallel platforms running imbalanced scientific applications without degrading their performance. Our research combines dynamic load balancing with DVFS techniques in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experimental results with benchmarks and a real-world application presented energy savings of up to 32% with our fine-grained variant that performs per-core DVFS, and of up to 34% with our coarsegrained variant that performs per-chip DVFS. Edson L. Padoin, Márcio Castro 0001, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
HiPC | 3 |
| 2014 | Improving the Performance of Seismic Wave Simulations with Dynamic Load BalancingabstractSeismic wave models provide a way to study the consequences of future earthquakes. When modeling a restricted region, these models require a boundary condition to absorb the energy that goes out of the simulated domain. To parallelize these models, the domain is decomposed into a grid of smaller subdomains which are mapped to different tasks. Due to the boundary condition, this division gives rise to load imbalance between the tasks that simulate border regions and those assigned center subdomains. To deal with this imbalance, and therefore improve the simulation's performance, we propose the use of dynamic load balancing. To evaluate our solution, we ported a seismic wave simulator to Adaptive MPI to profit from its load balancing framework. By using dynamic load balancers, we improved the performance of the application by 23.85% when compared to the original MPI implementation. We also show that load balancers are able to adapt to the variation of load imbalance during the application's execution. Rafael Keller Tesser, Laércio Lima Pilla, Fabrice Dupros, Philippe Olivier Alexandre Navaux, Jean-François Méhaut, Celso L. Mendes |
PDP | 2 |
| 2014 | Optimizing Memory Locality Using a Locality-Aware Page TableabstractOne of the main challenges for modern parallel shared-memory architectures are accesses to main memory. In current systems, the performance and energy efficiency of memory accesses depend on their locality: accesses to remote caches and NUMA nodes are more expensive than accesses to local ones. Increasing the locality requires knowledge about how the threads of a parallel application access memory pages. With this information, pages can be migrated to the NUMA nodes that access them (data mapping), as well as threads that access the same pages can be migrated to the same node such that locality can be improved even further (thread mapping). In this paper, we propose LAPT, a mechanism to store the memory access pattern of parallel applications in the page table, which is updated by the hardware during TLB misses. This information is used by the operating system to perform an optimized thread and data mapping during the execution of the parallel application. In contrast to previous work, LAPT does not require any previous information about the behavior of the applications, or changes to the application or runtime libraries. Extensive experiments with the NAS Parallel Benchmarks (NPB) and PARSEC showed performance and energy efficiency improvements of up to 19.2% and 15.7%, respectively, (6.7% and 5.3% on average). Eduardo Henrique Molina da Cruz, Matthias Diener, Marco A. Z. Alves, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 4 |
| 2014 | A topology-aware load balancing algorithm for clustered hierarchical multi-core machines
Laércio Lima Pilla, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Future Gener. Comput. Syst. | 1 |
| 2012 | Asymptotically Optimal Load Balancing for Hierarchical Multi-Core SystemsabstractCurrent multi-core machines feature a complex and hierarchical core topology, multiple levels of cache and memory subsystem with NUMA design. Although this design provides high processing power to parallel machines, it comes with the cost of asymmetric memory access latencies. Depending on the parallel application communication patterns, this asymmetry may reduce the overall performance of the system. Therefore, to achieve scalable performance in this environment, it becomes crucial to exploit the machine architecture while taking into account the application communication patterns. In this paper, we introduce a topology-aware load balancing algorithm named HWTOPOLB. It combines the machine topology characteristics with the communication patterns of the application to equalize the application load on the available cores while reducing latencies. We also present the proof that the algorithm is asymptotically optimal (Theorem 1). We have implemented our load balancing algorithm using the CHARM++ Parallel System and analyzed its performance using three different benchmarks. Our experimental results show that the HWTOPOLB can achieve average performance improvements of 24% when compared to existing load balancing strategies on three different multi-core machines. Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Jean-François Méhaut |
ICPADS | 1 |
| 2012 | A Hierarchical Approach for Load Balancing on Parallel Multi-core SystemsabstractMulti-core compute nodes with non-uniform memory access (NUMA) are now a common architecture in the assembly of large-scale parallel machines. On these machines, in addition to the network communication costs, the memory access costs within a compute node are also asymmetric. Ignoring this can lead to an increase in the data movement costs. Therefore, to fully exploit the potential of these nodes and reduce data access costs, it becomes crucial to have a complete view of the machine topology (i.e. the compute node topology and the interconnection network among the nodes). Furthermore, the parallel application behavior has an important role in determining how to utilize the machine efficiently. In this paper, we propose a hierarchical load balancing approach to improve the performance of applications on parallel multi-core systems. We introduce NucoLB, a topology-aware load balancer that focuses on redistributing work while reducing communication costs among and within compute nodes. NucoLB takes the asymmetric memory access costs present on NUMA multi-core compute nodes, the interconnection network overheads, and the application communication patterns into account in its balancing decisions. We have implemented NucoLB using the Charm++ parallel runtime system and evaluated its performance. Results show that our load balancer improves performance up to 20% when compared to state-of-the-art load balancers on three different NUMA parallel machines. Laércio Lima Pilla, Christiane Pousa Ribeiro, Daniel Cordeiro, Abhinav Bhatele, Philippe Olivier Alexandre Navaux, François Broquedis, Jean-François Méhaut, Laxmikant V. Kalé |
ICPP | 1 |
| 2011 | Combining Multiple Metrics to Control BSP Process Rescheduling in Response to Resource and Application DynamicsabstractThis article discusses MigBSP: a rescheduling model that acts on Bulk Synchronous Parallel applications running over computational Grids. It combines the metrics Computation, Communication and Memory to make migration decisions. MigBSP also offers efficient adaptations to reduce its overhead. Additionally, MigBSP is infrastructure and application independent and tries to handle dynamicity on both levels. MigBSP's results show application performance improvements of up to 16% on dynamic environments while maintaining a small overhead when migrations do not take place. Rodrigo da Rosa Righi, Lucas Graebin, Rafael Bohrer Ávila, Philippe Olivier Alexandre Navaux, Laércio Lima Pilla |
ICPADS | 5 |
| 2011 | Improving Performance on Atmospheric Models through a Hybrid OpenMP/MPI ImplementationabstractThis work shows how a Hybrid MPI/OpenMP implementation can improve the performance of the Ocean-Land-Atmosphere Model (OLAM) on a multi-core cluster environment, which is a typical HPC many small files workload application. Previous experiments have shown that the scalability of this application on clusters is limited by the performance of the output operations. We show that the Hybrid MPI/OpenMP version of OLAM decreases the number of output files, resulting in better performance for I/O operations. We also observe that the MPI version of OLAM performs better for unbalanced workloads and that further parallel optimizations should be included on the hybrid version in order to improve the parallel execution time of OLAM. Carla Osthoff, Pablo Javier Grunmann, Francieli Zanon Boito, Rodrigo Kassick, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Claudio Schepke, Jairo Panetta, Nicolas Maillard, Pedro Leite da Silva Dias, Robert L. Walko |
ISPA | 5 |
| 2010 | Supporting performance and adaptivity on BSP process reschedulingabstractIn this paper we will describe a model for BSP (Bulk Synchronous Parallel) process rescheduling called MigBSP. Considering the scope of BSP applications, its differential approach is the combination of three metrics - Memory, Computation and Communication - in order to measure the Potential of Migration of each BSP process. In this context, this paper addresses both the efficiency and the adaptivity perspectives of this model over our multi-cluster machine. Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß |
ISCC | 2 |
| 2009 | MigBSP: A Novel Migration Model for Bulk-Synchronous Parallel Processes ReschedulingabstractWe have developed a model called MigBSP that controls processes rescheduling in BSP (bulk synchronous parallel)applications. A BSP application is composed by one or more supersteps, each one containing both computation and communication phases followed by a synchronization barrier. Since the barrier waits for the slowest process, MigBSPpsilas final idea is to adjust the processes location in order to reduce the superstepspsila times. Considering the scope of the BSP model, the novel ideas of MigBSPare: (i) combination of three metrics - memory, computation and communication - to measure the potential of migration of each BSP process; (ii) use of both computation and communication patterns to control processespsila regularity;(iii) adaptation regarding the periodicity to launch the processes rescheduling. This paper describes MigBSP and presents some experimental results and related work. Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux, Hans-Ulrich Heiß |
HPCC | 2 |
| 2008 | Controlling Processes Reassignment in BSP ApplicationsabstractWe have developed a model for dynamic process scheduling in heterogeneous and non-dedicated environments. This model acts over a BSP (Bulk Synchronous Parallel) application, applying runtime processes reassignment to new processors. A BSP application is divided in one or more supersteps, each one containing both computation and communication phases followed by a barrier synchronization. In this context, the developed model combines three metrics - Memory, Computation and Communication - in order to measure the potential of migration of each BSP process. The final idea is to offer a mathematical formalism involving these metrics and to decide the following questions about the process migration: When? Where? Which? This paper presents the algorithms of our model, the parallel machine organization, some experimental results and related work. Rodrigo da Rosa Righi, Laércio Lima Pilla, Alexandre Carissimi, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 2 |