EDBT 2026 Demo / reviewers in the wild / expert
Lucas Leandro Nesi
dblp:223/1149
· DBLP profile ↗
10ranked-venue papers
9as first author
5since 2021 · last 2023
0000-0001-8874-1839ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 7 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Summarizing task-based applications behavior over many nodes through progression clusteringabstractVisualization strategies are a valuable tool in the performance evaluation of HPC applications. Although the traditional Gantt charts are a widespread and enlightening strategy, it presents scalability problems and may misguide the analysis by focusing on resource utilization alone. This paper proposes an overview strategy to indicate nodes of interest for further investigation with classical visualizations like Gantt charts. For this, it uses a progression metric that captures work done per node inferred from the task-based structure, a time-step clustering of those metrics to decrease redundant information, and a more scalable visualization technique. We demonstrate with six scenarios and two applications that such a strategy can indicate problematic nodes more straightforwardly while using the same visualization space. Also, we provide examples where it correctly captures application work progression, showing application problems earlier and as an easy way to compare nodes. At the same time that traditional methods are misleading. Lucas Leandro Nesi, Vinícius Garcia Pinto, Lucas Mello Schnorr, Arnaud Legrand |
PDP | 1 |
| 2023 | Asynchronous multi-phase task-based applications: Employing different nodes to design better distributions
Lucas Leandro Nesi, Arnaud Legrand, Lucas Mello Schnorr |
Future Gener. Comput. Syst. | 1 |
| 2022 | Multi-Phase Task-Based HPC Applications: Quickly Learning how to Run FastabstractParallel applications performance strongly depends on the number of resources. Although adding new nodes usually reduces execution time, excessive amounts are often detrimental as they incur substantial communication overhead, which is difficult to anticipate. Characteristics like network contention, data distribution methods, synchronizations, and how communications and computations overlap generally impact the performance. Finding the correct number of resources can thus be particularly tricky for multi-phase applications as each phase may have very different needs, and the popularization of hybrid ($C$PU+GPU) machines and heterogeneous partitions makes it even more difficult. In this paper, we study and propose, in the context of a task-based GeoStatistic application, strategies for the application to actively learn and adapt to the best set of heterogeneous nodes it has access to. We propose strategies that use the Gaussian Process method with trends, bound mechanisms for reducing the search space, and heterogeneous behavior modeling. We compare these methods with traditional exploration strategies in 16 different machines scenarios. In the end, the proposed strategies are able to gain up to ≈51% compared to the standard case of using all the nodes while having low overhead. Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
IPDPS | 1 |
| 2022 | Performance analysis of task-based multi-frontal sparse linear solvers: Structure matters
Marcelo Miletto, Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
Future Gener. Comput. Syst. | 2 |
| 2021 | Exploiting system level heterogeneity to improve the performance of a GeoStatistics multi-phase task-based applicationabstractHeterogeneity is part of HPC infrastructures, not only at the intra-node but at the system level. Applications with multiple phases with distinct resource necessities can take advantage of this inter-node heterogeneity to improve performance and reduce resource idleness. Such an application is ExaGeoStat, a task-based machine learning framework specifically designed for geostatistics data. This work presents strategies to efficiently distribute multi-phase applications in system-level heterogeneous resources. We both (1) improve application phase overlap by optimizing runtime and scheduling decisions and (2) compute the optimal distribution for all the phases using a linear program leveraging node heterogeneity while limiting communication overhead. The performance gains of our phase overlap improvements are between 36% and 50% compared to the original base synchronous and homogeneous execution. We show that by adding some slow nodes to a homogeneous set of fast nodes, we can improve the performance by another 25% compared to a standard block-cyclic distribution, thereby harnessing any machine. Lucas Leandro Nesi, Arnaud Legrand, Lucas Mello Schnorr |
ICPP | 1 |
| 2020 | Communication-Aware Load Balancing of the LU Factorization over Heterogeneous ClustersabstractSupercomputers are designed to be as homogeneous as possible but it is common that a few nodes exhibit variable performance capabilities due to processor manufacturing. It is also common to find partitions equipped with different types of accelerators. Data distribution over heterogeneous nodes is very challenging but essential to exploit all resources efficiently. In this article, we build upon task-based runtimes' flexibility of managing data to study the interplay between static communication-aware data distribution strategies and dynamic scheduling of the linear algebra LU factorization over heterogeneous sets of hybrid nodes. We propose two techniques derived from the state-of-the-art 1D×1D data distributions. First, to use fewer computing nodes towards the end to better match performance bounds and save computing power. Second, to carefully move a few blocks between nodes to optimize even further the load balancing among nodes. We also demonstrate how 1D×1D data distributions, tailored for heterogeneous nodes, can scale better with homogeneous clusters than classical block-cyclic distributions. Validation is carried out both in real and in simulated environments under homogeneous and heterogeneous platforms, demonstrating compelling performance improvements. Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
ICPADS | 1 |
| 2020 | Task-based parallel strategies for computational fluid dynamic application in heterogeneous CPU/GPU resourcesabstractSummary Parallel applications executing in contemporary heterogeneous clusters are complex to code and optimize. The task‐based programming model is an alternative to handle the coding complexity. This model consists of splitting the problem domain into tasks with dependencies through a directed acyclic graph, and submit the set of tasks to a runtime scheduler that maps each task dynamically to resources. We consider that computational fluid dynamics applications are typical in scientific computing but not enough exploited by designs that employ the task‐based programming model. This article presents task‐based parallel strategies for a simple CFD application that targets heterogeneous multi‐CPU/multi‐GPU computing resources. We design, develop, evaluate, and compare the performance of three parallel strategies (naive, ghost‐cells, and arrow) of a task‐based heterogeneous (CPU and GPU) application that simulates the flow of an incompressible Newtonian fluid with constant viscosity. All implementations rely on the StarPU runtime, and we use the StarVZ toolkit to conduct comprehensive performance analysis. Results indicate that the ghost cell strategy provides the best speedup (77×) considering the simulation time when the GPU resources still have available memory. However, the arrow strategy achieves better results when the simulation data increases. Lucas Leandro Nesi, Matheus S. Serpa, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Visual Performance Analysis of Memory Behavior in a Task-Based Runtime on Hybrid PlatformsabstractProgramming parallel applications for heterogeneous HPC platforms is much more straightforward when using the task-based programming paradigm. The simplicity exists because a runtime takes care of many activities usually carried out by the application developer, such as task mapping, load balancing, and memory management operations. In this paper, we present a visualization-based performance analysis methodology to investigate the CPU-GPU-Disk memory management of the StarPU runtime, a popular task-based middleware for HPC applications. We detail the design of novel graphical strategies that were fundamental to recognize performance problems in four study cases. We first identify poor management of data handles when GPU memory is saturated, leading to low application performance. Our experiments using the dense tiled-based Cholesky factorization show that our fix leads to performance gains of 66% and better scalability for larger input sizes. In the other three cases, we study scenarios where the main memory is insufficient to store all the application's data, forcing the runtime to store data out-of-core. Using our methodology, we pin-point different behavior among schedulers and how we have identified a crucial problem in the application code regarding initial block placement, which leads to poor performance. Lucas Leandro Nesi, Samuel Thibault, Luka Stanisic, Lucas Mello Schnorr |
CCGRID | 1 |
| 2018 | GPU-Accelerated Algorithms for Allocating Virtual Infrastructure in Cloud Data CentersabstractAllocating IT resources to Virtual Infrastructures (VIs) (i.e.groups of VMs, virtual switches, and their network interconnections) is an NP-hard problem. Most allocation algorithms designed to run on CPUs face scalability issues when considering current cloud data centers comprising thousands of servers. This work offers and evaluates a set of allocation algorithms refactored for Graphic Processing Units (GPUs). Experimental results demonstrate their ability to handle three large-scale data center topologies. Lucas Leandro Nesi, Maurício Aronne Pillon, Marcos Dias de Assunção, Guilherme P. Koslovski |
CCGrid | 1 |
| 2018 | Tackling Virtual Infrastructure Allocation in Cloud Data Centers: a GPU-Accelerated Framework
Lucas Leandro Nesi, Maurício Aronne Pillon, Marcos Dias de Assunção, Charles Miers, Guilherme P. Koslovski |
CNSM | 1 |