EDBT 2026 Demo / reviewers in the wild / expert
Lucas Mello Schnorr
dblp:57/3743
· DBLP profile ↗
31ranked-venue papers
8as first author
8since 2021 · last 2023
0000-0003-4828-9942ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 6 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Summarizing task-based applications behavior over many nodes through progression clusteringabstractVisualization strategies are a valuable tool in the performance evaluation of HPC applications. Although the traditional Gantt charts are a widespread and enlightening strategy, it presents scalability problems and may misguide the analysis by focusing on resource utilization alone. This paper proposes an overview strategy to indicate nodes of interest for further investigation with classical visualizations like Gantt charts. For this, it uses a progression metric that captures work done per node inferred from the task-based structure, a time-step clustering of those metrics to decrease redundant information, and a more scalable visualization technique. We demonstrate with six scenarios and two applications that such a strategy can indicate problematic nodes more straightforwardly while using the same visualization space. Also, we provide examples where it correctly captures application work progression, showing application problems earlier and as an easy way to compare nodes. At the same time that traditional methods are misleading. Lucas Leandro Nesi, Vinícius Garcia Pinto, Lucas Mello Schnorr, Arnaud Legrand |
PDP | 3 |
| 2023 | Performance Modeling of MARE2DEM's Adaptive Mesh Refinement for Makespan EstimationabstractAdaptive Mesh Refinement (AMR) is a widely known technique to adapt the accuracy of a solution in critical areas of the problem domain instead of using regular or irregular but static meshes. The MARE2DEM is a parallel application that employs the AMR technique to model 2D electromagnetics in oil and gas exploration. The modeling consists in iteratively applying a data inversion based on a set of measurements collected and registered by a survey on an area of interest. The parallelism of the MARE2DEM works by dividing the workload into a set of refinement groups that represent overlapping areas of the problem domain. Each refinement group can be computed independently of the others by a set of workers, carrying out the AMR in the meshes when necessary. The shape and compute performance of the refinement group depend directly of a set of user-defined parameters. In this article, we provide a method to estimate the MARE2DEM performance for all possible values that can be used in the influencing parameters of the application for a given case study. Our relatively cheap method enables the geologist to configure MARE2DEM correctly and extract the best performance for a given cluster configuration. We detail how the method works and evaluate its effectiveness with success, pinpointing the best values for the creating refinement groups using a real case study from the Marlim field on the coast of Rio de Janeiro, Brazil. Although we demonstrate our evaluation with this scenario, our method works for any input of MARE2DEM. Bruno da Silva Alves, Lucas Mello Schnorr |
SBAC-PAD | 2 |
| 2023 | Asynchronous multi-phase task-based applications: Employing different nodes to design better distributions
Lucas Leandro Nesi, Arnaud Legrand, Lucas Mello Schnorr |
Future Gener. Comput. Syst. | 3 |
| 2023 | Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos DumontabstractHigh-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues. Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
J. Parallel Distributed Comput. | 4 |
| 2022 | Multi-Phase Task-Based HPC Applications: Quickly Learning how to Run FastabstractParallel applications performance strongly depends on the number of resources. Although adding new nodes usually reduces execution time, excessive amounts are often detrimental as they incur substantial communication overhead, which is difficult to anticipate. Characteristics like network contention, data distribution methods, synchronizations, and how communications and computations overlap generally impact the performance. Finding the correct number of resources can thus be particularly tricky for multi-phase applications as each phase may have very different needs, and the popularization of hybrid ($C$PU+GPU) machines and heterogeneous partitions makes it even more difficult. In this paper, we study and propose, in the context of a task-based GeoStatistic application, strategies for the application to actively learn and adapt to the best set of heterogeneous nodes it has access to. We propose strategies that use the Gaussian Process method with trends, bound mechanisms for reducing the search space, and heterogeneous behavior modeling. We compare these methods with traditional exploration strategies in 16 different machines scenarios. In the end, the proposed strategies are able to gain up to ≈51% compared to the standard case of using all the nodes while having low overhead. Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
IPDPS | 2 |
| 2022 | Performance analysis of task-based multi-frontal sparse linear solvers: Structure matters
Marcelo Miletto, Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
Future Gener. Comput. Syst. | 3 |
| 2021 | Exploiting system level heterogeneity to improve the performance of a GeoStatistics multi-phase task-based applicationabstractHeterogeneity is part of HPC infrastructures, not only at the intra-node but at the system level. Applications with multiple phases with distinct resource necessities can take advantage of this inter-node heterogeneity to improve performance and reduce resource idleness. Such an application is ExaGeoStat, a task-based machine learning framework specifically designed for geostatistics data. This work presents strategies to efficiently distribute multi-phase applications in system-level heterogeneous resources. We both (1) improve application phase overlap by optimizing runtime and scheduling decisions and (2) compute the optimal distribution for all the phases using a linear program leveraging node heterogeneity while limiting communication overhead. The performance gains of our phase overlap improvements are between 36% and 50% compared to the original base synchronous and homogeneous execution. We show that by adding some slow nodes to a homogeneous set of fast nodes, we can improve the performance by another 25% compared to a standard block-cyclic distribution, thereby harnessing any machine. Lucas Leandro Nesi, Arnaud Legrand, Lucas Mello Schnorr |
ICPP | 3 |
| 2021 | HPC Data Storage at a Glance: The Santos Dumont ExperienceabstractHigh-Performance Computing (HPC) platforms are used to solve the most diverse scientific problems in research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific software, which have different requirements. These requirements include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle a mixed workload scenario when storing data from the applications. Knowledge of the application set and its performance running in a supercomputer is needed to understand the storage system's usage, pinpoint possible bottlenecks, and guide optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, where we were able to identify inefficient usage and problematic factors of performance. Andre Ramos Carneiro, Jean Luca Bez, Carla Osthoff, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 4 |
| 2020 | Communication-Aware Load Balancing of the LU Factorization over Heterogeneous ClustersabstractSupercomputers are designed to be as homogeneous as possible but it is common that a few nodes exhibit variable performance capabilities due to processor manufacturing. It is also common to find partitions equipped with different types of accelerators. Data distribution over heterogeneous nodes is very challenging but essential to exploit all resources efficiently. In this article, we build upon task-based runtimes' flexibility of managing data to study the interplay between static communication-aware data distribution strategies and dynamic scheduling of the linear algebra LU factorization over heterogeneous sets of hybrid nodes. We propose two techniques derived from the state-of-the-art 1D×1D data distributions. First, to use fewer computing nodes towards the end to better match performance bounds and save computing power. Second, to carefully move a few blocks between nodes to optimize even further the load balancing among nodes. We also demonstrate how 1D×1D data distributions, tailored for heterogeneous nodes, can scale better with homogeneous clusters than classical block-cyclic distributions. Validation is carried out both in real and in simulated environments under homogeneous and heterogeneous platforms, demonstrating compelling performance improvements. Lucas Leandro Nesi, Lucas Mello Schnorr, Arnaud Legrand |
ICPADS | 2 |
| 2020 | Task-based parallel strategies for computational fluid dynamic application in heterogeneous CPU/GPU resourcesabstractSummary Parallel applications executing in contemporary heterogeneous clusters are complex to code and optimize. The task‐based programming model is an alternative to handle the coding complexity. This model consists of splitting the problem domain into tasks with dependencies through a directed acyclic graph, and submit the set of tasks to a runtime scheduler that maps each task dynamically to resources. We consider that computational fluid dynamics applications are typical in scientific computing but not enough exploited by designs that employ the task‐based programming model. This article presents task‐based parallel strategies for a simple CFD application that targets heterogeneous multi‐CPU/multi‐GPU computing resources. We design, develop, evaluate, and compare the performance of three parallel strategies (naive, ghost‐cells, and arrow) of a task‐based heterogeneous (CPU and GPU) application that simulates the flow of an incompressible Newtonian fluid with constant viscosity. All implementations rely on the StarPU runtime, and we use the StarVZ toolkit to conduct comprehensive performance analysis. Results indicate that the ghost cell strategy provides the best speedup (77×) considering the simulation time when the GPU resources still have available memory. However, the arrow strategy achieves better results when the simulation data increases. Lucas Leandro Nesi, Matheus S. Serpa, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Visual Performance Analysis of Memory Behavior in a Task-Based Runtime on Hybrid PlatformsabstractProgramming parallel applications for heterogeneous HPC platforms is much more straightforward when using the task-based programming paradigm. The simplicity exists because a runtime takes care of many activities usually carried out by the application developer, such as task mapping, load balancing, and memory management operations. In this paper, we present a visualization-based performance analysis methodology to investigate the CPU-GPU-Disk memory management of the StarPU runtime, a popular task-based middleware for HPC applications. We detail the design of novel graphical strategies that were fundamental to recognize performance problems in four study cases. We first identify poor management of data handles when GPU memory is saturated, leading to low application performance. Our experiments using the dense tiled-based Cholesky factorization show that our fix leads to performance gains of 66% and better scalability for larger input sizes. In the other three cases, we study scenarios where the main memory is insufficient to store all the application's data, forcing the runtime to store data out-of-core. Using our methodology, we pin-point different behavior among schedulers and how we have identified a crucial problem in the application code regarding initial block placement, which leads to poor performance. Lucas Leandro Nesi, Samuel Thibault, Luka Stanisic, Lucas Mello Schnorr |
CCGRID | 4 |
| 2019 | Performance modeling of a geophysics application to accelerate over-decomposition parameter tuning through simulationabstractSummary Finite‐difference methods are commonplace in High Performance Computing applications. Despite their apparent regularity, they often exhibit load imbalance that damages their efficiency. We characterize the spatial and temporal load imbalance of Ondes3D, a typical finite‐differences application dedicated to earthquake modeling. Our analysis reveals imbalance originating from the structure of the input data, and from low‐level CPU optimizations. Ondes3D was successfully ported to AMPI/CHARM++ using over‐decomposition and MPI process migration techniques to dynamically rebalance the load. However, this approach requires careful selection of the over‐decomposition level, the load balancing algorithm, and its activation frequency. These choices are usually tied to application structure and platform characteristics. In this article, we propose a workflow that leverages the capabilities of SimGrid to conduct such study at low experimental cost. We rely on a combination of emulation, simulation, and application modeling that requires minimal code modification and manages to capture both spatial and temporal load imbalance to faithfully predict the performance of dynamic load balancing. We evaluate the quality of our simulation by comparing simulation results with the outcome of real executions and demonstrate how this approach can be used to quickly find the optimal load balancing configuration for a given application/hardware configuration. Rafael Keller Tesser, Lucas Mello Schnorr, Arnaud Legrand, Franz C. Heinrich, Fabrice Dupros, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | A visual performance analysis framework for task-based parallel applications running on hybrid clustersabstractSummary Programming paradigms in High‐Performance Computing have been shifting toward task‐based models that are capable of adapting readily to heterogeneous and scalable supercomputers. The performance of task‐based application heavily depends on the runtime scheduling heuristics and on its ability to exploit computing and communication resources. Unfortunately, the traditional performance analysis strategies are unfit to fully understand task‐based runtime systems and applications: they expect a regular behavior with communication and computation phases, while task‐based applications demonstrate no clear phases. Moreover, the finer granularity of task‐based applications typically induces a stochastic behavior that leads to irregular structures that are difficult to analyze. Furthermore, the combination of application structure, scheduler, and hardware information is generally essential to understand performance issues. This paper presents a flexible framework that enables one to combine several sources of information and to create custom visualization panels allowing to understand and pinpoint performance problems incurred by bad scheduling decisions in task‐based applications. Three case‐studies using StarPU‐MPI, a task‐based multi‐node runtime system, are detailed to show how our framework can be used to study the performance of the well‐known Cholesky factorization. Performance improvements include a better task partitioning among the multi‐(GPU, core) to get closer to theoretical lower bounds, improved MPI pipelining in multi‐(node, core, GPU) to reduce the slow start, and changes in the runtime system to increase MPI bandwidth, with gains of up to 13% in the total makespan. Vinícius Garcia Pinto, Lucas Mello Schnorr, Luka Stanisic, Arnaud Legrand, Samuel Thibault, Vincent Danjean |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Using Simulation to Evaluate and Tune the Performance of Dynamic Load Balancing of an Over-Decomposed Geophysics Application
Rafael Keller Tesser, Lucas Mello Schnorr, Arnaud Legrand, Fabrice Dupros, Philippe Olivier Alexandre Navaux |
Euro-Par | 2 |
| 2017 | TWINS: Server Access Coordination in the I/O Forwarding LayerabstractThis paper presents a study of I/O scheduling techniques applied to the I/O forwarding layer. In high-performance computing environments, applications rely on parallel file systems (PFS) to obtain good I/O performance even when handling large amounts of data. To alleviate the concurrency caused by thousands of nodes accessing a significantly smaller number of PFS servers, intermediate I/O nodes are typically applied between processing nodes and the file system. Each intermediate node forwards requests from multiple clients to the system, a setup which gives this component the opportunity to perform optimizations like I/O scheduling. We evaluate scheduling techniques that improve spatiality and request size of the access patterns. We show they are only partially effective because the access pattern is not the main factor for read performance in the I/O forwarding layer. A new scheduling algorithm, TWINS, is presented to coordinate the access of intermediate I/O nodes to the data servers. Our proposal decreases concurrency at the data servers, a factor previously proven to negatively affect performance. The proposed algorithm is able to improve read performance from shared files by up to 28% over other scheduling algorithms and by up to 50% over not forwarding I/O. Jean Luca Bez, Francieli Zanon Boito, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
PDP | 3 |
| 2017 | Performance and energy efficiency analysis of HPC physics simulation applications in a cluster of ARM processorsabstractSummary We analyze the feasibility and energy efficiency of using an unconventional cluster of low‐power Advanced RISC Machines processors to execute two scientific parallel applications. For this purpose, we have selected two applications that present high computational and communication cost: the Ondes3D that simulates geophysical events, and the all‐pairs N‐Body that simulates astrophysical events. We compare and discuss the impact of different compilation directives and processor frequency and how they interfere in Time‐to‐Solution and Energy‐to‐Solution. Our results demonstrate that by correctly tuning the application at compile time, for the Advanced RISC Machines architecture, we can considerably reduce the execution time and the energy spent by computing simulations. Furthermore, we observe reductions of up to 54.14% in Time‐to‐Solution and gains of up to 53.65% in Energy‐to‐Solution with two cores. Additionally, we consider the impact of two processor frequency governors on these metrics. Results indicate that the powersave governor presents a smaller instantaneous power consumption. However, it spends more time executing tasks, increasing the energy needed to achieve the solution. Finally, we correlate the energy consumption with the execution time in the experimental results using Pareto. These findings suggest that it is possible to explore low‐powered clusters for high‐performance computing applications by tuning application and hardware configuration to achieve energy efficiency. Copyright © 2016 John Wiley & Sons, Ltd. Jean Luca Bez, Eliezer E. Bernart, Fernando Santos 0001, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | PhenoVis - A tool for visual phenological analysis of digital camera images using chronological percentage mapsabstractPhenoVis is framework for the visual phenological analysis of forest ecosystems. It contains the chronological percentage maps (CPM), a novel representation that is capable of discovering additional patterns by encoding percentage distributions of the data. Two types of masks are used in PhenoVis: a community mask , which considers all plant species in the image; and a species mask , associated with a given plant species. Among the several images taken at different times of the day, the image taken at noon is preferred for the analysis because it minimizes shadow effects. Therefore, only one image per day is used. The analysis considers the chromatic co- efficients associated with each pixel in the image. In PhenoVis we associate different colors with each bucket of the percentage histogram. The histogram granularity defines the size of a given bucket of the percentage distribution. The number of buckets is given by the number of colors available, and the range of the distribution is given by the IOI. The percentage map of a single input image consists of a normalized stacked bar chart. The chronological percentage map consists of a sequence of percentage maps stacked in chronological order, from top to bottom (portrait) or left to right (landscape). Roger A. Leite, Lucas Mello Schnorr, Jurandy Almeida, Bruna Alberton, Leonor Patricia C. Morellato, Ricardo da Silva Torres, João Luiz Dihl Comba |
Inf. Sci. | 2 |
| 2014 | A spatiotemporal data aggregation technique for performance analysis of large-scale execution tracesabstractAnalysts commonly use execution traces collected at runtime to understand the behavior of an application running on distributed and parallel systems. These traces are inspected post mortem using various visualization techniques that, however, do not scale properly for a large number of events. This issue, mainly due to human perception limitations, is also the result of bounded screen resolutions preventing the proper drawing of many graphical objects. This paper proposes a new visualization technique overcoming such limitations by providing a concise overview of the trace behavior as the result of a spatiotemporal data aggregation process. The experimental results show that this approach can help the quick and accurate detection of anomalies in traces containing up to two hundred million events. Damien Dosimont, Robin Lamarche-Perrin, Lucas Mello Schnorr, Guillaume Huard, Jean-Marc Vincent |
CLUSTER | 3 |
| 2014 | Evaluating trace aggregation for performance visualization of large distributed systemsabstractPerformance analysis through visualization techniques usually suffers semantic limitations due to the size of parallel applications. Most performance visualization tools rely on data aggregation to work at scale, without any attempt to evaluate the loss of information caused by such aggregations. This paper proposes a technique to evaluate the quality of aggregated representations - using measures from information theory - and to optimize such measures in order to build consistent multiresolution representations of large execution traces. Robin Lamarche-Perrin, Lucas Mello Schnorr, Jean-Marc Vincent, Yves Demazeau |
ISPASS | 2 |
| 2014 | Best of SBAC-PAD 2012
Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
Parallel Comput. | 1 |
| 2013 | Interactive analysis of large distributed systems with scalable topology-based visualizationabstractThe performance of parallel and distributed applications is highly dependent on the characteristics of the execution environment. In such environments, the network topology and characteristics tell how fast data can be transmitted and placed in the resources. These are key phenomena to understand the behavior of such applications and possibly improve it. Unfortunately few visualization available to the analyst are capable of accounting for such phenomena. In this paper, we propose an interactive topology-based visualization technique based on data aggregation that enables to correlate network characteristics, such as bandwidth and topology, with application performance traces. We show that such kind of visualization enables to explore and understand non trivial behavior that are impossible to grasp with classical visualization techniques. We also show that the combination of multi-scale aggregation and dynamic graph layout allows our visualization technique to scale seamlessly to large distributed systems. These results are validated through a detailed analysis of a high performance computing scenario and of a grid computing scenario. Lucas Mello Schnorr, Arnaud Legrand, Jean-Marc Vincent |
ISPASS | 1 |
| 2012 | Detection and analysis of resource usage anomalies in large distributed systems through multi-scale visualizationabstractSUMMARY Understanding the behavior of large scale distributed systems is generally extremely difficult as it requires to observe a very large number of components over very large time. Most analysis tools for distributed systems gather basic information such as individual processor or network utilization. Although scalable because of the data reduction techniques applied before the analysis, these tools are often insufficient to detect or fully understand anomalies in the dynamic behavior of resource utilization and their influence on the applications performance. In this paper, we propose a methodology for detecting resource usage anomalies in large scale distributed systems. The methodology relies on four functionalities: characterized trace collection, multi‐scale data aggregation, specifically tailored user interaction techniques, and visualization techniques. We show the efficiency of this approach through the analysis of simulations of the volunteer computing Berkeley Open Infrastructure for Network Computing architecture. Three scenarios are analyzed in this paper: analysis of the resource sharing mechanism, resource usage considering response time instead of throughput, and the evaluation of input file size on Berkeley Open Infrastructure for Network Computing architecture. The results show that our methodology enables to easily identify resource usage anomalies, such as unfair resource sharing, contention, moving network bottlenecks, and harmful short‐term resource sharing. Copyright © 2011 John Wiley & Sons, Ltd. Lucas Mello Schnorr, Arnaud Legrand, Jean-Marc Vincent |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | A hierarchical aggregation model to achieve visualization scalability in the analysis of parallel applications
Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux |
Parallel Comput. | 1 |
| 2010 | Impact of Parallel Workloads on NoC Architecture DesignabstractDue to the multi-core processors, the importance of parallel workloads has increased considerably. However, many-core chips demand new interconnection strategies, since traditional crossbars or buses, common for current multi-core processors, have problems related to wires and scalability. For this reason, Networks-on-Chip (NoCs) have been developed in order to support the performance and parallelism focused on several workloads. Although a Network-on-Chip is a good option, most designs consist of a large number of routers. These routers are responsible for forwarding packets, and consequently, for supporting message-passing workloads. In this context, the NoC performance is a problem. Therefore, the main goal of this paper is to evaluate the impact of well-known parallel workloads on NoC architecture design. In order to achieve high performance, the results point out to parallel workloads with small packets and cluster-based NoCs with circuit switching and adaptable topologies. Henrique Cota de Freitas, Lucas Mello Schnorr, Marco A. Z. Alves, Philippe Olivier Alexandre Navaux |
PDP | 2 |
| 2010 | Triva: Interactive 3D visualization for performance analysis of parallel applications
Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux |
Future Gener. Comput. Syst. | 1 |
| 2009 | Towards Visualization Scalability through Time Intervals and Hierarchical Organization of Monitoring DataabstractHighly distributed systems such as grids are used today to the execution of large-scale parallel applications. The behavior analysis of these applications is not trivial. The complexity appears because of the event correlation among processes, external influences like time-sharing mechanisms and saturation of network links, and also the amount of data that registers the application behavior. Almost all visualization tools to analysis of parallel applications offer a space-time representation of the application behavior. This paper presents a novel technique that combines traces from grid applications with a treemap visualization of the data. With this combination, we dynamically create an annotated hierarchical structure that represents the application behavior for the selected time interval. The experiments in the grid show that we can readily use our technique to the analysis of large-scale parallel applications with thousands of processes. Lucas Mello Schnorr, Guillaume Huard, Philippe Olivier Alexandre Navaux |
CCGRID | 1 |
| 2009 | Performance Evaluation of NoC Architectures for Parallel WorkloadsabstractNetwork-on-Chip is the state-of-the-art approach to interconnect many processing cores in the next generation of general-purpose processors. In this context, the problem is to choose NoC architectures capable of achieving high performance for parallel programs. Therefore, the main goal of this paper is to evaluate the performance of three NoC architectures using well-known parallel workloads. Henrique Cota de Freitas, Marco A. Z. Alves, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux |
NOCS | 3 |
| 2009 | Visual Mapping of Program Components to Resources Representation: A 3D Analysis of Grid Parallel ApplicationsabstractHighly distributed systems such as Grids are usually interconnected by a hierarchical organization of different types of network. This strong network hierarchy influences directly the behavior of parallel applications. In order to obtain a good understanding of parallel application's behavior, the performance analysis must take into account a correspondence between application and network characteristics. This paper presents a novel way to analyze parallel applications, by using a three dimensional visualization and a technique to visually map the application's components to the used resources. This mapping technique is able to handle different types of resources description: different system logical organizations and the description of the network or system interconnection in different levels. The technique, implemented in our prototype Triva, is evaluated through a series of visual representations of the monitoring data obtained through real executions of parallel applications in a grid. The resulting visualizations enable an alternative and powerful way to view and understand various aspects of application behavior together with the network topology. Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Guillaume Huard |
SBAC-PAD | 1 |
| 2006 | ICE: A Service Oriented Approach to Uniform the Access and Management of Cluster Environments
Clarissa Cassales Marquezan, Rodrigo da Rosa Righi, Lucas Mello Schnorr, Alexandre Carissimi, Nicolas Maillard, Philippe Olivier Alexandre Navaux |
CCGRID | 3 |
| 2006 | DIMVisual: Data Integration Model for Visualization of Parallel Programs BehaviorabstractThe development of high performance parallel applications for clusters is considered a complex task. This can happen because the influence of the execution environment and the non-deterministic natural behavior of this kind of applications. In such development, the programmer uses application traces and cluster monitoring tools to register the events of the application and the underlying execution environment. Generally, the analysis of the information from each source is made independently, making the correlation of events from the application with events from the execution environment difficult. This paper presents DIMVisual, a Data Integration Model which addresses this problem by integrating information from different sources and providing a unified visualization. An implementation of this model is also presented, using as data sources traces from MPI and DECK applications, events from Ganglia and Performance Co-Pilot cluster monitoring tools and operating system context switches. The results show the information gathered by these data sources integrated and visualized together in the generic visualization tool Paj´e, allowing the programmer a more complete view of his application behavior. Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Benhur de Oliveira Stein |
CCGRID | 1 |
| 2003 | JRastro: A Trace Agent for Debugging Multithreaded and Distributed Java ProgramsabstractProgram tracing is one of the most used techniques to debug parallel and distributed programs. In this technique, events are recorded in trace files during the execution of the program for post mortem visualization of its behavior. We describe JRastro, a trace agent capable of tracing Java programs. The agent was designed to cover three key features: to be transparent to the application developer, to use unmodified Java virtual machines and to observe remote method invocations. By integrating these three features, JRastro differentiates itself from similar tools. Unfortunately, for a complete and clean implementation of RMI visualization, additional support on the Java monitoring system is needed. Gabriela Jacques-Silva, Lucas Mello Schnorr, Benhur de Oliveira Stein |
SBAC-PAD | 2 |