EDBT 2026 Demo / reviewers in the wild / expert
Jean-François Méhaut
dblp:63/4133
· DBLP profile ↗
54ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0003-1047-7462ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 3 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | LWMPI: An MPI library for NoC-based lightweight manycore processors with on-chip memory constraintsabstractAbstract Lightweight manycore processors deliver high performance and energy efficiency by bundling hundreds of low‐power cores, a distributed memory architecture with small local memories and Networks‐on‐Chip in a single die. However, the lack of rich and portable programming models for these processors makes software development a challenging task. Currently, two approaches are employed to address programmability in lightweight manycores: Operating Systems (OSes) and baremetal runtime libraries. The former provides portability but exposes complex Operating System (OS)‐level programming interfaces to developers. The latter focuses on providing rich and high performance interfaces, which are vendor‐specific and yield to non‐portable software. In this work, we address these programmability and portability challenges by combining a rich OS with a well‐known standard for parallel programming. We propose a portable and lightweight Message Passing Interface (MPI) library (LWMPI) designed from scratch to cope with restrictions and intricacies of lightweight manycores. We integrated LWMPI into Nanvix, an open‐source distributed OS that runs on silicon lightweight manycores. The results obtained with a synthetic benchmark and a subset of the CAP Bench applications running on Kalray MPPA‐256 unveil that LWMPI not only delivers a lightweight and richer programming interface but also presents good performance and scalability results. João Fellipe Uller, João Vicente Souto, Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 6 |
| 2021 | Memory allocation anomalies in high-performance computing applications: A study with numerical simulationsabstractSummary A memory allocation anomaly occurs when the allocation of a set of heap blocks imposes an unnecessary overhead on the execution of an application. This overhead is particularly disturbing for high‐performance computing (HPC) applications running on shared resources—for example, numerical simulations running on clusters or clouds—because it may increase either the execution time of the application (contributing to a reduction on the overall efficiency of the shared resource) or its memory consumption (eventually inhibiting its capacity to handle larger problems). In this article, we propose a method for identifying, locating, characterizing and fixing allocation anomalies, and a tool for developers to apply the method. We experiment our method and tool with a numerical simulator aimed at approximating the solutions to partial differential equations using a finite element method. We show that taming allocation anomalies in this simulator reduces both its execution time and the memory footprint of its processes, irrespective of the specific heap allocator being employed with it. We conclude that the developer of HPC applications can benefit from the method and tool during the software development cycle. Antônio Tadeu A. Gomes, Enzo Molion, Roberto Pinto Souto, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Inter-kernel communication facility of a distributed operating system for NoC-based lightweight manycores
Pedro Henrique de Mello Morado Penna, João Vicente Souto, João Fellipe Uller, Márcio Castro 0001, Henrique Cota de Freitas, Jean-François Méhaut |
J. Parallel Distributed Comput. | 6 |
| 2021 | ARTful: A model for user-defined schedulers targeting multiple high-performance computing runtime systemsabstractAbstract Global schedulers are components in parallel runtime libraries that distribute the application's workload across physical resources. More often than not, applications showcase dynamic load imbalance and require customized scheduling solutions to avoid wasting resources. Some libraries lack support for user‐defined schedulers and developers resort to unofficial extensions that are harder to reuse and maintain. We propose a global scheduler software design, entitled ARTful model, to create user‐defined solutions with minimal alterations in the runtime library. Our model uses a component‐based design to separate components from the runtime library and the scheduling policy implementation. The ARTful modeldescribes the interface of a portable scheduler library, allowing policies to operate on different runtime libraries. We study the overhead induced by our design through our ARTful library implementation metaprogramming‐oriented global scheduling library using workload‐aware scheduling policies. We experiment with two different policies from OpenMP and Charm++ runtime systems, also presenting evaluations of the policies outside of their original library context. We observe that our portable schedulers can sometimes perform decisions faster than their native counterparts with negligible overhead in the execution times of synthetic applications and molecular dynamics kernels. Alexandre de Limas Santana, Vinicius Freitas, Márcio Castro 0001, Laércio Lima Pilla, Jean-François Méhaut |
Softw. Pract. Exp. | 5 |
| 2020 | Characterization Research on I/O Improvements Targeting DISC and HPC ApplicationsabstractImprovements in I/O architectures are becoming increasingly required nowadays. This is an essential point to complex and data intensive scalable applications. Data-Intensive Scalable Computing (DISC) and High-Performance Computing (HPC) applications frequently need to transfer data between storage resources. In the scientific and industrial fields, the storage component is a key element, because usually those applications employ a huge amount of data. Therefore, the performance of these applications commonly depends on some factors related to time spent in execution of the I/O operations. However, researchers, through their works, are proposing different approaches targeting improvements on the storage layer, thus, reducing the gap between processing and storage. Some solutions combine different hardware technologies to achieve high performance, while others develop solutions on the software layer. This paper aims to present a characterization model for classifying research works on I/O performance improvements for large scale computing facilities. Analysis over 36 different scenarios using a synthetic I/O benchmark demonstrates how the latency parameter behaves when performing different I/O operations using distinct storage technologies and approaches. Laércio Pioli, Eduardo Camilo Inacio, Douglas Dyllon Jeronimo de Macedo, Victor Ströele A. Menezes, José Maria N. David, Jean-François Méhaut, Mario A. R. Dantas |
IECON | 6 |
| 2020 | A Parallel Graph Partitioning Approach to Enhance Community Detection in Social NetworksabstractDealing with complex networks is often a challenge due to the high computational cost in analyzing a huge amount of data. Partitioning methods can decrease the complexity of large structures by reducing them to smaller, less connected parts. Also, the data splitting allows the use of multiprocessing to accelerate the execution of data procedures with simultaneity and parallelism. In this paper, we propose a new parallel partitioning algorithm with a focus on assisting in community detection in social networks. The algorithm uses a subtree-splitting strategy, as well as boundaries defined, in order to cut the network into n balanced subnetworks. Our proposal stands out for the focus on aiding density-based approaches, such as the NetSCAN clustering algorithm, considering two particulars: (i) keeping the partitions connectivity, and; (ii) allowing node overlapping between partitions. Experiments were carried out with different instances intending to investigate the partitions obtained and evaluate our proposal. Furthermore, the algorithm performance analysis in a large network is employed, sequential and parallel implementations are compared in terms of execution time and memory consumption. Evidence was provided that the proposed algorithm is able to split an extensive data set into balanced partitions with optimistic performance results. Tales Lopes, Victor Ströele A. Menezes, Mario A. R. Dantas, Regina Braga 0001, Jean-François Méhaut |
ISCC | 5 |
| 2020 | Increasing the efficiency of Fog Nodes through of Priority-based Load BalancingabstractThe continuous growth and the heterogeneity of the Internet of Things devices is an increasing concern in Fog Computing, where the nodes tend to stay overloaded, which compromises the response times. In response to this challenge, we propose an Architecture Model for Fog Computing and a new Priority Load Balancer that aims to increase the fog nodes’ efficiency. Our research combines task information and computational dynamics load in order to reduce the response time of the Fog Computing. Results show that the proposed solution has the best response time compared to the scenario with direct and round-robin strategies. Our proposed load balancer was able to reduce the response time of high priority tasks by more than 56% compared to other balancers. Eder Paulo Pereira, Edson L. Padoin, Roseclea Duarte Medina, Jean-François Méhaut |
ISCC | 4 |
| 2019 | Managing Power Demand and Load Imbalance to Save Energy on Systems with Heterogeneous CPU SpeedsabstractDifferent simulations of real problems have been executed in High Performance Computing systems. However, the power consumption of these systems is an increasing concern once more energy are consumed to large simulations. In this context, load balancers emerge as a promising alternative for supporting the computational science methods. In response to this challenge, we developed a new heterogeneous energy-aware load balancer called H-ENERGYLB to reduce the average power demand of systems with heterogeneous processors and save energy when scientific applications with imbalanced load are executed. Our new load balancing strategy combines dynamic load balancing with DVFS techniques to mitigate the imbalanced workloads in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experiments with three applications on two different heterogeneous architectures show that H-ENERGYLB results in power reductions of 7.14% in average with the energy saving of 36.6% in average compared to others load balancers. Edson L. Padoin, Matthias Diener, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
SBAC-PAD | 4 |
| 2019 | Energy efficiency and I/O performance of low-power architecturesabstractSummary This paper presents an energy efficiency and I/O performance analysis of low‐power architectures when compared to conventional architectures, with the goal of studying the viability of using them as storage servers. Our results show that despite the fact the power demand of the storage device amounts for a small fraction of the power demand of the whole system, significant increases in power demand are observed when accessing the storage device. We investigate the access pattern impact on power demand, looking at the whole system and at the storage device by itself, and compare all tested configurations regarding energy efficiency. Then we extrapolate the conclusions from this research to provide guidelines for when considering the replacement of traditional storage servers by low‐power alternatives. We show the choice depends on the expected workload, estimates of power demand of the systems, and factors limiting performance. These guidelines can be applied for other architectures than the ones used in this work. Pablo J. Pavan, Ricardo K. Lorenzoni, Vinícius Machado 0002, Jean Luca Bez, Edson L. Padoin, Francieli Zanon Boito, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 8 |
| 2019 | A comprehensive performance evaluation of the BinLPT workload-aware loop schedulerabstractSummary Workload‐aware loop schedulers were introduced to deliver better performance than classical loop scheduling strategies. However, they presented limitations such as inflexible built‐in workload estimators and suboptimal chunk scheduling. Targeting these challenges, we proposed previously a workload‐aware scheduling strategy called BinLPT, which relies on three features: (i) user‐supplied estimations of the workload of the loop; (ii) a greedy heuristic that adaptively partitions the iteration space in several chunks; and (iii) a scheduling scheme based on the Longest Processing Time (LPT) rule and on‐demand technique. In this paper, we present two new contributions to the state‐of‐the‐art. First, we introduce a multiloop support feature to BinLPT, which enables the reuse of estimations across loops. Based on this feature, we integrated BinLPT into a real‐world elastodynamics application, and we evaluated it running on a supercomputer. Second, we present an evaluation of BinLPT using simulations as well as synthetic and application kernels. We carried out this analysis on a large‐scale NUMA machine under a variety of workloads. Our results revealed that BinLPT is better at balancing the workloads of the loop iterations and this behavior improves as the algorithmic complexity of the loop increases. Overall, BinLPT delivers up to 37.15% and 9.11% better performance than well‐known loop scheduling strategies, for the application kernels and the elastodynamics simulation, respectively. Pedro Henrique de Mello Morado Penna, Antônio Tadeu A. Gomes, Márcio Castro 0001, Patricia Della Méa Plentz, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 7 |
| 2018 | Design Space Exploration of Energy Efficient NoC-and Cache-Based Many-Core ArchitectureabstractPerformance of parallel scientific applications on many-core processor architectures is a challenge that increases every day, especially when energy efficiency is concerned. To achieve this, it is necessary to explore architectures with high processing power composed by a network-on-chip to integrate many processing cores and other components. In this context, this paper presents a design space exploration over NoC-based manycore processor architectures with distributed and shared caches, using full-system simulations. We evaluate bottlenecks in such architectures with regard to energy efficiency, using different parallel scientific applications and considering aspects from caches and NoCs jointly. Five applications from NAS Parallel Benchmarks were executed over the proposed architectures, which vary in number of cores; in L2 cache size; and in 12 types of NoC topologies. A clustered topology was set up, in which we obtain performance gains up to 30.56% and reduction in energy consumption up to 38.53%, when compared to a traditional one. Matheus Alcântara Souza, Henrique Cota de Freitas, Jean-François Méhaut |
SBAC-PAD | 3 |
| 2018 | An autonomic-computing approach on mapping threads to multi-cores for software transactional memoryabstractSummary A parallel program needs to manage the trade‐off between the time spent in synchronisation and computation. This trade‐off is significantly affected by its parallelism degree. A high parallelism degree may decrease computing time while increasing synchronisation cost. Furthermore, thread placement on processor cores may impact program performance, as the data access time can vary from one core to another due to intricacies of the underlying memory architecture. Alas, there is no universal rule to decide thread parallelism and its mapping to cores from an offline view, especially for a program with online behaviour variation. Moreover, offline tuning is less precise. We present our work on dynamic control of thread parallelism and mapping. We address concurrency issues via Software Transactional Memory (STM). STM bypasses locks to tackle synchronisation through transactions. Autonomic computing offers designers a framework of methods and techniques to build autonomic systems with well‐mastered behaviours. Its key idea is to implement feedback control loops to design safe, efficient, and predictable controllers, which enable monitoring and adjusting controlled systems dynamically while keeping overhead low. We implement feedback control loops to automate management of threads and diminish program execution time. Naweiluo Zhou, Gwenaël Delaval, Bogdan Robu, Éric Rutten, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Interactive Runtime Verification - When Interactive Debugging Meets Runtime VerificationabstractRuntime Verification consists in studying a system at runtime, looking for input and output events to discover, check or enforce behavioral properties. Interactive debugging consists in studying a system at runtime in order to discover and understand its bugs and fix them, inspecting interactively its internal state.Interactive Runtime Verification (i-RV) combines runtime verification and interactive debugging. We define an efficient and convenient way to check behavioral properties automatically on a program using a debugger. We aim at helping bug discovery and understanding by guiding classical interactive debugging techniques using runtime verification. Raphaël Jakse, Yliès Falcone, Jean-François Méhaut, Kevin Pouget |
ISSRE | 3 |
| 2017 | TWINS: Server Access Coordination in the I/O Forwarding LayerabstractThis paper presents a study of I/O scheduling techniques applied to the I/O forwarding layer. In high-performance computing environments, applications rely on parallel file systems (PFS) to obtain good I/O performance even when handling large amounts of data. To alleviate the concurrency caused by thousands of nodes accessing a significantly smaller number of PFS servers, intermediate I/O nodes are typically applied between processing nodes and the file system. Each intermediate node forwards requests from multiple clients to the system, a setup which gives this component the opportunity to perform optimizations like I/O scheduling. We evaluate scheduling techniques that improve spatiality and request size of the access patterns. We show they are only partially effective because the access pattern is not the main factor for read performance in the I/O forwarding layer. A new scheduling algorithm, TWINS, is presented to coordinate the access of intermediate I/O nodes to the data servers. Our proposal decreases concurrency at the data servers, a factor previously proven to negatively affect performance. The proposed algorithm is able to improve read performance from shared files by up to 28% over other scheduling algorithms and by up to 50% over not forwarding I/O. Jean Luca Bez, Francieli Zanon Boito, Lucas Mello Schnorr, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
PDP | 5 |
| 2017 | Design methodology for workload-aware loop scheduling strategies based on genetic algorithm and simulationabstractSummary In high‐performance computing, the application's workload must be evenly balanced among threads to deliver cutting‐edge performance and scalability. In OpenMP, the load balancing problem arises when scheduling loop iterations to threads. In this context, several scheduling strategies have been proposed, but they do not take into account the input workload of the application and thus turn out to be suboptimal. In this work, we introduce a design methodology to propose, study, and assess the performance of workload‐aware loop scheduling strategies. In this methodology, a genetic algorithm is employed to explore the state space solution of the problem itself and to guide the design of new loop scheduling strategies, and a simulator is used to evaluate their performance. As a proof of concept, we show how the proposed methodology was used to propose and study a new workload‐aware loop scheduling strategy named smart round‐robin (SRR). We implemented this strategy into GNU Compiler Collection's OpenMP runtime. We carry out several experiments to validate the simulator and to evaluate the performance of SRR. Our experimental results show that SRR may deliver up to 37.89%and 14.10%better performance than OpenMP's dynamic loop scheduling strategy in the simulated environment and in a real‐world application kernel, respectively. Copyright © 2016 John Wiley & Sons, Ltd. Pedro Henrique de Mello Morado Penna, Márcio Castro 0001, Henrique Cota de Freitas, François Broquedis, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | CAP Bench: a benchmark suite for performance and energy evaluation of low-power many-core processorsabstractSummary The constant need for faster and more energy‐efficient processors has been stimulating the development of new architectures, such as low‐power many‐core architectures. Researchers aiming to study these architectures are challenged by peculiar characteristics of some components such as networks‐on‐chip and lack of specific tools to evaluate their performance. In this context, the goal of this paper is to present a benchmark suite to evaluate state‐of‐the‐art low‐power many‐core architectures such as the Kalray MPPA‐256 low‐power processor, which features 256 compute cores in a single chip. The benchmark was designed and used to highlight important aspects and details that need to be considered when developing parallel applications for emerging low‐power many‐core architectures. As a result, this paper demonstrates that the benchmark offers a diverse suite of programs with regard to parallel patterns, job types, communication intensity, and task load strategies suitable for a broad understanding of performance and energy consumption of MPPA‐256 and upcoming many‐core architectures. Copyright © 2016 John Wiley & Sons, Ltd. Matheus Alcântara Souza, Pedro Henrique de Mello Morado Penna, Matheus M. Queiroz, Alyson D. Pereira, Fabrício Góes, Henrique Cota de Freitas, Márcio Castro 0001, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 9 |
| 2016 | The mont-blanc prototype: an alternative approach for HPC systemsabstractHigh-performance computing (HPC) is recognized as one of the pillars for further progress in science, industry, medicine, and education. Current HPC systems are being developed to overcome emerging architectural challenges in order to reach Exascale level of performance, projected for the year 2020. The much larger embedded and mobile market allows for rapid development of intellectual property (IP) blocks and provides more flexibility in designing an application-specific system-on-chip (SoC), in turn providing the possibility in balancing performance, energy-efficiency, and cost. In the Mont-Blanc project, we advocate for HPC systems being built from such commodity IP blocks, currently used in embedded and mobile SoCs. As a first demonstrator of such an approach, we present the Mont-Blanc prototype; the first HPC system built with commodity SoCs, memories, and network interface cards (NICs) from the embedded and mobile domain, and off-the-shelf HPC networking, storage, cooling, and integration solutions. We present the system's architecture and evaluate both performance and energy efficiency. Further, we compare the system's abilities against a production level supercomputer. At the end, we discuss parallel scalability and estimate the maximum scalability point of this approach across a set of applications. Nikola Rajovic, Alejandro Rico, Filippo Mantovani, Daniel Ruiz 0003, Josep Oriol Vilarrubi, Constantino Gómez, Luna Backes, Diego Nieto, Harald Servat, Xavier Martorell, Jesús Labarta, Eduard Ayguadé, Chris Adeniyi-Jones, Said Derradji, Hervé Gloaguen, Piero Lanucara, Nico Sanna, Jean-François Méhaut, Kevin Pouget, Brice Videau, Eric Boyer, Momme Allalen, Axel Auweter, David Brayford, Daniele Tafani, Volker Weinberg, Dirk Brömmel, René Halver, Jan H. Meinke, Ramón Beivide, Mariano Benito, Enrique Vallejo 0001, Mateo Valero, Alex Ramírez |
SC | 18 |
| 2016 | Seismic wave propagation simulations on low-power and performance-centric manycores
Márcio Castro 0001, Emilio Francesquini, Fabrice Dupros, Hideo Aochi, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Parallel Comput. | 6 |
| 2015 | Reducing trace size in multimedia applications endurance tests
Serge Vladimir Emteu Tchagou, Alexandre Termier, Jean-François Méhaut, Brice Videau, Miguel Santana, Rene Quiniou |
DATE | 3 |
| 2015 | Data mining approach to temporal debugging of embedded streaming applicationsabstractOne of the greatest challenges in the embedded systems area is to empower software developers with tools that speed up the debugging of QoS properties in applications. Typical streaming applications, such as multimedia (audio/video) decoding, fulfill the QoS properties by respecting the real-time deadlines. A perfectly functional application, when missing these deadlines, may lead to cracks in the sound or perceptible artifacts in the image. We start from the premise that most of the streaming applications that run on embedded systems can be expressed under a data ow model of computation, where the application is represented as a directed graph of the data flowing through computational units called actors. It has been shown that in order to meet real-time constraints the actors should be scheduled in a periodic manner. We exploit this property to propose SATM - a novel approach based on data mining techniques that automatically analyzes execution traces of streaming applications, and discovers significant breaks in the periodicity of actors, as well as potential causes of these breaks. We show on a real use case that our debugging approach can uncover important defects and pinpoint their location to the application developer. Oleg Iegorov, Vincent Leroy 0001, Alexandre Termier, Jean-François Méhaut, Miguel Santana |
EMSOFT | 4 |
| 2015 | Faithful performance prediction of a dynamic task-based runtime system for heterogeneous multi-core architecturesabstractSummary Multi‐core architectures comprising several graphics processing units (GPUs) have become mainstream in the field of high‐performance computing. However, obtaining the maximum performance of such heterogeneous machines is challenging as it requires to carefully off‐load computations and manage data movements between the different processing units. The most promising and successful approaches so far build on task‐based runtimes that abstract the machine and rely on opportunistic scheduling algorithms. As a consequence, the problem gets shifted to choosing the task granularity, task graph structure, and optimizing the scheduling strategies. Trying different combinations of these different alternatives is also itself a challenge. Indeed, obtaining accurate measurements requires reserving the target system for the whole duration of experiments. Furthermore, observations are limited to the few available systems at hand and may be difficult to generalize. In this article, we show how we crafted a coarse‐grain hybrid simulation/emulation of StarPU, a dynamic runtime for hybrid architectures, over SimGrid, a versatile simulator of distributed systems. This approach allows to obtain performance predictions of classical dense linear algebra kernels accurate within a few percents and in a matter of seconds, which allows both runtime and application designers to quickly decide which optimization to enable or whether it is worth investing in higher‐end graphics processing units or not. Additionally, it allows to conduct robust and extensive scheduling studies in a controlled environment whose characteristics are very close to real platforms while having reproducible behavior. Copyright © 2015 John Wiley & Sons, Ltd. Luka Stanisic, Samuel Thibault, Arnaud Legrand, Brice Videau, Jean-François Méhaut |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | On the energy efficiency and performance of irregular application executions on multicore, NUMA and manycore platforms
Emilio Francesquini, Márcio Castro 0001, Pedro Henrique de Mello Morado Penna, Fabrice Dupros, Henrique Cota de Freitas, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
J. Parallel Distributed Comput. | 7 |
| 2014 | Modeling and Simulation of a Dynamic Task-Based Runtime System for Heterogeneous Multi-core Architectures
Luka Stanisic, Samuel Thibault, Arnaud Legrand, Brice Videau, Jean-François Méhaut |
Euro-Par | 5 |
| 2014 | Saving energy by exploiting residual imbalances on iterative applicationsabstractThe power consumption of High Performance Computing (HPC) systems is an increasing concern as large-scale systems grow in size and, consequently, consume more energy. In response to this challenge, we propose two variants of a new energy-aware load balancer that aim at reducing the energy consumption of parallel platforms running imbalanced scientific applications without degrading their performance. Our research combines dynamic load balancing with DVFS techniques in order to reduce the clock frequency of underloaded computing cores which experience some residual imbalance even after tasks are remapped. Experimental results with benchmarks and a real-world application presented energy savings of up to 32% with our fine-grained variant that performs per-core DVFS, and of up to 34% with our coarsegrained variant that performs per-chip DVFS. Edson L. Padoin, Márcio Castro 0001, Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
HiPC | 5 |
| 2014 | Preliminary results of SEU fault-injection on multicore processors in AMP modeabstractThe current technological challenge for computing systems is to use multicore processors in order to ensure reliability, improve performance, and reduce power consumption. This paper presents a method and preliminary results of SEU fault-injection campaigns performed on multicore systems in asymmetric multi-processing mode. The target used for this purpose was a Quad-core processor. This work aims at validating the efficiency of TMR fault-tolerance method and to show the weakest variables of a given application running over multicore systems. Vanessa Vargas, Pablo Ramos, Wassim Mansour, Raoul Velazco, Nacer-Eddine Zergainoh, Jean-François Méhaut |
IOLTS | 6 |
| 2014 | Improving the Performance of Seismic Wave Simulations with Dynamic Load BalancingabstractSeismic wave models provide a way to study the consequences of future earthquakes. When modeling a restricted region, these models require a boundary condition to absorb the energy that goes out of the simulated domain. To parallelize these models, the domain is decomposed into a grid of smaller subdomains which are mapped to different tasks. Due to the boundary condition, this division gives rise to load imbalance between the tasks that simulate border regions and those assigned center subdomains. To deal with this imbalance, and therefore improve the simulation's performance, we propose the use of dynamic load balancing. To evaluate our solution, we ported a seismic wave simulator to Adaptive MPI to profit from its load balancing framework. By using dynamic load balancers, we improved the performance of the application by 23.85% when compared to the original MPI implementation. We also show that load balancers are able to adapt to the variation of load imbalance during the application's execution. Rafael Keller Tesser, Laércio Lima Pilla, Fabrice Dupros, Philippe Olivier Alexandre Navaux, Jean-François Méhaut, Celso L. Mendes |
PDP | 5 |
| 2014 | Energy Efficient Seismic Wave Propagation Simulation on a Low-Power Manycore ProcessorabstractLarge-scale simulation of seismic wave propagation is an active research topic. Its high demand for processing power makes it a good match for High Performance Computing (HPC). Although we have observed a steady increase on the processing capabilities of HPC platforms, their energy efficiency is still lacking behind. In this paper, we analyze the use of a low-power manycore processor, the MPPA-256, for seismic wave propagation simulations. First we look at its peculiar characteristics such as limited amount of on-chip memory and describe the intricate solution we brought forth to deal with this processor's idiosyncrasies. Next, we compare the performance and energy efficiency of seismic wave propagation on MPPA-256 to other commonplace platforms such as general-purpose processors and a GPU. Finally, we wrap up with the conclusion that, even if MPPA-256 presents an increased software development complexity, it can indeed be used as an energy efficient alternative to current HPC platforms, resulting in up to 71% and 81% less energy than a GPU and a general-purpose processor, respectively. Márcio Castro 0001, Fabrice Dupros, Emilio Francesquini, Jean-François Méhaut, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 4 |
| 2014 | Para Miner: a generic pattern mining algorithm for multi-core architectures
Benjamin Négrevergne, Alexandre Termier, Marie-Christine Rousset, Jean-François Méhaut |
Data Min. Knowl. Discov. | 4 |
| 2014 | A topology-aware load balancing algorithm for clustered hierarchical multi-core machines
Laércio Lima Pilla, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Philippe Olivier Alexandre Navaux, Jean-François Méhaut |
Future Gener. Comput. Syst. | 7 |
| 2014 | Adaptive thread mapping strategies for transactional memory applications
Márcio Castro 0001, Fabrício Góes, Jean-François Méhaut |
J. Parallel Distributed Comput. | 3 |
| 2013 | Performance analysis of HPC applications on low-power embedded platformsabstractThis paper presents performance evaluation and analysis of well-known HPC applications and benchmarks running on low-power embedded platforms. The performance to power consumption ratios are compared to classical x86 systems. Scalability studies have been conducted on the Mont-Blanc Tibidabo cluster.We have also investigated optimization opportunities and pitfalls induced by the use of these new platforms, and proposed optimization strategies based on auto-tuning. Luka Stanisic, Brice Videau, Johan Cronsioe, Augustin Degomme, Vania Marangozova-Martin, Arnaud Legrand, Jean-François Méhaut |
DATE | 7 |
| 2013 | A NUMA-Aware Runtime Environment for the Actor ModelabstractThe actor model is present in several mission-critical systems, such as those supporting WhatsApp and Twitter. These systems serve thousands of clients simultaneously, therefore demanding substantial computing resources usually provided by multiprocessor and multicore platforms. Non-Uniform Memory Access (NUMA) architectures account for an important share of these platforms. Yet, little or no research has been done on the suitability of the current actor runtime environments for these machines. Current runtime environments assume a flat memory space, thus not performing as well as they could. The NUMA environment presents challenges to the actor model runtime environment in fields varying from memory management to scheduling and load-balancing. In this document we analyze and characterize actor based applications to, in light of the above, propose improvements to actor runtime environments. As a proof of concept, we have applied our ideas in a real actor runtime environment, the Erlang virtual machine. This modified virtual machine uses the NUMA characteristics and the application knowledge to take better memory management, scheduling and load-balancing decisions. We have evaluated this modified runtime environment using standard benchmarks and, taking the default virtual machine as a baseline, we improved the performance of the tested applications by a factor of 2.50 on the best case while limiting our slowdown on the worst case by a factor of 1.09. Emilio Francesquini, Alfredo Goldman, Jean-François Méhaut |
ICPP | 3 |
| 2012 | Debugging embedded multimedia application traces through periodic pattern miningabstractIncreasing complexity in both the software and the underlying hardware, and ever tighter time-to-market pressures are some of the key challenges faced when designing multimedia embedded systems. Optimizing the debugging phase can help to reduce development time significantly. A powerful approach used extensively during this phase is the analysis of execution traces. However, huge trace volumes make manual trace analysis unmanageable. In such situations, Data Mining can help by automatically discovering interesting patterns in large amounts of data. In this paper, we are interested in discovering periodic behaviors in multimedia applications. Therefore, we propose a new pattern mining approach for automatically discovering all periodic patterns occurring in a multimedia application execution trace. Patricia López Cueva, Aurélie Bertaux, Alexandre Termier, Jean-François Méhaut, Miguel Santana |
EMSOFT | 4 |
| 2012 | Dynamic Thread Mapping Based on Machine Learning for Transactional Memory Applications
Márcio Castro 0001, Fabrício Góes, Luiz Gustavo Fernandes, Jean-François Méhaut |
Euro-Par | 4 |
| 2012 | Asymptotically Optimal Load Balancing for Hierarchical Multi-Core SystemsabstractCurrent multi-core machines feature a complex and hierarchical core topology, multiple levels of cache and memory subsystem with NUMA design. Although this design provides high processing power to parallel machines, it comes with the cost of asymmetric memory access latencies. Depending on the parallel application communication patterns, this asymmetry may reduce the overall performance of the system. Therefore, to achieve scalable performance in this environment, it becomes crucial to exploit the machine architecture while taking into account the application communication patterns. In this paper, we introduce a topology-aware load balancing algorithm named HWTOPOLB. It combines the machine topology characteristics with the communication patterns of the application to equalize the application load on the available cores while reducing latencies. We also present the proof that the algorithm is asymptotically optimal (Theorem 1). We have implemented our load balancing algorithm using the CHARM++ Parallel System and analyzed its performance using three different benchmarks. Our experimental results show that the HWTOPOLB can achieve average performance improvements of 24% when compared to existing load balancing strategies on three different multi-core machines. Laércio Lima Pilla, Philippe Olivier Alexandre Navaux, Christiane Pousa Ribeiro, Pierre Coucheney, François Broquedis, Bruno Gaujal, Jean-François Méhaut |
ICPADS | 7 |
| 2012 | A Hierarchical Approach for Load Balancing on Parallel Multi-core SystemsabstractMulti-core compute nodes with non-uniform memory access (NUMA) are now a common architecture in the assembly of large-scale parallel machines. On these machines, in addition to the network communication costs, the memory access costs within a compute node are also asymmetric. Ignoring this can lead to an increase in the data movement costs. Therefore, to fully exploit the potential of these nodes and reduce data access costs, it becomes crucial to have a complete view of the machine topology (i.e. the compute node topology and the interconnection network among the nodes). Furthermore, the parallel application behavior has an important role in determining how to utilize the machine efficiently. In this paper, we propose a hierarchical load balancing approach to improve the performance of applications on parallel multi-core systems. We introduce NucoLB, a topology-aware load balancer that focuses on redistributing work while reducing communication costs among and within compute nodes. NucoLB takes the asymmetric memory access costs present on NUMA multi-core compute nodes, the interconnection network overheads, and the application communication patterns into account in its balancing decisions. We have implemented NucoLB using the Charm++ parallel runtime system and evaluated its performance. Results show that our load balancer improves performance up to 20% when compared to state-of-the-art load balancers on three different NUMA parallel machines. Laércio Lima Pilla, Christiane Pousa Ribeiro, Daniel Cordeiro, Abhinav Bhatele, Philippe Olivier Alexandre Navaux, François Broquedis, Jean-François Méhaut, Laxmikant V. Kalé |
ICPP | 8 |
| 2011 | Resource Management of Virtual Infrastructure for On-demand SaaS Services
Rodrigue Chakode, Blaise Omer Yenke, Jean-François Méhaut |
CLOSER | 3 |
| 2011 | A Generic Component-Based Approach to MPSoC ObservationabstractObserving the execution of MPSoC is a key activity in debugging and optimizing embedded applications. However, providing suitable observation tools becomes a real challenge with the fast evolution of embedded platforms. We propose a generic component-based approach which allows for partial and configurable MPSoC observation. Genericity is obtained by encapsulating specific embedded features and exporting a generic observation API. Partial observation is achieved by instantiating and attaching observation components only to the entities of interest: application modules, HW components or different levels of the SW stack. Configurability is related to the possibility to organize and parametrize observation treatments (data collection, filtering, storing) according to the target system. The approach has been validated in several application and platform contexts including an embedded platform built by STMicroelectronics. Carlos Hernan Prada-Rojas, Vania Marangozova-Martin, Jean-François Méhaut, Miguel Santana |
EUC | 3 |
| 2011 | A Contention-Aware Performance Model for HPC-Based Networks: A Case Study of the InfiniBand Network
Maxime Martinasso, Jean-François Méhaut |
Euro-Par (1) | 2 |
| 2011 | Introduction
Sabri Pllana, Jean-François Méhaut, Eduard Ayguadé, Herbert Cornelius, Jacob Barhen |
Euro-Par (2) | 2 |
| 2011 | A machine learning-based approach for thread mapping on transactional memory applicationsabstractThread mapping has been extensively used as a technique to efficiently exploit memory hierarchy on modern chip-multiprocessors. It places threads on cores in order to amortize memory latency and/or to reduce memory contention. However, efficient thread mapping relies upon matching application behavior with system characteristics. Particularly, Software Transactional Memory (STM) applications introduce another dimension due to its runtime system support. Existing STM systems implement several conflict detection and resolution mechanisms, which leads STM applications to behave differently for each combination of these mechanisms. In this paper we propose a machine learning-based approach to automatically infer a suitable thread mapping strategy for transactional memory applications. First, we profile several STM applications from the STAMP benchmark suite considering application, STM system and platform features to build a set of input instances. Then, such data feeds a machine learning algorithm, which produces a decision tree able to predict the most suitable thread mapping strategy for new unobserved instances. Results show that our approach improves performance up to 18.46% compared to the worst case and up to 6.37% over the Linux default thread mapping strategy. Márcio Castro 0001, Fabrício Góes, Christiane Pousa Ribeiro, Murray Cole, Marcelo Cintra, Jean-François Méhaut |
HiPC | 6 |
| 2011 | Analysis and Tracing of Applications Based on Software Transactional Memory on Multicore ArchitecturesabstractTransactional Memory (TM) is a new programming paradigm that offers an alternative to traditional lock-based concurrency mechanisms. It offers a higher-level programming interface and promises to greatly simplify the development of correct concurrent applications on multicore architectures. However, simplicity often comes with an important performance deterioration and given the variety of TM implementations it is still a challenge to know what kind of applications can really take advantage of TM. In order to gain some insight on these issues, helping developers to understand and improve the performance of TM applications, we propose a generic approach for collecting and tracing relevant information about transactions. Our solution can be applied to different Software Transactional Memory (STM) libraries and applications as it does not modify neither the target application nor the STM library source codes. We show that the collected information can be helpful in order to comprehend the performance of TM applications. Márcio Castro 0001, Kiril Georgiev, Vania Marangozova-Martin, Jean-François Méhaut, Luiz Gustavo Fernandes, Miguel Santana |
PDP | 4 |
| 2011 | Scheduling of Computing Services on Intranet NetworksabstractNowadays, enterprises can provide computing services through their intranet networks by letting their available resources be used as virtual clusters for scientific computation during idle periods such as nights, weekends, and holidays. Generally, these idle periods do not permit to carry out the computations completely. It is therefore necessary to save the context of uncompleted applications for possible restart. This checkpointing mechanism is subject to resource constraints: the network bandwidth, the disk bandwidth, and the delay T imposed for releasing the workstations. We first introduce a function bw that gives the bandwidth bw(m,V) of a system during the checkpointing of m applications with aggregated memory requirement V. Assuming that this bandwidth is shared equitably among the applications, the scheduling problem becomes a sequence of knapsack problems with nonlinear constraints for which we propose approximate solutions. Experiments carried out on Grid5000 show that the running time of this algorithm is negligible compared to the delay T which is of the order of few minutes. This means that the proposed scheduling algorithm does not induce a significant overhead on the checkpointing process. As a consequence, our mechanism can be incorporated in a batch scheduler. Blaise Omer Yenke, Jean-François Méhaut, Maurice Tchuenté |
IEEE Trans. Serv. Comput. | 2 |
| 2010 | High Performance Computing on Demand: Sharing and Mutualization of ClustersabstractFor software vendors who need to provide their softwares as services via Internet, an infrastructure of high performance computing (HPC) such as clusters is required. However, for small and medium enterprises (SMEs) and/or startup businesses, owning a cluster is generally out of reach. Indeed, the cost of a cluster can be very high, since enterprises have to deal with acquisition costs, as well as many operating costs (engineering, power supply, air conditioning, etc.). The emergence of infrastructure providers – like Amazon, Google, IBM, Sun, etc. – allows businesses to use remote infrastructures. However, in the long term, renting those infrastructures can also be expensive. In order to lower the cost, an alternative solution might be for small businesses to join in order to purchase and maintain a common infrastructure that would be shared among them. In this case, each partner has to have the guarantee that their use of the infrastructure would be equitable, rational and proportional to their investment. Additionally, customers would expect the service/application to be cheaper and to have a good performance. In principle, infrastructure sharing is not simple to manage. In this paper, we have defined two approaches concerning the equitable sharing of a cluster among several concurrent softwares hosted as services. The first approach is based on the static partitioning of resources, and the second approach is based on the dynamic resource allocation with dynamic priorities among applications. This work has been carried out in cooperation with an industrial project named CILOE1. Such a project aims at providing a shared computing cluster to small editors of electronic design automation (EDA) and embedded softwares. Rodrigue Chakode, Jean-François Méhaut, Francois Charlet |
AINA | 2 |
| 2009 | PaSTeL: Parallel Runtime and Algorithms for Small DatasetsabstractIn this paper, we put forward PaSTeL, an engine dedicated to parallel algorithms. PaSTeL offers both a programming model, to build parallel algorithms and an execution model based on work-stealing. Special care has been taken on using optimized thread activation and synchronization mechanisms. In order to illustrate the use of PaSTeL a subset of the STL's algorithms was implemented, which were also used on performance experiments. PaSTeL's performance is evaluated on a laptop computer using two cores, but also on a 16 cores platform. PaSTeL shows better performance than other implementations of the STL, especially on small datasets. Brice Videau, Erik Saule, Jean-François Méhaut |
CISIS | 3 |
| 2009 | NUMA-ICTM: A parallel version of ICTM exploiting memory placement strategies for NUMA machinesabstractIn geophysics, the appropriate subdivision of a region into segments is extremely important. ICTM (interval categorizer tesselation model) is an application that categorizes geographic regions using information extracted from satellite images. The categorization of large regions is a computational intensive problem, what justifies the proposal and development of parallel solutions in order to improve its applicability. Recent advances in multiprocessor architectures lead to the emergence of NUMA (non-uniform memory access) machines. In this work, we present NUMA-ICTM: a parallel solution of ICTM for NUMA machines. First, we parallelize ICTM using OpenMP. After, we improve the OpenMP solution using the MAI (memory affinity interface) library, which allows a control of memory allocation in NUMA machines. The results show that the optimization of memory allocation leads to significant performance gains over the pure OpenMP parallel solution. Márcio Castro 0001, Luiz Gustavo Fernandes, Christiane Pousa Ribeiro, Jean-François Méhaut, Marilton S. de Aguiar |
IPDPS | 4 |
| 2009 | Memory Affinity for Hierarchical Shared Memory MultiprocessorsabstractCurrently, parallel platforms based on large scale hierarchical shared memory multiprocessors with Non-Uniform Memory Access (NUMA) are becoming a trend in scientific High Performance Computing (HPC). Due to their memory access constraints, these platforms require a very careful data distribution. Many solutions were proposed to resolve this issue. However, most of these solutions did not include optimizations for numerical scientific data (array data structures) and portability issues. Besides, these solutions provide a restrict set of memory policies to deal with data placement. In this paper, we describe an user-level interface named Memory Affinity interface (MAi), which allows memory affinity control on Linux based cache-coherent NUMA (ccNUMA) platforms. Its main goals are, fine data control, flexibility and portability. The performance of MAi is evaluated on three ccNUMA platforms using numerical scientific HPC applications, the NAS Parallel Benchmarks and a Geophysics application. The results show important gains (up to 31\%) when compared to Linux default solution. Christiane Pousa Ribeiro, Jean-François Méhaut, Alexandre Carissimi, Márcio Castro 0001, Luiz Gustavo Fernandes |
SBAC-PAD | 2 |
| 2008 | Scheduling Deadline-Constrained Checkpointing on Virtual ClustersabstractWe consider a context where the available resources of the Intranet of a company are used as a virtual cluster for scientific computation, during the idle periods (nights, weekends, holidays, ). Generally, these idle periods do not permit to carry out completely the computations. For instance, a workstation mobilized during the night must be released in the morning to make it available for the employee, even if the application running on it is not completed. It is therefore necessary to save the context of uncompleted applications for possible restart. Hereafter, we assume that the computations running on the workstations are independent from each other. The checkpointing mechanism which ensures the continuity of applications is subject to resource constraints : the network bandwidth, the disk bandwidth and the delay T imposed for releasing the workstations. We first show that the designing of a scheduling strategy which optimizes resource consumption while taking into account the above constraints, can be formalized as a variant of the classical 0/1 knapsack problem. We then propose an algorithm whose implementation does not have a significant overhead on checkpointing mechanisms. Experiments carried out on a real cluster show that this algorithm performs better than the naive scheduling algorithm which selects the applications one after the other in order of decreasing amount of resource consumption. Blaise Omer Yenke, Jean-François Méhaut, Maurice Tchuenté |
APSCC | 2 |
| 2008 | Predictive models for bandwidth sharing in high performance clustersabstractUsing MPI as communication interface, one or several applications may introduce complex communication behaviors over the network cluster. This effect is increased when nodes of the cluster are multi-processors, and where communications can income or outgo from the same node with a common interval time. Our goal is to understand those behaviors to build a class of predictive models of bandwidth sharing, knowing, on the one hand the flow control mechanisms and, on the other hand, a set of experimental results. This paper present experiences that show how is shared the bandwidth on gigabit Ethernet, Myrinet 2000 and Infiniband network before to introduce the models for Gigabit Ethernet and Myrinet 2000 networks. Jérôme Vienne, Maxime Martinasso, Jean-Marc Vincent, Jean-François Méhaut |
CLUSTER | 4 |
| 2002 | Madeleine II: a portable and efficient communication library for high-performance cluster computing
Olivier Aumage, Luc Bougé, Jean-François Méhaut, Raymond Namyst |
Parallel Comput. | 3 |
| 2000 | Madeleine II: a Portable and Efficient Communication Library for High-Performance Cluster ComputingabstractThis paper introduces Madeleine II, a new adaptive and portable multi-protocol implementation of the Madeleine communication library. Madeleine II has the ability to control multiple network interfaces (BIP, SISCI, VIA) and multiple network adapters (Ethernet, Myrinet, SCI) within the same application session. Moreover it includes advanced mechanisms to dynamically select the most appropriate transfer method for a given network protocol according to various parameters such as data size or responsiveness user requirements. We report on performance measurements obtained using BIP/Myrinet and SISCI/SCI and we present preliminary results about our Nexus/Madeleine II and MPICH/Madeleine II ports. Olivier Aumage, Luc Bougé, Alexandre Denis 0001, Jean-François Méhaut, Guillaume Mercier, Raymond Namyst, Loïc Prylli |
CLUSTER | 4 |
| 1999 | Multi-protocol Communications and High Speed Networks
Benoît Planquelle, Jean-François Méhaut, Nathalie Revol |
Euro-Par | 2 |
| 1992 | Communicating active components: An environment for concurrent applications on parallel machines
Luc Courtrai, Jean-François Roos, Jean-Marc Geib, Jean-François Méhaut |
Microprocess. Microprogramming | 4 |
| 1991 | An analysis of communication and multiprogramming in the Helios operating system
Fred Hemery, Dominique Lazure, Eric Delattre, Jean-François Méhaut |
Microprocessing and Microprogramming | 4 |