EDBT 2026 Demo / reviewers in the wild / expert
Brice Goglin
dblp:59/4382
· DBLP profile ↗
35ranked-venue papers
15as first author
10since 2021 · last 2026
0000-0002-8671-4615ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 12 first-author · 9 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Impact of Interference from Concurrent Jobs on Checkpointing PerformanceabstractI/O has been identified as one of the main bottlenecks in HPC. Among the most I/O-intensive operations is checkpointing, which is necessary to save the state of an application and allow it to be restarted at an advanced stage of computation. However, near the parallel file system, concurrency prevents checkpoint phases from reaching the best I/O performance. In this paper, we study I/O interference in this specific context: we look at performance of a checkpoint phase when faced with different interference patterns, exploring aspects such as scale, number of processes, operation, number of files, etc. Through an extensive experimentation, in two systems, we show the impact of these aspects on checkpoint. Moreover, we show that some configurations — e.g., an application that does random accesses — lead to degraded system I/O performance. This paper provides an important background for any effort into mitigating I/O interference and into improving checkpointing performance. Méline Trochon, Jean-Thomas Acquaviva, Francieli Zanon Boito, Brice Goglin, Francois Tessier, Luan Teylo |
SSDBM | 4 |
| 2025 | Communication Notification Through User-Level Interrupts for the BXI NetworkabstractTo reduce the cost of communications in highperformance computing, it is possible to overlap communications with computations. Some communication protocols, such as rendez-vous, multi-chunk messages, and collectives, may require a completion notification to be processed before they can further progress. With active polling or passive waiting, completions are not processed while the application is busy with computation, and thus communication does not progress. However, with an eventbased method like interrupts, it is expected to be much more reactive. Nevertheless, using interrupts usually involves system calls, which are avoided with high-performance networks. The Intel Sapphire Rapids processors introduced user-level interrupts (UINTR), hardware interrupts designed to be used directly in user space, without going through the kernel. However, their current implementation is limited to inter-process communication. They cannot be triggered from a device. In this paper, we propose new mechanisms to extend the scope of user-level interrupts, so as to be able to trigger them from a device and not only from a CPU. We have implemented these mechanisms in the BXI network from Eviden. We have evaluated their performance: we obtain a latency only 2.4 times higher than active polling (v.s. 6 times higher for interrupts with system calls). We have assessed their ability to make communication progress when overlapped with computation; we observe a near-perfect computation/communication overlap. Charles Goedefroit, Alexandre Denis 0001, Mathieu Barbe, Brice Goglin, Gregoire Pichon |
CLUSTER | 4 |
| 2025 | Performance Projection for Design-Space Exploration on future HPC ArchitecturesabstractTo address the growing need for performance from future HPC machines, their processor designs are constantly evolving. Assessing the impact of changes in hardware, software stack, and applications on performance is crucial in every step of a codesign process. Here, we propose a performance projection workflow to facilitate the exploration of design space for multicore nodes and multi-threaded applications. For this purpose, we analyze the architectural efficiency of an accessible source machine and determine the maximum sustainable flop/s performance of a hypothetical target machine based on its software stack on a per-thread basis. Finally, we use these characterizations to project the performance evolution from the source machine to the target machine. In this work, we assess the strengths and weaknesses of our approach by integrating it into the Fugaku-Next Feasibility Study. We compare the accuracy and overhead of our approach with the gem5 cycle-level simulations and a fast exploration methodology based on Machine Code Analyzer (MCA), using NAS Parallel benchmarks and CCS-QCD, a quantum chromodynamics miniapp. The study demonstrates that, compared to gem5, our approach has a prediction deviation of 5% for most cases and up to 30% for extreme cases. Additionally, it exhibits an execution overhead an order of magnitude bigger than MCA but orders of magnitude smaller than gem5. Finally, we demonstrate our approach's capability to study larger scale and more representative applications than gem5, such as QWS and Genesis, two applications of RIKEN optimized for Fugaku. Clément Gavoille, Hugo Taboada, Jens Domke, Brice Goglin, Emmanuel Jeannot |
IPDPS | 4 |
| 2024 | Phase-Based Data Placement Optimization in Heterogeneous MemoryabstractWhile scientific applications show increasing demand for memory speed and capacity, the performance gap between compute cores and the memory subsystem continues to spread. In response, heterogeneous memory systems integrating high-bandwidth memory (HBM) and non-volatile memory (NVM) alongside traditional DRAM on the CPU side are gaining traction. Despite the potential benefits of optimized memory selection for improved performance and efficiency, adapting applications to leverage diverse memory types often requires extensive modifications. Moreover, applications often comprise multiple execution phases with varying data access patterns. Since the capacity of the “fastest” memory is limited, relying solely on fixed data placement decisions may not yield optimal performance. Thus, considering allocation lifetimes and dynamically migrating data between memory types becomes imperative to ensure that performance-critical data for each phase resides in fast memory. To address these challenges, we developed a workflow incorporating memory access profiling, optimization techniques and a runtime system, which selects initial data placement for allocations and performs data migration during execution, considering the platform's memory subsystem characteristics and capacities. We formalize the optimization problems for initial and phase-based data placement and propose heuristics derived from memory profiling metrics to solve it. Additionally, we outline the implementation of these approaches, including allocation interception to enforce placement decisions. Experiments conducted with several applications on an Intel Ice Lake$(\text{DRAM}+\text{NVM})$and Sapphire Rapids$(\text{HBM}+\text{DRAM})$system demonstrate that our methodology can effectively bridge the performance gap between slow and fast memory in heterogeneous memory environments. Jannis Klinkenberg, Clément Foyer, Pierre Clouzet, Brice Goglin, Emmanuel Jeannot, Christian Terboven, Anara Kozhokanova |
CLUSTER | 4 |
| 2023 | H2M: Exploiting Heterogeneous Shared Memory ArchitecturesabstractOver the past decades, the performance gap between the memory subsystem and compute capabilities continued to spread. However, scientific applications and simulations show increasing demand for both memory speed and capacity. To tackle these demands, new technologies such as high-bandwidth memory (HBM) or non-volatile memory (NVM) emerged, which are usually combined with classical DRAM. The resulting architecture is a heterogeneous memory system in which no single memory is “best”. HBM is smaller but offers higher bandwidth than DRAM, whereas NVM provides larger capacity than DRAM at a reasonable cost and less energy consumption. Despite that, in several cases, DRAM still offers the best latency out of all three technologies. In order to use different kinds of memory, applications typically have to be modified to a great extent. Consequently, vendor-agnostic solutions are desirable. First, they should offer the functionality to identify kinds of memory, and second, to allocate data on it. In addition, because memory capacities may be limited, decisions about data placement regarding the different memory kinds have to be made. Finally, in making these decisions, changes over time in data that is accessed, and the actual access pattern, should be considered for initial data placement and be respected in data migration at run-time. In this paper, we introduce a new methodology that aims to provide portable tools and methods for managing data placement in systems with heterogeneous memory. Our approach allows programmers to provide traits (hints) for allocations that describe how data is used and accessed. Combined with characteristics of the platforms’ memory subsystem, these traits are exploited by heuristics to decide where to place data items. We also discuss methodologies for analyzing and identifying memory access characteristics of existing applications, and for recommending allocation traits. In our evaluation, we conduct experiments with several kernels and two proxy applications on Intel Knights Landing (HBM + DRAM) and Intel Ice Lake with Intel Optane DC Persistent Memory (DRAM + NVM) systems. We demonstrate that our methodology can bridge the performance gap between slow and fast memory by applying heuristics for initial data placement. Jannis Klinkenberg, Anara Kozhokanova, Christian Terboven, Clément Foyer, Brice Goglin, Emmanuel Jeannot |
Future Gener. Comput. Syst. | 5 |
| 2023 | A survey of software techniques to emulate heterogeneous memory systems in high-performance computing
Clément Foyer, Brice Goglin, Andrès Rubio Proaño |
Parallel Comput. | 2 |
| 2022 | H2M: Towards Heuristics for Heterogeneous MemoryabstractFor the past years, scientific applications and simulations show increasing demand for both memory speed and capacity. The performance gap between compute units and the memory subsystem continues to spread which led to redesigns and the emergence of new technologies. Recent architectures already comprise, next to classical DRAM, portions of High Bandwidth Memory (HBM) that has less capacity than DRAM and is solving only one of the requirements. The newly introduced Non- Volatile Memory (NVM) shows performance closer to DRAM, while providing terabytes of capacity, consuming less power and having a better price per byte ratio. Clément Foyer, Brice Goglin, Emmanuel Jeannot, Jannis Klinkenberg, Anara Kozhokanova, Christian Terboven |
CLUSTER | 2 |
| 2022 | Relative Performance Projection on Arm Architectures
Clément Gavoille, Hugo Taboada, Patrick Carribault, Fabrice Dupros, Brice Goglin, Emmanuel Jeannot |
Euro-Par | 5 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 29 |
| 2021 | Profiles of Upcoming HPC Applications and Their Impact on Reservation StrategiesabstractWith the expected convergence between HPC, BigData and AI, new applications with different profiles are coming to HPC infrastructures. We aim at better understanding the features and needs of these applications in order to be able to run them efficiently on HPC platforms. The approach followed is bottom-up: we study thoroughly an emerging application, Spatially Localized Atlas Network Tiles (SLANT, originating from the neuroscience community) to understand its behavior. Based on these observations, we derive a generic, yet simple, application model (namely, a linear sequence of stochastic jobs). We expect this model to be representative for a large set of upcoming applications from emerging fields that start to require the computational power of HPC clusters without fitting the typical behavior of large-scale traditional applications. In a second step, we show how one can use this generic model in a scheduling framework. Specifically we consider the problem of making reservations (both time and memory) for an execution on an HPC platform based on the application expected resource requirements. We derive solutions using the model provided by the first step of this work. We experimentally show the robustness of the model, even with very few data points or using another application, to generate the model, and provide performance gains with regards to standard and more recent approaches used in the neuroscience community. Ana Gainaru, Brice Goglin, Valentin Honoré, Guillaume Pallez |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Reservation and Checkpointing Strategies for Stochastic JobsabstractIn this paper, we are interested in scheduling and checkpointing stochastic jobs on a reservation-based platform, whose cost depends both (i) on the reservation made, and (ii) on the actual execution time of the job. Stochastic jobs are jobs whose execution time cannot be determined easily. They arise from the heterogeneous, dynamic and data-intensive requirements of new emerging fields such as neuroscience. In this study, we assume that jobs can be interrupted at any time to take a checkpoint, and that job execution times follow a known probability distribution. Based on past experience, the user has to determine a sequence of fixed-length reservation requests, and to decide whether the state of the execution should be checkpointed at the end of each request. The objective is to minimize the expected cost of a successful execution of the jobs. We provide an optimal strategy for discrete probability distributions of job execution times, and we design fully polynomial-time approximation strategies for continuous distributions with bounded support. These strategies are then experimentally evaluated and compared to standard approaches such as periodic-length reservations and simple checkpointing strategies (either checkpoint all reservations, or none). The impact of an imprecise knowledge of checkpoint and restart costs is also assessed experimentally. Ana Gainaru, Brice Goglin, Valentin Honoré, Guillaume Pallez, Padma Raghavan, Yves Robert, Hongyang Sun 0001 |
IPDPS | 2 |
| 2019 | Data and Thread Placement in NUMA Architectures: A Statistical Learning ApproachabstractNowadays, NUMA architectures are common in compute-intensive systems. Achieving high performance for multi-threaded application requires both a careful placement of threads on computing units and a thorough allocation of data in memory. Finding such a placement is a hard problem to solve, because performance depends on complex interactions in several layers of the memory hierarchy. In this paper we propose a black-box approach to decide if an application execution time can be impacted by the placement of its threads and data, and in such a case, to choose the best placement strategy to adopt. We show that it is possible to reach near-optimal placement policy selection. Furthermore, solutions work across several recent processor architectures and decisions can be taken with a single run of low overhead profiling. Nicolas Denoyelle, Brice Goglin, Emmanuel Jeannot, Thomas Ropars |
ICPP | 2 |
| 2019 | Modeling Non-Uniform Memory Access on Large Compute Nodes with the Cache-Aware Roofline ModelabstractNUMA platforms, emerging memory architectures with on-package high bandwidth memories bring new opportunities and challenges to bridge the gap between computing power and memory performance. Heterogeneous memory machines feature several performance trade-offs, depending on the kind of memory used, when writing or reading it. Finding memory performance upper-bounds subject to such trade-offs aligns with the numerous interests of measuring computing system performance. In particular, representing applications performance with respect to the platform performance bounds has been addressed in the state-of-the-art Cache-Aware Roofline Model (CARM) to troubleshoot performance issues. In this paper, we present a Locality-Aware extension (LARM) of the CARM to model NUMA platforms bottlenecks, such as contention and remote access. On top of this, the new contribution of this paper is the design and validation of a novel hybrid memory bandwidth model. This new hybrid model quantifies the achievable bandwidth upper-bound under above-described trade-offs with less than 3 percent error. Hence, when comparing applications performance with the maximum attainable performance, software designers can now rely on more accurate information. Nicolas Denoyelle, Brice Goglin, Aleksandar Ilic, Emmanuel Jeannot, Leonel Sousa |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Co-Scheduling HPC Workloads on Cache-Partitioned CMP PlatformsabstractCo-scheduling techniques are used to improve the throughput of applications on chip multiprocessors (CMP), but sharing resources often generates critical interferences. We focus on the interferences in the last level of cache (LLC) and use the Cache Allocation Technology (CAT) recently provided by Intel to partition the LLC and give each co-scheduled application their own cache area. We consider m iterative HPC applications running concurrently and answer the following questions: (i) how to precisely model the behavior of these applications on the cache partitioned platform? and (ii) how many cores and cache fractions should be assigned to each application to maximize the platform efficiency? Here, platform efficiency is defined as maximizing the performance either globally, or as guaranteeing a fixed ratio of iterations per second for each application. Through extensive experiments using CAT, we demonstrate the impact of cache partitioning when multiple HPC application are co-scheduled onto CMP platforms. Guillaume Pallez, Anne Benoit, Brice Goglin, Loïc Pottier, Yves Robert |
CLUSTER | 3 |
| 2018 | Hardware topology management in MPI applications through hierarchical communicators
Brice Goglin, Emmanuel Jeannot, Farouk Mansouri, Guillaume Mercier |
Parallel Comput. | 1 |
| 2017 | On the Overhead of Topology Discovery for Locality-Aware Scheduling in HPCabstractThe increasing complexity of parallel computing platforms requires a deep knowledge of the hardware and of the application needs. Locality a key criteria for performance optimization. It involves software tools to expose information about the hardware topology to high performance runtime libraries. We show that the overhead of gathering such information from the operating system is significant on large computing nodes that run Linux. This overhead also increases more than linearly with the number of processes that perform it simultaneously. We then study the actual needs of the HPC software ecosystem in terms of topology information. We propose some ways to avoid multiple expensive topology discovery and to share topology information between components such as the resource manager or the runtime libraries. Brice Goglin |
PDP | 1 |
| 2013 | KNEM: A generic and scalable kernel-assisted intra-node MPI communication framework
Brice Goglin, Stéphanie Moreaud |
J. Parallel Distributed Comput. | 1 |
| 2011 | Introduction
Jesper Larsson Träff, Brice Goglin, Ulrich Brüning 0001, Fabrizio Petrini |
Euro-Par (2) | 2 |
| 2011 | Kernel Assisted Collective Intra-node MPI Communication among Multi-Core and Many-Core CPUsabstractShared memory is among the most common approaches to implementing message passing within multicorenodes. However, current shared memory techniques donot scale with increasing numbers of cores and expanding memory hierarchies -- most notably when handling large data transfers and collective communication. Neglecting the underlying hardware topology, using copy-in/copy-out memory transfer operations, and overloading the memory subsystem using one-to-many types of operations are some of the most common mistakes in today's shared memory implementations. Unfortunately, they all negatively impact the performance and scalability of MPI libraries -- and therefore applications. In this paper, we present several kernel-assisted intra-node collective communication techniques that address these three issues on many-core systems. We also present a new OpenMPI collective communication component that uses the KNEMLinux module for direct inter-process memory copying. Our Open MPI component implements several novel strategies to decrease the number of intermediate memory copies and improve data locality in order to diminish both cache pollution and memory pressure. Experimental results show that our KNEM-enabled Open MPI collective component can outperform state-of-art MPI libraries (Open MPI and MPICH2) on synthetic benchmarks, resulting in a significant improvement for a typical graph application. George Bosilca, Aurelien Bouteiller, Brice Goglin, Jeffrey M. Squyres, Jack J. Dongarra |
ICPP | 4 |
| 2011 | NIC-assisted cache-efficient receive stack for message passing over EthernetabstractAbstract High‐speed networking in clusters usually relies on advanced hardware features in the NICs, such as zero‐copy capability. Open‐MX is a high‐performance message‐passing stack tailored for regular Ethernet hardware without such capabilities. We present the addition of a multiqueue support in the Open‐MX receive stack so that all incoming packets for the same process are handled on the same core. We then introduce the idea of binding the target end process near its dedicated receive queue. This model leads to a more cache‐efficient receive stack for Open‐MX. It also proves that very simple and stateless hardware features may have a significant impact on message‐passing performance over Ethernet. The implementation of this model in a firmware reveals that it may not be as efficient as some manually tuned micro‐benchmarks. But our multiqueue receive stack generally performs better than the original single queue stack, especially on large communication patterns where multiple processes are involved and manual binding is difficult. Copyright © 2010 John Wiley & Sons, Ltd. Brice Goglin |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | High-performance message-passing over generic Ethernet hardware with Open-MX
Brice Goglin |
Parallel Comput. | 1 |
| 2010 | Structuring the execution of OpenMP applications for multicore architecturesabstractThe now commonplace multi-core chips have introduced, by design, a deep hierarchy of memory and cache banks within parallel computers as a tradeoff between the user friendliness of shared memory on the one side, and memory access scalability and efficiency on the other side. However, to get high performance out of such machines requires a dynamic mapping of application tasks and data onto the underlying architecture. Moreover, depending on the application behavior, this mapping should favor cache affinity, memory bandwidth, computation synchrony, or a combination of these. The great challenge is then to perform this hardware-dependent mapping in a portable, abstract way. To meet this need, we propose a new, hierarchical approach to the execution of OpenMP threads onto multicore machines. Our ForestGOMP runtime system dynamically generates structured trees out of OpenMP programs. It collects relationship information about threads and data as well. This information is used together with scheduling hints and hardware counter feedback by the scheduler to select the most appropriate threads and data distribution. ForestGOMP features a highlevel platform for developing and tuning portable threads schedulers. We present several applications for which we developed specific scheduling policies that achieve excellent speedups on 16-core machines. François Broquedis, Olivier Aumage, Brice Goglin, Samuel Thibault, Pierre-André Wacrenier, Raymond Namyst |
IPDPS | 3 |
| 2010 | hwloc: A Generic Framework for Managing Hardware Affinities in HPC ApplicationsabstractThe increasing numbers of cores, shared caches and memory nodes within machines introduces a complex hardware topology. High-performance computing applications now have to carefully adapt their placement and behavior according to the underlying hierarchy of hardware resources and their software affinities. We introduce the Hardware Locality (hwloc) software which gathers hardware information about processors, caches, memory nodes and more, and exposes it to applications and runtime systems in a abstracted and portable hierarchical manner. hwloc may significantly help performance by having runtime systems place their tasks or adapt their communication strategies depending on hardware affinities. We show that hwloc can already be used by popular high-performance OpenMP or MPI software. Indeed, scheduling OpenMP threads according to their affinities or placing MPI processes according to their communication patterns shows interesting performance improvement thanks to hwloc. An optimized MPI communication strategy may also be dynamically chosen according to the location of the communicating processes in the machine and its hardware characteristics. François Broquedis, Jérôme Clet-Ortega, Stéphanie Moreaud, Nathalie Furmento, Brice Goglin, Guillaume Mercier, Samuel Thibault, Raymond Namyst |
PDP | 5 |
| 2010 | Adaptive MPI Multirail Tuning for Non-uniform Input/Output Access
Stéphanie Moreaud, Brice Goglin, Raymond Namyst |
EuroMPI | 2 |
| 2009 | Finding a tradeoff between host interrupt load and MPI latency over EthernetabstractAchieving high-performance message passing on top of generic Ethernet hardware suffers from the NIC interrupt-driven model where coalescing is usually involved. We present an in-depth study of the impact of interrupt coalescing on the Open-MX performance. It shows that disabling coalescing may not be relevant for most metrics except small-message latency. Two new coalescing strategies are then presented so as to efficiently support both latency-friendly and coalescing-friendly workloads thanks to the NIC looking at Open-MX messages and streams before deciding when to raise interrupts. The implementation of these strategies in the firmware of Myri-10G NICs shows that Open-MX is now able to achieve a low small-message latency, a high large-message throughput, and a satisfying message rate without having to manually tune the coalescing delay depending on the benchmark. Real application performance evaluation further shows that our modifications even improve the NAS parallel benchmark IS execution time by 7-8% thanks to our NIC firmware raising up to 20% of additional interrupts at the correct time. Brice Goglin, Nathalie Furmento |
CLUSTER | 1 |
| 2009 | NIC-Assisted Cache-Efficient Receive Stack for Message Passing over Ethernet
Brice Goglin |
Euro-Par | 1 |
| 2009 | Cache-Efficient, Intranode, Large-Message MPI Communication with MPICH2-NemesisabstractThe emergence of multicore processors raises the need to efficiently transfer large amounts of data between local processes. MPICH2 is a highly portable MPI implementation whose large-message communication schemes suffer from high CPU utilization and cache pollution because of the use of a double-buffering strategy, common to many MPI implementations. We introduce two strategies offering a kernel-assisted, single-copy model with support for noncontiguous and asynchronous transfers. The first one uses the now widely available vmsplice Linux system call; the second one further improves performance thanks to a custom kernel module called KNEM. The latter also offers I/OAT copy offload, which is dynamically enabled depending on both hardware cache characteristics and message size. These new solutions outperform the standard transfer method in the MPICH2 implementation when no cache is shared between the processing cores or when very large messages are being transferred. Collective communication operations show a dramatic improvement, and the IS NAS parallel benchmark shows a 25% speedup and better cache efficiency. Darius Buntinas, Brice Goglin, David Goodell, Guillaume Mercier, Stéphanie Moreaud |
ICPP | 2 |
| 2009 | Decoupling memory pinning from the application with overlapped on-demand pinning and MMU notifiersabstractHigh-performance cluster networks achieve very high throughput thanks to zero-copy techniques that require pinning of application buffers in physical memory. The Open-MX stack implements message passing over generic Ethernet hardware with similar needs. We present the design of an innovative pinning model in Open-MX based on the decoupling of memory pinning from the application. This idea eases the implementation of a reliable pinning cache in the kernel and enables full overlap of pinning with communication. The pinning cache enables performance improvement when the application reuses the same buffers multiple times, while overlapped pinning is also applicable to other applications. Performance evaluation shows that both these optimizations bring from 5 up to 20% throughput improvements depending on the host and network performance. Brice Goglin |
IPDPS | 1 |
| 2009 | Enabling high-performance memory migration for multithreaded applications on LINUXabstractAs the number of cores per machine increases, memory architectures are being redesigned to avoid bus contention and sustain higher throughput needs. The emergence of Non-Uniform Memory Access (NUMA) constraints has caused affinities between threads and buffers to become an important decision criterion for schedulers. Memory migration dynamically enables the joint distribution of work and data across the machine but requires high-performance data transfers as well as a convenient programming interface. We present improvements of the LINUX migration primitives and the implementation of a Next-touch policy in the kernel to provide multithreaded applications with an easy way to dynamically maintain thread-data affinity. Microbenchmarks show that our work enables a high-performance, synchronous and lazy memory migration within multithreaded applications. A threaded LU factorization then reveals the large improvement that our Next-touch policy model may bring in applications with complex access patterns. Brice Goglin, Nathalie Furmento |
IPDPS | 1 |
| 2009 | High Throughput Intra-Node MPI Communication with Open-MXabstractThe increasing number of cores per node in high-performance computing requires an efficient intra-node MPI communication subsystem. Most existing MPI implementations rely on two copies across a shared memory-mapped file. Open-MX offers a single-copy mechanism that is tightly integrated in its regular communication stack, making it transparently available to the MX backend of many MPI layers. We describe this implementation and its offloaded copy backend using I/OAT hardware. Memory pinning requirements are then discussed, and overlapped pinning is introduced to enable the start of Open-MX intra-node data transfer earlier. Performance evaluation shows that this local communication stack performs better than MPICH2 and Open-MPI for large messages, reaching up to 70% better throughput in micro-benchmarks when using I/OAT copy offload. Thanks to a single-copy being involved, the Open-MX intra-node communication throughput also does not heavily depend on cache sharing between processing cores, making these performance improvements easier to observe in real applications. Brice Goglin |
PDP | 1 |
| 2008 | Improving message passing over Ethernet with I/OAT copy offload in Open-MXabstractOpen-MX is a new message passing layer implemented on top of the generic Ethernet stack of the Linux kernel. Open-MX works on all Ethernet hardware, but it suffers from expensive memory copy requirements on the receiver side due to the hardwarepsilas inability to deposit messages directly in the target application buffers. Brice Goglin |
CLUSTER | 1 |
| 2008 | Design and implementation of Open-MX: High-performance message passing over generic Ethernet hardwareabstractOpen-MX is a new message passing layer implemented on top of the generic Ethernet stack of the Linux kernel. It provides high-performance communication on top of any Ethernet hardware while exhibiting the Myrinet Express application interface. Open-MX also enables wire- interoperability with Myricom's MXoE hosts. This article presents the design of the Open-MX stack which reproduces the MX firmware in a Linux driver. MPICH-MX and PVFS2 layers are already able to work flawlessly on Open-MX. The first performance evaluation shows interesting latency and bandwidth results on 1 and 10 gigabit hardware. Brice Goglin |
IPDPS | 1 |
| 2005 | An Efficient Network API for in-Kernel Applications in ClustersabstractRunning parallel applications on clusters with highspeed local networks requires fast communication between computing nodes but also low latency and high bandwidth file access. However, the application programming interfaces of high-speed local networks were designed for MPI communication and do not always meet the requirements of other applications like distributed file systems. In this paper, we explore several solutions to improve the use of high-speed network for in-kernel applications. Distributed file systems implemented on top of the GM interface of MYRINET are first examined to demonstrate how hard it is to get an efficient interaction between such applications and the network. Then, we propose solutions to simplify and improve this interaction and integrate them into the kernel interface of the new MYRINET driver, MX. Performance comparisons between MX and GM, and their usage in both a distributed file system and a zero-copy protocol show nice improvements. Moreover, we are able to improve the performance of the flexible kernel API we designed in MX that allows to remove some intermediate copy Brice Goglin, Olivier Glück, Pascale Vicat-Blanc Primet |
CLUSTER | 1 |
| 2004 | Performance Analysis of Remote File System Access over a High-Speed Local NetworkabstractSummary form only given. We study the performance of remote file system access in a cluster environment. Our experimental lightweight system is called ORFA and aims at providing high performance for general purpose file access. Using a simple protocol allows us to get the best performance from the underlying communication subsystem. Our user-level implementation avoids several kernel constraints and thus provides better performance than NFS. Moreover, we explore several optimization techniques to reduce the server overhead, which usually is the bottleneck. Performance evaluation shows that data access may saturate the network link while the server remains scalable. Brice Goglin, Loïc Prylli |
IPDPS | 1 |
| 2004 | Optimizations of Client's Side Communications in a Distributed File System within a Myrinet ClusterabstractHigh performance applications running on high-speed interconnects require both efficient communication between computing nodes and fast access to the storage system. Making the most of these networks to access remote files requires a good interaction between their highly specific software interface and the special requirements of distributed file systems. We show how non-buffered and buffered remote file access may be improved by modifying the network programming interface and firmware and by adding the required infrastructure in the operating system. Our modifications in Myrinet/GM show no performance penalty while the network usage in our ORFA (optimized remote file access) protocol is improved. Brice Goglin, Loïc Prylli, Olivier Glück |
LCN | 1 |