VLDB 2026 Research / reviewers in the wild / expert
Emmanuel Jeannot
dblp:46/5717
· DBLP profile ↗
71ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0002-3956-2997ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 13 first-author · 16 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Looking for (Genomic) Needles in a Haystack: Sparsity-Driven Search for Identifying Correlated Genetic Mutations in Cancer
Ritvik Prabhu, Emil Vatai, Bernard Moussad, Emmanuel Jeannot, Ramu Anandakrishnan, Wu-chun Feng, Mohamed Wahib |
IPDPS | 4 |
| 2025 | Performance Projection for Design-Space Exploration on future HPC ArchitecturesabstractTo address the growing need for performance from future HPC machines, their processor designs are constantly evolving. Assessing the impact of changes in hardware, software stack, and applications on performance is crucial in every step of a codesign process. Here, we propose a performance projection workflow to facilitate the exploration of design space for multicore nodes and multi-threaded applications. For this purpose, we analyze the architectural efficiency of an accessible source machine and determine the maximum sustainable flop/s performance of a hypothetical target machine based on its software stack on a per-thread basis. Finally, we use these characterizations to project the performance evolution from the source machine to the target machine. In this work, we assess the strengths and weaknesses of our approach by integrating it into the Fugaku-Next Feasibility Study. We compare the accuracy and overhead of our approach with the gem5 cycle-level simulations and a fast exploration methodology based on Machine Code Analyzer (MCA), using NAS Parallel benchmarks and CCS-QCD, a quantum chromodynamics miniapp. The study demonstrates that, compared to gem5, our approach has a prediction deviation of 5% for most cases and up to 30% for extreme cases. Additionally, it exhibits an execution overhead an order of magnitude bigger than MCA but orders of magnitude smaller than gem5. Finally, we demonstrate our approach's capability to study larger scale and more representative applications than gem5, such as QWS and Genesis, two applications of RIKEN optimized for Fugaku. Clément Gavoille, Hugo Taboada, Jens Domke, Brice Goglin, Emmanuel Jeannot |
IPDPS | 5 |
| 2024 | Phase-Based Data Placement Optimization in Heterogeneous MemoryabstractWhile scientific applications show increasing demand for memory speed and capacity, the performance gap between compute cores and the memory subsystem continues to spread. In response, heterogeneous memory systems integrating high-bandwidth memory (HBM) and non-volatile memory (NVM) alongside traditional DRAM on the CPU side are gaining traction. Despite the potential benefits of optimized memory selection for improved performance and efficiency, adapting applications to leverage diverse memory types often requires extensive modifications. Moreover, applications often comprise multiple execution phases with varying data access patterns. Since the capacity of the “fastest” memory is limited, relying solely on fixed data placement decisions may not yield optimal performance. Thus, considering allocation lifetimes and dynamically migrating data between memory types becomes imperative to ensure that performance-critical data for each phase resides in fast memory. To address these challenges, we developed a workflow incorporating memory access profiling, optimization techniques and a runtime system, which selects initial data placement for allocations and performs data migration during execution, considering the platform's memory subsystem characteristics and capacities. We formalize the optimization problems for initial and phase-based data placement and propose heuristics derived from memory profiling metrics to solve it. Additionally, we outline the implementation of these approaches, including allocation interception to enforce placement decisions. Experiments conducted with several applications on an Intel Ice Lake$(\text{DRAM}+\text{NVM})$and Sapphire Rapids$(\text{HBM}+\text{DRAM})$system demonstrate that our methodology can effectively bridge the performance gap between slow and fast memory in heterogeneous memory environments. Jannis Klinkenberg, Clément Foyer, Pierre Clouzet, Brice Goglin, Emmanuel Jeannot, Christian Terboven, Anara Kozhokanova |
CLUSTER | 5 |
| 2024 | Tracing task-based runtime systems: Feedbacks from the StarPU caseabstractSummary Given the complexity of current supercomputers and applications, being able to trace application executions to understand their behavior is not a luxury. As constraints, tracing systems have to be as little intrusive as possible in the application code and performances, and be precise enough in the collected data. In this article, we present how works the tracing system used by the task‐based runtime systemStarPU. We study the different sources of performance overhead coming from the tracing system and how to reduce these overheads. Then, we evaluate the accuracy of distributed traces with different clock synchronization techniques. Finally, we summarize our experiments and conclusions with the lessons we learned to efficiently trace applications, and the list of characteristics each tracing system should feature to be competitive. The reported experiments and implementation details comprise a feedback of integrating into a task‐based runtime system state‐of‐the‐art techniques to efficiently and precisely trace application executions. We highlight the points every application developer or end‐user should be aware of to seamlessly integrate a tracing system or just trace application executions. Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | Adding topology and memory awareness in data aggregation algorithmsabstractWith the growing gap between computing power and the ability of large-scale systems to ingest data, I/O is becoming the bottleneck for many scientific applications. Improving read and write performance thus becomes decisive, and requires consideration of the complexity of architectures. In this paper, we introduce TAPIOCA, an architecture-aware data aggregation library. TAPIOCA offers an optimized implementation of the two-phase I/O scheme for collective I/O operations, taking advantage of the many levels of memory and storage that populate modern HPC systems, and leveraging network topology. We show that TAPIOCA can significantly improve the I/O bandwidth of synthetic benchmarks and I/O kernels of scientific applications running on leading supercomputers. For example, on HACC-IO, a cosmology code, TAPIOCA improves data writing by a factor of 13 on nearly a third of the target supercomputer. Francois Tessier, Venkatram Vishwanath, Emmanuel Jeannot |
Future Gener. Comput. Syst. | 3 |
| 2023 | Adaptive multi-tier intelligent data manager for ExascaleabstractThe main objective of the ADMIRE project1 is the creation of an active I/O stack that dynamically adjusts computation and storage requirements through intelligent global coordination, the elasticity of computation and I/O, and the scheduling of storage resources along all levels of the storage hierarchy, while offering quality-of-service (QoS), energy efficiency, and resilience for accessing extremely large data sets in very heterogeneous computing and storage environments. We have developed a framework prototype that is able to dynamically adjust computation and storage requirements through intelligent global coordination, separated control, and data paths, the malleability of computation and I/O, the scheduling of storage resources along all levels of the storage hierarchy, and scalable monitoring techniques. The leading idea in ADMIRE is to co-design applications with ad-hoc storage systems that can be deployed with the application and adapt their computing and I/O behaviour on runtime, using malleability techniques, to increase the performance of applications and the throughput of the applications. Jesús Carretero 0001, Francisco Javier García Blas, Marco Aldinucci, Jean-Baptiste Besnard, Jean-Thomas Acquaviva, André Brinkmann, Marc-Andre Vef, Emmanuel Jeannot, Alberto Miranda, Ramon Nou, Morris Riedel, Massimo Torquati, Felix Wolf 0001 |
CF | 8 |
| 2023 | H2M: Exploiting Heterogeneous Shared Memory ArchitecturesabstractOver the past decades, the performance gap between the memory subsystem and compute capabilities continued to spread. However, scientific applications and simulations show increasing demand for both memory speed and capacity. To tackle these demands, new technologies such as high-bandwidth memory (HBM) or non-volatile memory (NVM) emerged, which are usually combined with classical DRAM. The resulting architecture is a heterogeneous memory system in which no single memory is “best”. HBM is smaller but offers higher bandwidth than DRAM, whereas NVM provides larger capacity than DRAM at a reasonable cost and less energy consumption. Despite that, in several cases, DRAM still offers the best latency out of all three technologies. In order to use different kinds of memory, applications typically have to be modified to a great extent. Consequently, vendor-agnostic solutions are desirable. First, they should offer the functionality to identify kinds of memory, and second, to allocate data on it. In addition, because memory capacities may be limited, decisions about data placement regarding the different memory kinds have to be made. Finally, in making these decisions, changes over time in data that is accessed, and the actual access pattern, should be considered for initial data placement and be respected in data migration at run-time. In this paper, we introduce a new methodology that aims to provide portable tools and methods for managing data placement in systems with heterogeneous memory. Our approach allows programmers to provide traits (hints) for allocations that describe how data is used and accessed. Combined with characteristics of the platforms’ memory subsystem, these traits are exploited by heuristics to decide where to place data items. We also discuss methodologies for analyzing and identifying memory access characteristics of existing applications, and for recommending allocation traits. In our evaluation, we conduct experiments with several kernels and two proxy applications on Intel Knights Landing (HBM + DRAM) and Intel Ice Lake with Intel Optane DC Persistent Memory (DRAM + NVM) systems. We demonstrate that our methodology can bridge the performance gap between slow and fast memory by applying heuristics for initial data placement. Jannis Klinkenberg, Anara Kozhokanova, Christian Terboven, Clément Foyer, Brice Goglin, Emmanuel Jeannot |
Future Gener. Comput. Syst. | 6 |
| 2023 | An introspection monitoring library to improve MPI communication time
Emmanuel Jeannot, Richard Sartori |
J. Supercomput. | 1 |
| 2022 | One core dedicated to MPI nonblocking communication progression? A model to assess whether it is worth itabstractOverlapping communications with computation is an efficient way to amortize the cost of communications of an HPC application. To do so, it is possible to utilize MPI nonblocking primitives so that communications run in back-ground alongside computation. However, these mechanisms rely on communications actually making progress in the background, which may not be true for all MPI libraries. Some MPI libraries leverage a core dedicated to communications to ensure communication progression. However, taking a core away from the application for such purpose may have a negative impact on the overall execution time. It may be difficult to know when such dedicated core is actually helpful. In this paper, we propose a model for the performance of applications using MPI nonblocking primitives running on top of an MPI library with a dedicated core for communications. This model is used to understand the compromise between computation slowdown due to the communication core not being available for computation, and the communication speed-up thanks to the dedicated core; evaluate whether nonblocking communication is actually obtaining the expected performance in the context of the given application; predict the performance of a given application if ran with a dedicated core. We describe the performance model and evaluate it on different applications. We compare the predictions of the model with actual executions. Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Florian Reynier |
CCGRID | 3 |
| 2022 | H2M: Towards Heuristics for Heterogeneous MemoryabstractFor the past years, scientific applications and simulations show increasing demand for both memory speed and capacity. The performance gap between compute units and the memory subsystem continues to spread which led to redesigns and the emergence of new technologies. Recent architectures already comprise, next to classical DRAM, portions of High Bandwidth Memory (HBM) that has less capacity than DRAM and is solving only one of the requirements. The newly introduced Non- Volatile Memory (NVM) shows performance closer to DRAM, while providing terabytes of capacity, consuming less power and having a better price per byte ratio. Clément Foyer, Brice Goglin, Emmanuel Jeannot, Jannis Klinkenberg, Anara Kozhokanova, Christian Terboven |
CLUSTER | 3 |
| 2022 | Relative Performance Projection on Arm Architectures
Clément Gavoille, Hugo Taboada, Patrick Carribault, Fabrice Dupros, Brice Goglin, Emmanuel Jeannot |
Euro-Par | 6 |
| 2022 | A methodology for assessing computation/communication overlap of MPI nonblocking collectivesabstractSummary By allowing computation/communication overlap, MPI nonblocking collectives (NBC) are supposed to improve application scalability and performance. However, it is known that to actually get overlap, the MPI library has to implement progression mechanisms in software or rely on the network hardware. These mechanisms may be present or not, adequate or perfectible, they may have an impact on communication performance or may interfere with computation by stealing CPU cycles. From a user point of view, assessing and understanding the behavior of an MPI library concerning computation/communication overlap is difficult. In this article, we propose a methodology to assess the computation/communication overlap of NBC. We propose new metrics to measure how much communication and computation do overlap, and to evaluate how they interfere with each other. We integrate these metrics into a complete methodology. We compare our methodology with state of the art metrics and benchmarks, and show that ours provides more meaningful informations. We perform experiments on a large panel of MPI implementations and network hardware and show when and why overlap is efficient, nonexistent or even degrades performance. Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Florian Reynier |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Process mapping on any topology with TopoMatch
Emmanuel Jeannot |
J. Parallel Distributed Comput. | 1 |
| 2021 | READYS: A Reinforcement Learning Based Strategy for Heterogeneous Dynamic SchedulingabstractIn this paper, we propose READYS, a reinforcement learning algorithm for the dynamic scheduling of computations modeled as a Directed Acyclic Graph (DAGs). Our goal is to develop a scheduling algorithm in which allocation and scheduling decisions are made at runtime, based on the state of the system, as performed in runtime systems such as StarPU or ParSEC. Reinforcement Learning is a natural candidate to achieve this task, since its general principle is to build step by step a strategy that, given the state of the system (the state of the resources and a view of the ready tasks and their successors in our case), makes a decision to optimize a global criterion. Moreover, the use of Reinforcement Learning is natural in a context where the duration of tasks (and communications) is stochastic. We propose READYS that combines Graph Convolutional Networks (GCN) with an Actor-Critic Algorithm (A2C): it builds an adaptive representation of the scheduling problem on the fly and learns a scheduling strategy, aiming at minimizing the makespan. A crucial point is that READYS builds a general scheduling strategy which is neither limited to only one specific application or task graph nor one particular problem size, and that can be used to schedule any DAG. We focus on different types of task graphs originating from linear algebra factorization kernels (CHOLESKY, LU, QR) and we consider heterogeneous platforms made of a few CPUs and GPUs. We first propose to analyze the performance of READYS when learning is performed on a given (platform, kernel, problem size) combination. Using simulations, we show that the scheduling agent obtains performances very similar or even superior to algorithms from the literature, and that it is especially powerful when the scheduling environment contains a lot of uncertainty. We additionally demonstrate that our agent exhibits very promising generalization capabilities. To the best of our knowledge, this is the first paper which shows that reinforcement learning can really be used for dynamic DAG scheduling on heterogeneous resources. Nathan Grinsztajn, Olivier Beaumont, Emmanuel Jeannot, Philippe Preux |
CLUSTER | 3 |
| 2021 | Interferences between Communications and Computations in Distributed HPC SystemsabstractInternational audience Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher |
ICPP | 2 |
| 2021 | An international survey on MPI users
Atsushi Hori, Emmanuel Jeannot, George Bosilca, Takahiro Ogura, Balazs Gerofi, Yutaka Ishikawa |
Parallel Comput. | 2 |
| 2020 | Using Dynamic Broadcasts to Improve Task-Based Runtime Performances
Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Euro-Par | 2 |
| 2020 | Mapping and scheduling HPC applications for optimizing I/OabstractIn HPC platforms, concurrent applications are sharing the same file system. This can lead to conflicts, especially as applications are more and more data intensive. I/O contention can represent a performance bottleneck. The access to bandwidth can be split in two complementary yet distinct problems. The mapping problem and the scheduling problem. The mapping problem consists in selecting the set of applications that are in competition for the I/O resource. The scheduling problem consists then, given I/O requests on the same resource, in determining the order to these accesses to minimize the I/O time. In this work we propose to couple a novel bandwidth-aware mapping algorithm to I/O list-scheduling policies to develop a cross-layer optimization solution. Jesús Carretero 0001, Emmanuel Jeannot, Guillaume Pallez, David E. Singh, Nicolas Vidal 0001 |
ICS | 2 |
| 2019 | Towards Portable Online Prediction of Network Utilization Using MPI-Level Monitoring
Shu-Mei Tseng, Bogdan Nicolae, George Bosilca, Emmanuel Jeannot, Aparna Chandramowlishwaran, Franck Cappello |
Euro-Par | 4 |
| 2019 | Data and Thread Placement in NUMA Architectures: A Statistical Learning ApproachabstractNowadays, NUMA architectures are common in compute-intensive systems. Achieving high performance for multi-threaded application requires both a careful placement of threads on computing units and a thorough allocation of data in memory. Finding such a placement is a hard problem to solve, because performance depends on complex interactions in several layers of the memory hierarchy. In this paper we propose a black-box approach to decide if an application execution time can be impacted by the placement of its threads and data, and in such a case, to choose the best placement strategy to adopt. We show that it is possible to reach near-optimal placement policy selection. Furthermore, solutions work across several recent processor architectures and decisions can be taken with a single run of low overhead profiling. Nicolas Denoyelle, Brice Goglin, Emmanuel Jeannot, Thomas Ropars |
ICPP | 3 |
| 2019 | Modeling Non-Uniform Memory Access on Large Compute Nodes with the Cache-Aware Roofline ModelabstractNUMA platforms, emerging memory architectures with on-package high bandwidth memories bring new opportunities and challenges to bridge the gap between computing power and memory performance. Heterogeneous memory machines feature several performance trade-offs, depending on the kind of memory used, when writing or reading it. Finding memory performance upper-bounds subject to such trade-offs aligns with the numerous interests of measuring computing system performance. In particular, representing applications performance with respect to the platform performance bounds has been addressed in the state-of-the-art Cache-Aware Roofline Model (CARM) to troubleshoot performance issues. In this paper, we present a Locality-Aware extension (LARM) of the CARM to model NUMA platforms bottlenecks, such as contention and remote access. On top of this, the new contribution of this paper is the design and validation of a novel hybrid memory bandwidth model. This new hybrid model quantifies the achievable bandwidth upper-bound under above-described trade-offs with less than 3 percent error. Hence, when comparing applications performance with the maximum attainable performance, software designers can now rely on more accurate information. Nicolas Denoyelle, Brice Goglin, Aleksandar Ilic, Emmanuel Jeannot, Leonel Sousa |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2018 | Process Affinity, Metrics and Impact on Performance: An Empirical StudyabstractProcess placement, also called topology mapping, is a well-known strategy to improve parallel program execution by reducing the communication cost between processes. It requires two inputs: the topology of the target machine and a measure of the affinity between processes. In the literature, the dominant affinity measure is the communication matrix that describes the amount of communication between processes. The goal of this paper is to study the accuracy of the communication matrix as a measure of affinity. We have done an extensive set of tests with two fat-tree machines and a 3d-torus machine to evaluate several hypotheses that are often made in the literature and to discuss their validity. First, we check the correlation between algorithmic metrics and the performance of the application. Then, we check whether a good generic process placement algorithm never degrades performance. And finally, we see whether the structure of the communication matrix can be used to predict gain. Cyril Bordage, Emmanuel Jeannot |
CCGrid | 2 |
| 2018 | Dynamic Placement of Progress Thread for Overlapping MPI Non-blocking Collectives on Manycore Processor
Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Marc Pérache, Hugo Taboada |
Euro-Par | 3 |
| 2018 | The Twenty Sixth International Heterogeneity in Computing Workshop (HCW) and to the Fifteenth International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar)abstracteditorial Jorge G. Barbosa, Emmanuel Jeannot |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | Hardware topology management in MPI applications through hierarchical communicators
Brice Goglin, Emmanuel Jeannot, Farouk Mansouri, Guillaume Mercier |
Parallel Comput. | 2 |
| 2017 | Automatic, Abstracted and Portable Topology-Aware Thread PlacementabstractEfficiently programming shared-memory machines is a difficult challenge because mapping application threads onto the memory hierarchy has a strong impact on the performance. However, optimizing such thread placement is difficult: architectures become increasingly complex and application behavior changes with implementations and input parameters, e.g problem size and number of threads. In this work, we propose a fully automatic, abstracted and portable affinity module. It produces and implements an optimized affinity strategy that combines knowledge about application characteristics and the platform topology. Implemented in the back-end of our runtime system (ORWL), our approach was used to enhance the performance and the scalability of several unmodified ORWL-coded applications: matrix multiplication, a 2D stencil (Livermore Kernel 23), and a video tracking real world application. On two SMP machines with quite different hardware characteristics, our tests show spectacular performance improvements for these unmodified application codes due to a dramatic decrease of cache misses and pipeline stalls. A comparison to reference implementations using OpenMP confirms this performance gain of almost one order of magnitude. Jens Gustedt, Emmanuel Jeannot, Farouk Mansouri |
CLUSTER | 2 |
| 2017 | TAPIOCA: An I/O Library for Optimized Topology-Aware Data Aggregation on Large-Scale SupercomputersabstractReading and writing data efficiently from storage system is necessary for most scientific simulations to achieve good performance at scale. Many software solutions have been developed to decrease the I/O bottleneck. One well-known strategy, in the context of collective I/O operations, is the two-phase I/O scheme. This strategy consists of selecting a subset of processes to aggregate contiguous pieces of data before performing reads/writes. In this paper, we present TAPIOCA, an MPI-based library implementing an efficient topology-aware two-phase I/O algorithm. We show how TAPIOCA can take advantage of double-buffering and one-sided communication to reduce as much as possible the idle time during data aggregation. We also introduce our cost model leading to a topology-aware aggregator placement optimizing the movements of data. We validate our approach at large scale on two leadership-class supercomputers: Mira (IBM BG/Q) and Theta (Cray XC40). We present the results obtained with TAPIOCA on a micro-benchmark and the I/O kernel of a large-scale simulation. On both architectures, we show a substantial improvement of I/O performance compared with the default MPI I/O implementation. On BG/Q+GPFS, for instance, our algorithm leads to a performance improvement by a factor of twelve while on the Cray XC40 system associated with a Lustre filesystem, we achieve an improvement of four. Francois Tessier, Venkatram Vishwanath, Emmanuel Jeannot |
CLUSTER | 3 |
| 2017 | Online Dynamic Monitoring of MPI Communications
George Bosilca, Clément Foyer, Emmanuel Jeannot, Guillaume Mercier, Guillaume Papauré |
Euro-Par | 3 |
| 2017 | Trends in Data Locality Abstractions for HPC SystemsabstractThe cost of data movement has always been an important concern in high performance computing (HPC) systems. It has now become the dominant factor in terms of both energy consumption and performance. Support for expression of data locality has been explored in the past, but those efforts have had only modest success in being adopted in HPC applications for various reasons. them However, with the increasing complexity of the memory hierarchy and higher parallelism in emerging HPC systems, locality management has acquired a new urgency. Developers can no longer limit themselves to low-level solutions and ignore the potential for productivity and performance portability obtained by using locality abstractions. Fortunately, the trend emerging in recent literature on the topic alleviates many of the concerns that got in the way of their adoption by application developers. Data locality abstractions are available in the forms of libraries, data structures, languages and runtime systems; a common theme is increasing productivity without sacrificing performance. This paper examines these trends and identifies commonalities that can combine various locality concepts to develop a comprehensive approach to expressing and managing data locality on future large-scale high-performance computing systems. Didem Unat, Anshu Dubey, Torsten Hoefler, John Shalf, Mark James Abraham, Mauro Bianco, Bradford L. Chamberlain, Romain Cledat, H. Carter Edwards, Hal Finkel, Karl Fürlinger, Frank Hannig, Emmanuel Jeannot, Amir Kamil, Jeff Keasler, Paul H. J. Kelly, Vitus J. Leung, Hatem Ltaief, Naoya Maruyama, Chris J. Newburn, Miquel Pericàs |
IEEE Trans. Parallel Distributed Syst. | 13 |
| 2016 | Optimizing Locality by Topology-Aware Placement for a Task Based Programming ModelabstractThe ordered read-write lock model (ORWL) is a modern framework that proposes high level abstractions for the decomposition of an application and for the management of synchronizations and communications. The implementation of the model reaches high performances thanks to a decentralized event-based runtime. In this paper, we propose to enrich ORWL by proposing a topology-aware placement module that is based on the Hardware Locality framework, HWLOC. The aim is double. On one hand we increase the abstraction and the portability of the framework, and on the other hand we enhance the performance of the model's runtime. We propose a placement policy, that takes the characteristics of the application, of the runtime and of the architecture into account. We validate and compare our approach with the Livermore kernel23 benchmarks. Jens Gustedt, Emmanuel Jeannot, Farouk Mansouri |
CLUSTER | 2 |
| 2016 | DKPN: A Composite Dataflow/Kahn Process Networks Execution ModelabstractTo address the high level of dynamism and variability in modern streaming applications (e.g. video decoding) as well as the difficulties in programming heterogeneous MPSoCs, we propose a novel execution model based upon both dataflow and Kahn process networks. This paper presents the semantics and properties of this hierarchical and parametric model, called DKPN. Parameters are classified and it is shown that hints can be derived to improve the execution. A scheduler framework and policies to back the model are also exposed. Experiments illustrate the benefits of our approach. Paul-Antoine Arras, Didier Fuin, Emmanuel Jeannot, Samuel Thibault |
PDP | 3 |
| 2016 | HeteroPar 2014, APCIE 2014, and TASUS 2014 Special IssueabstractThese workshops were organized by members of the Nesus Cost Action IC 1305: Network for Sustainable Ultrascale Computing, which is a follow-up of COST Actions IC0804 and IC0805 1. The goal of the NESUS Action is to establish an open European research network targeting sustainable solutions for ultrascale computing aiming at cross fertilization among HPC, large-scale distributed systems, and big data management. This network aims at contributing to glue disparate researchers working across different areas and provide a meeting ground for researchers in these separate areas to exchange ideas, to identify synergies, and to pursue common activities in research topics such as sustainable software solutions (applications and system software stack), data management, energy efficiency, and resilience. The selected papers cover very important scientific issues encountered nowadays such as the following: CPU/GPU execution, system-on-chip programming, parallel algorithms taking into account various constraints (energy, communication, etc.), programming models, and so on. We really hope that the reader will enjoy this high-quality issue, and we are sure that she/he will find it highly relevant to the state-of-the-art of today's heterogeneous and parallel computing. Jesús Carretero 0001, Raimondas Ciegis, Emmanuel Jeannot, Laurent Lefèvre, Gudula Rünger, Domenico Talia, Julius Zilinskas |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Correlation-Aware Heuristics for Evaluating the Distribution of the Longest Path Length of a DAG with Random WeightsabstractCoping with uncertainties when scheduling task graphs on parallel machines requires to perform non-trivial evaluations. When considering that each computation and communication duration is a random variable, evaluating the distribution of the critical path length of such graphs involves computing maximums and sums of possibly dependent random variables. The discrete version of this evaluation problem is known to be #P-hard. Here, we propose two heuristics, CorLCA and Cordyn, to compute such lengths. They approximate the input random variables and the intermediate ones as normal random variables, and they precisely take into account correlations with two distinct mechanisms: through lowest common ancestor queries for CorLCA and with a dynamic programming approach for Cordyn. Moreover, we empirically compare some classical methods from the literature and confront them to our solutions. Simulations on a large set of cases indicate that CorLCA and Cordyn constitute each a new relevant trade-off in terms of rapidity and precision. Louis-Claude Canon, Emmanuel Jeannot |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | SPAGHETtI: Scheduling/Placement Approach for Task-Graphs on HETerogeneous archItecture
Denis Barthou, Emmanuel Jeannot |
Euro-Par | 2 |
| 2014 | Process Placement in Multicore Clusters: Algorithmic Issues and Practical TechniquesabstractCurrent generations of NUMA node clusters feature multicore or manycore processors. Programming such architectures efficiently is a challenge because numerous hardware characteristics have to be taken into account, especially the memory hierarchy. One appealing idea to improve the performance of parallel applications is to decrease their communication costs by matching the communication pattern to the underlying hardware architecture. In this paper, we detail the algorithm and techniques proposed to achieve such a result: first, we gather both the communication pattern information and the hardware details. Then we compute a relevant reordering of the various process ranks of the application. Finally, those new ranks are used to reduce the communication costs of the application. Emmanuel Jeannot, Guillaume Mercier, Francois Tessier |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Communication and topology-aware load balancing in Charm++ with TreeMatchabstractProgramming multicore or manycore architectures is a hard challenge particularly if one wants to fully take advantage of their computing power. Moreover, a hierarchical topology implies that communication performance is heterogeneous and this characteristic should also be exploited. We developed two load balancers for Charm++ that take into account both aspects, depending on the fact that the application is compute-bound or communication-bound. This work is based on our TREEMATCH library that computes process placement in order to reduce an application communication costs based on the hardware topology. We show that the proposed load-balancing schemes manage to improve the execution times for the two aforementioned classes of parallel applications. Emmanuel Jeannot, Esteban Meneses, Guillaume Mercier, Francois Tessier, Gengbin Zheng |
CLUSTER | 1 |
| 2013 | List Scheduling in Embedded Systems under Memory ConstraintsabstractVideo decoding and image processing in embedded systems are subject to strong resource constraints, particularly in terms of memory. List-scheduling heuristics with static priorities (HEFT, SDC, etc.) being the often-cited solutions due to both their good performance and their low complexity, we propose a method aimed at introducing the notion of memory into them. Moreover, we show that through appropriate adjustment of task priorities and judicious resort to insertion-based policy, speedups up to 20% can be achieved. Lastly, we show that our technique allows to prevent deadlock and to substantially reduce the required memory footprint compared to classic list-scheduling heuristics. Paul-Antoine Arras, Didier Fuin, Emmanuel Jeannot, Arthur Stoutchinin, Samuel Thibault |
SBAC-PAD | 3 |
| 2012 | Optimizing performance and reliability on heterogeneous parallel systems: Approximation algorithms and heuristics
Emmanuel Jeannot, Erik Saule, Denis Trystram |
J. Parallel Distributed Comput. | 1 |
| 2011 | Towards Real-Time, Volunteer Distributed ComputingabstractMany large-scale distributed computing applications demand real-time responses by soft deadlines. To enable such real-time task distribution and execution on the volunteer resources, we previously proposed the design of the real-time volunteer computing platform called RT-BOINC. The system gives low O(1) worst-case execution time for task management operations, such as task scheduling, state transitioning, and validation. In this work, we present a full implementation RT-BOINC, adding new features including deadline timer and parameter-based admission control. We evaluate RT-BOINC at large scale using two real-time applications, namely, the games Go and Chess. The results of our case study show that RT-BOINC provides much better performance than the original BOINC in terms of average and worst-case response time, scalability and efficiency. Sangho Yi, Emmanuel Jeannot, Derrick Kondo, David P. Anderson |
CCGRID | 2 |
| 2011 | A Scheduling and Certification Algorithm for Defeating Collusion in Desktop GridsabstractBy exploiting idle time on volunteer machines, desktop grids provide a way to execute large sets of tasks with negligible maintenance and low cost. Although desktop grids are attractive for their scalability and low cost, relying on external resources may compromise the correctness of application execution due to the well-known unreliability of nodes. In this paper, we consider a very challenging threat model: correlated errors caused either by organized groups of cheaters that may collude to produce incorrect results, or by buggy or so-called "unofficial" clients. By using a previously described on-line algorithm for detecting collusion and characterizing the participant behaviors, we propose a scheduling and result certification algorithm that tackles collusion. Using several real-life traces, we show that our approach minimizes both replication overhead and the number of incorrectly certified results. Louis-Claude Canon, Emmanuel Jeannot, Jon B. Weissman |
ICDCS | 2 |
| 2011 | Improving MPI Applications Performance on Multicore Clusters with Rank Reordering
Guillaume Mercier, Emmanuel Jeannot |
EuroMPI | 2 |
| 2010 | Near-Optimal Placement of MPI Processes on Hierarchical NUMA Architectures
Emmanuel Jeannot, Guillaume Mercier |
Euro-Par (2) | 1 |
| 2010 | A dynamic approach for characterizing collusion in desktop gridsabstractBy exploiting idle time on volunteer machines, desktop grids provide a way to execute large sets of tasks with negligible maintenance and low cost. Although desktop grids are attractive for cost-conscious projects, relying on external resources may compromise the correctness of application execution due to the well-known unreliability of nodes. In this paper, we consider the most challenging threat model: organized groups of cheaters that may collude to produce incorrect results. We propose two on-line algorithms for detecting collusion and characterizing the participant behaviors. Using several real-life traces, we show that our approach is accurate and efficient in identifying collusion and in estimating group behavior. Louis-Claude Canon, Emmanuel Jeannot, Jon B. Weissman |
IPDPS | 2 |
| 2010 | Defining and controlling the heterogeneity of a cluster: The Wrekavoc tool
Louis-Claude Canon, Olivier Dubuisson, Jens Gustedt, Emmanuel Jeannot |
J. Syst. Softw. | 4 |
| 2010 | Evaluation and Optimization of the Robustness of DAG Schedules in Heterogeneous EnvironmentsabstractA schedule is said to be robust if it is able to absorb some degree of uncertainty in task or communication durations while maintaining a stable solution. This intuitive notion of robustness has led to a lot of different metrics and almost no heuristics. In this paper, we perform an experimental study of these different metrics and show how they are correlated to each other. Additionally, we propose different strategies for minimizing the makespan while maximizing the robustness: from an evolutionary metaheuristic (best solutions but longer computation time) to more simple heuristics making approximations (medium quality solutions but fast computation time). We compare these different approaches experimentally and show that we are able to find different approximations of the Pareto front for this bicriteria problem. Louis-Claude Canon, Emmanuel Jeannot |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Introduction
Emmanuel Jeannot, Ramin Yahyapour, Daniel Grosu, Helen D. Karatza |
Euro-Par | 1 |
| 2009 | Validating Wrekavoc: A tool for heterogeneity emulationabstractExperimental validation and testing of solutions designed for heterogeneous environment is a challenging issue. Wrekavoc is a tool for performing such validation. It runs unmodified applications on emulated multisite heterogeneous platforms. Therefore it downgrades the performance of the nodes (CPU and memory) and the interconnection network in a prescribed way. We report on new strategies to improve the accuracy of the network and memory models. Then, we present an experimental validation of the tool that compares executions of a variety of application code. The comparison of a real heterogeneous platform is done against the emulation of that platform with Wrekavoc. The measurements show that our approach allows for a close reproduction of the real measurements in the emulator. Olivier Dubuisson, Jens Gustedt, Emmanuel Jeannot |
IPDPS | 3 |
| 2008 | Bi-objective Approximation Scheme for Makespan and Reliability Optimization on Uniform Parallel Machines
Emmanuel Jeannot, Erik Saule, Denis Trystram |
Euro-Par | 1 |
| 2008 | Scheduling strategies for the bicriteria optimization of the robustness and makespanabstractIn this paper we study the problem of scheduling a stochastic task graph with the objective of minimizing the makespan and maximizing the robustness. As these two metrics are not equivalent, we need a bicriteria approach to solve this problem. Moreover, as computing these two criteria is very time consuming we propose different approaches: from an evolutionary meta-heuristic (best solutions but longer computation time) to more simple heuristics making approximations (bad quality solutions but fast computation time). We compare these different strategies experimentally and show that we are able to find different approximations of the Pareto front of this bicriteria problem. Louis-Claude Canon, Emmanuel Jeannot |
IPDPS | 2 |
| 2008 | Experimental validation of grid algorithms: A comparison of methodologiesabstractThe increasing complexity of available infrastructures with specific features (caches, hyper- threading, dual core, etc.) or with complex architectures (hierarchical, parallel, distributed, etc.) makes models either extremely difficult to build or intractable. Hence, it raises the question: how to validate algorithms if a realistic analytic analysis is not possible any longer? As for some other sciences (physics, chemistry, biology, etc.), the answer partly falls in experimental validation. Nevertheless, experiment in computer science is a difficult subject that opens many questions: what an experiment is able to validate? What is a "good experiments"? How to build an experimental environment that allows for "good experiments"? etc. In this paper we will provide some hints on this subject and show how some tools can help in performing "good experiments". More precisely we will focus on three main experimental methodologies, namely real-scale experiments (with an emphasis on PlanetLab and Grid'5000), Emulation (with an emphasis on Wrekavoc: http://wrekavoc.gforge.inria.fr) and simulation (with an emphasis on SimGRID and Grid-Sim). We will provide a comparison of these tools and methodologies from a quantitative but also qualitative point of view. Emmanuel Jeannot |
IPDPS | 1 |
| 2007 | A Comparison of robustness metrics for scheduling DAGs on heterogeneous systemsabstractA schedule is said robust if it is able to absorb some degree of uncertainty in tasks duration while maintaining a stable solution. This intuitive notion of robustness has led to a lot of different interpretations and metrics. However, no comparison of these different metrics have ever been preformed. In this paper, we perform an experimental study of these different metrics and show how they are correlated to each other in the case of task scheduling, with dependencies between tasks. Louis-Claude Canon, Emmanuel Jeannot |
CLUSTER | 2 |
| 2007 | Fast and Efficient Total Exchange on Two Clusters
Emmanuel Jeannot, Luiz Angelo Steffenel |
Euro-Par | 1 |
| 2007 | Topic 9 Parallel and Distributed Programming
Luc Moreau 0001, Emmanuel Jeannot, George Bosilca, Antonio Plaza |
Euro-Par | 2 |
| 2007 | Bi-objective scheduling algorithms for optimizing makespan and reliability on heterogeneous systemsabstractWe tackle the problem of scheduling task graphs onto a heterogeneous set of machines, where each processor has a probability of failure governed by an exponential law. The goal is to design algorithms that optimize both makespan and reliability. First, we provide an optimal scheduling algorithm for independent unitary tasks where the objective is to maximize the reliability subject to makespan minimization. For the bi-criteria case, we provide an algorithm that approximates the Pareto-curve. Next, for independent non-unitary tasks, we show that the product {failure rate}x {unitary instruction execution time} is crucial to distinguish processors in this context. Based on these results we are able to let the user choose a trade-off between reliability maximization and makespan minimization. For general task graphs we provide a method for converting scheduling heuristics on heterogeneous cluster into heuristics that take reliability into account. Here again, we show how we can help the user to select a trade-off between makespan and reliability. Jack J. Dongarra, Emmanuel Jeannot, Erik Saule, Zhiao Shi |
SPAA | 2 |
| 2006 | Modeling, Predicting and Optimizing Redistribution between Clusters on Low Latency NetworksabstractIn this paper we study the problem of scheduling messages between two parallel machines connected by a low latency network during a data redistribution. We compare two approaches. In the first approach no scheduling is performed. Since all the messages cannot be transmitted at the same time, the transport layer has to manage the congestion. In the second approach we use two higher-level scheduling algorithms proposed in our previous work [E. Jeannot et al., (2004)] called GGP and OGGP. The contribution of this paper is the following: we show that the redistribution time with scheduling is always better than the brute-force approach (up to 30%). As this speedup depends on the input redistribution pattern, we propose a modelization of the behavior of both approaches and show that we are able to accurately predict the redistribution time with or without scheduling and thus able to choose for each pattern whether or not to schedule the communications. Emmanuel Jeannot, Frédéric Wagner |
AINA (2) | 1 |
| 2006 | Robust task scheduling in non-deterministic heterogeneous computing systemsabstractThe paper addresses the problem of matching and scheduling of DAG-structured application to both minimize the makespan and maximize the robustness in a heterogeneous computing system. Due to the conflict of the two objectives, it is usually impossible to achieve both goals at the same time. We give two definitions of robustness of a schedule based on tardiness and miss rate. Slack is proved to be an effective metric to be used to adjust the robustness. We employ epsiv-constraint method to solve the bi-objective optimization problem where minimizing the makespan and maximizing the slack are the two objectives. Overall performance of a schedule considering both makespan and robustness is defined such that user have the flexibility to put emphasis on either objective. Experiment results are presented to validate the performance of the proposed algorithm Zhiao Shi, Emmanuel Jeannot, Jack J. Dongarra |
CLUSTER | 2 |
| 2006 | A Practical Approach of Diffusion Load Balancing Algorithms
Emmanuel Jeannot, Flavien Vernier |
Euro-Par | 1 |
| 2006 | A probabilistic approach for fault tolerant multiprocessor real-time schedulingabstractIn this paper we tackle the problem of scheduling a periodic real time system on identical multiprocessor platforms, moreover the tasks considered may fail with a given probability. For each task we compute its duplication rate in order to (1) given a maximum tolerated probability of failure, minimize the size of the platform such at least one replica of each job meets its deadline (and does not fail) using a variant of EDF namely EDF(k)or (2) given the size of the platform, achieve the best possible reliability with the same constraints. Thanks to our probabilistic approach, no assumption is made on the number of failures which can occur. We propose several approaches to duplicate tasks and we show that we are able to find solutions always very close to the optimal one Vandy Berten, Joël Goossens, Emmanuel Jeannot |
IPDPS | 3 |
| 2006 | Wrekavoc: a tool for emulating heterogeneityabstractComputer science and especially heterogeneous distributed computing is an experimental science. Simulation, emulation, or in-situ implementation are complementary methodologies to conduct experiments in this context. In this paper, we address the problem of defining and controlling the heterogeneity of a platform. We evaluate the proposed solution, called Wrekavoc, with micro-benchmark and by implementing algorithms of the literature. Louis-Claude Canon, Emmanuel Jeannot |
IPDPS | 2 |
| 2006 | On the Distribution of Sequential Jobs in Random Brokering for Heterogeneous Computational GridsabstractScheduling stochastic workloads is a difficult task. In order to design efficient scheduling algorithms for such workloads, it is required to have a good in-depth knowledge of basic random scheduling strategies. This paper analyzes the distribution of sequential jobs and the system behavior in heterogeneous computational grid environments where the brokering is done in such a way that each computing element has a probability to be chosen proportional to its number of CPUs and (new from the previous paper) its relative speed. We provide the asymptotic behavior for several metrics (queue-sizes, slowdowns, etc.) or, in some cases, an approximation of this behavior. We study these metrics for a variety of workload configurations (load, distribution, etc.). We compare our probabilistic analysis to simulations in order to validate our results. These results provide a good understanding of the system behavior for each metric proposed. This enables us to design advanced and efficient algorithms for more complex cases. Vandy Berten, Joël Goossens, Emmanuel Jeannot |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2006 | Messages Scheduling for Parallel Data Redistribution between ClustersabstractWe study the problem of redistributing data between clusters interconnected by a backbone. We suppose that at most k communications can be performed at the same time (the value of k depending on the characteristics of the platform). Given a set of messages, we aim at minimizing the total communication time assuming that communications can be preempted and that preemption comes with an extra cost. Our problem, called k-preemptive bipartite scheduling (KPBS) is proven to be NP-hard. We study its lower bound. We propose two 8/3-approximation algorithms with low complexity and fast heuristics. Simulation results show that both algorithms perform very well compared to the optimal solution and to the heuristics. Experimental results, based on an MPI implementation of these algorithms, show that both algorithms outperform a brute-force TCP-based solution, where no scheduling of the messages is performed Johanne Cohen, Emmanuel Jeannot, Nicolas Padoy, Frédéric Wagner |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2005 | Grid-enabling medical image analysisabstractDigital medical image processing is a promising application area for grids. Given the volume of data, the sensitivity of medical information, and the joint complexity of medical datasets and computations expected in clinical practice, the challenge is to fill the gap between the grid middleware and the requirements of clinical applications. The research project AGIR (Grid Analysis of Radiological Data) presented in this paper addresses this challenge through a combined approach: on one hand, leveraging the grid middleware through core grid medical services which target the requirements of medical data processing applications; on the other hand, grid-enabling a panel of applications ranging from algorithmic research to clinical applications. Cécile Germain, Vincent Breton, Patrick Clarysse, Yann Gaudeau, Tristan Glatard, Emmanuel Jeannot, Yannick Legré, Charles Loomis, Johan Montagnat, Jean-Marie Moureaux, Angel Osorio, Xavier Pennec, Romain Texier |
CCGRID | 6 |
| 2004 | Experimental Study of Multi-criteria Scheduling Heuristics for GridRPC Systems
Yves Caniou, Emmanuel Jeannot |
Euro-Par | 2 |
| 2004 | Efficient Scheduling Heuristics for GridRPC Systems
Yves Caniou, Emmanuel Jeannot |
ICPADS | 2 |
| 2004 | Two Fast and Efficient Message Scheduling Algorithms for Data Redistribution through a BackboneabstractSummary form only given. We study the problem of redistributing in parallel data between clusters interconnected by a backbone. This problem is a generalization of the well-known redistribution problem that appears in parallelism. We suppose that at most k communications can be performed at the same time (the value of k depending on the characteristics of the platform). We use the knowledge of the application in order to schedule the messages and perform a control of the congestion by ourselves. Previous results show that this problem is NP-complete. We propose and study two fast and efficient algorithms for this problem. We prove that these algorithms are 2-approximation algorithms. Simulation results show that both algorithms perform very well compared to the optimal solution. These algorithms have been implemented using MPI. Experimental results show that both algorithms outperform a brute-force TCP based solution, where no scheduling of the messages is performed. Emmanuel Jeannot, Frédéric Wagner |
IPDPS | 1 |
| 2004 | Compact DAG representation and its symbolic scheduling
Michel Cosnard, Emmanuel Jeannot |
J. Parallel Distributed Comput. | 2 |
| 2002 | Adaptive Online Data CompressionabstractQuickly transmitting large datasets in the context of distributed computing on wide area networks can be achieved by compressing data before transmission, However such an approach is not efficient when dealing with higher speed networks. Indeed, the time to compress a large file and to send it is greater than the time to send the uncompressed file. In this paper we explore and enhance an algorithm that allows us to overlap communications with compression and to automatically adapt the compression effort to currently available network and processor resources. Emmanuel Jeannot, Björn Knutsson, Mats Björkman |
HPDC | 1 |
| 2001 | SCILAB to SCILAB//: The OURAGAN project
Eddy Caron, Serge Chaumette, Sylvain Contassot-Vivier, Frédéric Desprez, Eric Fleury, Claude Gomez, Maurice Goursat, Martin Quinson, Emmanuel Jeannot, Dominique Lazure, Frédéric Lombard, Jean-Marc Nicod, Laurent Philippe 0001, Pierre Ramet, Jean Roman, Frank Rubi, Serge Steer, Frédéric Suter, Gil Utard |
Parallel Comput. | 9 |
| 1999 | SLC: Symbolic Scheduling for Executing Parameterized Task Graphs on MultiprocessorsabstractTask graph scheduling has been found effective in performance prediction and optimization of parallel applications. A number of static scheduling algorithms have been proposed for task graph execution on distributed memory machines. Such an approach cannot be adapted to changes in values of program parameters and the number of processors and also it cannot handle large task graphs. In this paper, we model parallel computation using parameterized task graphs which represent coarse-grain parallelism independent of the problem size. We present a scheduling algorithm for a parameterized task graph which first derives symbolic linear clusters and then assigns task clusters to processors. The runtime system executes clusters on each processor in a multi-threaded fashion. We evaluate our method using various compute-intensive kernels that can be found in scientific applications. Michel Cosnard, Emmanuel Jeannot |
ICPP | 2 |
| 1999 | Compact DAG Representation and Its Dynamic Scheduling
Michel Cosnard, Emmanuel Jeannot |
J. Parallel Distributed Comput. | 2 |
| 1998 | Symbolic Partitioning and Scheduling of Parameterized Task GraphsabstractThe DAG based task graph model has been found effective in scheduling for performance prediction and optimization of parallel applications. However the scheduling complexity and solution normally depend on the problem size. We propose a symbolic scheduling scheme for a parameterized task graph which models coarse grain DAG parallelism, independent of the problem size. The algorithm first derives symbolic clusters to a group of tasks in order to minimize communication while preserving parallelism, and then it evenly assigns task clusters to processors. The run time system executes clusters on each processor in a multithreaded fashion. The paper also presents preliminary experimental results to demonstrate the effectiveness of our techniques. Michel Cosnard, Emmanuel Jeannot |
ICPADS | 2 |