Denis Barthou

dblp:66/2364 · DBLP profile ↗
← Back
40ranked-venue papers
6as first author
7since 2021 · last 2026
0009-0000-8547-5395ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Energy-aware scheduling strategies for partially-replicable task chains on heterogeneous processors
Yacine Idouar, Adrien Cassagne, Laércio Lima Pilla, Julien Sopena, Manuel Bouyer, Diane Orhan, Lionel Lacassagne, Dimitri Galayko, Denis Barthou, Christophe Jégo
Parallel Comput.9
2025 Optimal scheduling algorithms for software-defined radio pipelined and replicated task chains on multicore architectures
abstract
Software-Defined Radio (SDR) represents a move from dedicated hardware to software implementations of digital communication standards. This approach offers flexibility, shorter time to market, maintainability , and lower costs, but it requires an optimized distribution tasks in order to meet performance requirements. Thus, we study the problem of scheduling SDR linear task chains of stateless and stateful tasks for streaming processing. We model this problem as a pipelined workflow scheduling problem based on pipelined and replicated parallelism on homogeneous resources. We propose an optimal dynamic programming solution and an optimal greedy algorithm named OTAC for maximizing throughput while also minimizing resource utilization . Moreover, the optimality of the proposed scheduling algorithm is proved. We evaluate our solutions and compare their execution times and schedules to other algorithms using synthetic task chains and an implementation of the DVB-S2 communication standard on the AFF3CT SDR Domain Specific Language . Our results demonstrate how OTAC quickly finds optimal schedules, leading consistently to better results than other algorithms, or equivalent results with fewer resources.
Diane Orhan, Laércio Lima Pilla, Denis Barthou, Adrien Cassagne, Olivier Aumage, Romain Tajan, Christophe Jégo, Camille Leroux
J. Parallel Distributed Comput.3
2025 Performance portability of generated cardiac simulation kernels through automatic dimensioning and load balancing on heterogeneous nodes
Vincent Alba, Olivier Aumage, Denis Barthou, Marie Christine Counilh, Amina Guermouche
J. Supercomput.3
2024 PolyTOPS: Reconfigurable and Flexible Polyhedral Scheduler
abstract
Polyhedral techniques have been widely used for automatic code optimization in low-level compilers and higher-level processes. Loop optimization is central to this technique, and several polyhedral schedulers like Feautrier, Pluto, isl and Tensor Scheduler have been proposed, each of them targeting a different architecture, parallelism model, or application scenario. The need for scenario-specific optimization is growing due to the heterogeneity of architectures. One of the most critical cases is represented by NPUs (Neural Processing Units) used for AI, which may require loop optimization with different objectives. Another factor to be considered is the framework or compiler in which polyhedral optimization takes place. Different scenarios, depending on the target architecture, compilation environment, and application domain, may require different kinds of optimization to best exploit the architecture feature set. We introduce a new configurable polyhedral scheduler, PolyTOPS, that can be adjusted to various scenarios with straightforward, high-level configurations. This scheduler allows the creation of diverse scheduling strategies that can be both scenario-specific (like state-of-the-art schedulers) and kernel-specific, breaking the concept of a one-size-fits-all scheduler approach. PolyTOPS has been used with isl and CLooG as code generators and has been integrated in MindSpore AKG deep learning compiler. Experimental results in different scenarios show good performance: a geomean speedup of 7.66x on MindSpore (for the NPU Ascend architecture) hybrid custom operators over isl scheduling, a geomean speedup up to 1.80× on PolyBench on different multicore architectures over Pluto scheduling. Finally, some comparisons with different state-of-the-art tools are presented in the PolyMage scenario.
Gianpietro Consolaro, Harenome Razanajato, Nelson Lossing, Nassim Tchoulak, Adilla Susungi, Artur Cesar Araujo Alves, Renwei Zhang, Denis Barthou, Corinne Ancourt, Cédric Bastoul
CGO9
2023 A DSEL for high throughput and low latency software-defined radio on multicore CPUs
abstract
Summary This article presents a new Domain Specific Embedded Language (DSEL) dedicated to Software‐Defined Radio (SDR). From a set of carefully designed components, it enables to build efficient software digital communication systems, able to take advantage of the parallelism of modern processor architectures, in a straightforward and safe manner for the programmer. In particular, proposed DSEL enables the combination of pipelining and sequence duplication techniques to extract both temporal and spatial parallelism from digital communication systems. We leverage the DSEL capabilities on a real use case: a fully digital transceiver for the widely used DVB‐S2 standard designed entirely in software. Through evaluation, we show how proposed software DVB‐S2 transceiver is able to get the most from modern, high‐end multicore CPU targets.
Adrien Cassagne, Romain Tajan, Olivier Aumage, Camille Leroux, Denis Barthou, Christophe Jégo
Concurr. Comput. Pract. Exp.5
2022 Exploring Scheduling Algorithms for Parallel Task Graphs: A Modern Game Engine Case Study
Mustapha Regragui, Baptiste Coye, Laércio Lima Pilla, Raymond Namyst, Denis Barthou
Euro-Par5
2022 MPI detach - Towards automatic asynchronous local completion
Joachim Jenke, Marc-André Hermanns, Matthias S. Müller, Van Man Nguyen, Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou
Parallel Comput.8
2019 Multi-valued Expression Analysis for Collective Checking
Pierre Huchant, Emmanuelle Saillard, Denis Barthou, Patrick Carribault
Euro-Par3
2018 Adaptive Partitioning for Iterated Sequences of Irregular OpenCL Kernels
abstract
OpenCL defines a common parallel programming language for all devices, although writing tasks adapted to the devices, managing communication and load-balancing issues are left to the programmer. We propose in this paper a static/dynamic approach for the execution of an iterated sequence of data-dependent kernels on a multi-device heterogeneous architecture. The method allows to automatically distribute irregular kernels onto multiple devices and tackles, without training, both load balancing and data transfers issues coming from hardware heterogeneity, load imbalance within the application itself and load variations between repeated executions of the sequence.
Pierre Huchant, Denis Barthou, Marie Christine Counilh
SBAC-PAD2
2017 Rewriting System for Profile-Guided Data Layout Transformations on Binaries
Christopher Haine, Olivier Aumage, Denis Barthou
Euro-Par3
2016 Automatic OpenCL Task Adaptation for Heterogeneous Architectures
Pierre Huchant, Marie Christine Counilh, Denis Barthou
Euro-Par3
2016 Specific Read-Only Data Management for Memory System Optimization
abstract
This paper proposes a new way of managing the cache by exploiting the difference of behavior in the memory system between read-only data and read-write data. A division of the existing cache-based memory hierarchy is proposed in order to create a dedicated data path for read-only data. In order to justify this approach, an analysis performed on a set of benchmarks shows that read-only data count for significant part of the working set and are less reused than read-write data. A transparent solution is proposed based on specific compilation support to separate automatically the memory accesses of read-only data at L1-level. This organization exploits the properties of the different sub-workloads in order to increase the overall data locality and data reuse. Simulated in a multicore environment, the evaluation of the new memory organization shows reduction of L1 misses up to 28.5%. Moreover, the messages issued on the interconnection network can be reduced up to 14.7% without any penalty on the performance.
Gregory Vaumourin, Alexandre Guerre, Thomas Dombek, Denis Barthou
PDP4
2015 MPI Thread-Level Checking for MPI+OpenMP Applications
Emmanuelle Saillard, Patrick Carribault, Denis Barthou
Euro-Par3
2015 Static/Dynamic validation of MPI collective communications in multi-threaded context
abstract
Scientific applications mainly rely on the MPI parallel programming model to reach high performance on supercomputers. The advent of manycore architectures (larger number of cores and lower amount of memory per core) leads to mix MPI with a thread-based model like OpenMP. But integrating two different programming models inside the same application can be tricky and generate complex bugs. Thus, the correctness of hybrid programs requires a special care regarding MPI calls location. For example, identical MPI collective operations cannot be performed by multiple non-synchronized threads. To tackle this issue, this paper proposes a static analysis and a reduced dynamic instrumentation to detect bugs related to misuse of MPI collective operations inside or outside threaded regions. This work extends PARCOACH designed for MPI-only applications and keeps the compatibility with these algorithms. We validated our method on multiple hybrid benchmarks and applications with a low overhead.
Emmanuelle Saillard, Patrick Carribault, Denis Barthou
PPoPP3
2015 Correctness Analysis of MPI-3 Non-Blocking Communications in PARCOACH
abstract
MPI-3 provide functions for non-blocking collectives. To help programmers introduce non-blocking collectives to existing MPI programs, we improve the PARCOACH tool for checking correctness of MPI call sequences. These enhancements focus on correct call sequences of all flavor of collective calls, and on the presence of completion calls for all non-blocking communications. The evaluation shows an overhead under 10% of original compilation time.
Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou
EuroMPI4
2014 SPAGHETtI: Scheduling/Placement Approach for Task-Graphs on HETerogeneous archItecture
Denis Barthou, Emmanuel Jeannot
Euro-Par1
2014 Toward OpenCL Automatic Multi-Device Support
Sylvain Henry, Alexandre Denis 0001, Denis Barthou, Marie Christine Counilh, Raymond Namyst
Euro-Par3
2013 Hydra: Automatic algorithm exploration from linear algebra equations
abstract
Hydra accepts an equation written in terms of operations on matrices and automatically produces highly efficient code to solve these equations. Processing of the equation starts by tiling the matrices. This transforms the equation into either a single new equation containing terms involving tiles or into multiple equations some of which can be solved in parallel with each other. Hydra continues transforming the equations using tiling and seeking terms that Hydra knows how to compute or equations it knows how to solve. The end result is that by transforming the equations Hydra can produce multiple solvers with different locality behavior and/or different parallel execution profiles. Next, Hydra applies empirical search over this space of possible solvers to identify the most efficient version. In this way, Hydra enables the automatic production of efficient solvers requiring very little or no coding at all and delivering performance approximating that of the highly tuned library routines such as Intel's MKL.
Alexandre Duchateau, David A. Padua, Denis Barthou
CGO3
2013 Topic 4: High-Performance Architectures and Compilers - (Introduction)
Denis Barthou, Wolfgang Karl, Ramón Doallo, Evelyn Duesterwald, Sami Yehia
Euro-Par1
2013 Dynamic Thread Pinning for Phase-Based OpenMP Programs
Abdelhafid Mazouz, Sid Ahmed Ali Touati, Denis Barthou
Euro-Par3
2013 MIL: A language to build program analysis tools through static binary instrumentation
abstract
As software complexity increases, the analysis of code behavior during its execution is becoming more important. Instrumentation techniques, through the insertion of code directly into binaries, are essential for program analyses used in debugging, runtime profiling, and performance evaluation. In the context of high-performance parallel applications, building an instrumentation framework is quite challenging. One of the difficulties is due to the necessity to capture both coarse-grain behavior, such as the execution time of different functions, as well as finer-grain actions, in order to pinpoint performance issues. In this paper, we propose a language, MIL, for the development of program analysis tools based on static binary instrumentation. The key feature of MIL is to ease the integration of static, global program analysis with instrumentation. We will show how this enables both a precise targeting of the code regions to analyze and a better understanding of the optimized program behavior.
Andres Charif Rubial, Denis Barthou, Cédric Valensi, Sameer Shende, Allen D. Malony, William Jalby
HiPC2
2013 Combining static and dynamic validation of MPI collective communications
abstract
Collective MPI communications have to be executed in the same order by all processes in their communicator and the same number of times, otherwise a deadlock occurs. As soon as the control-flow involving these collective operations becomes more complex, in particular including conditionals on process ranks, ensuring the correction of such code is error-prone. We propose in this paper a static analysis to detect when such situation occurs, combined with a code transformation that prevents from deadlocking. We show on several benchmarks the small impact on performance and the ease of integration of our techniques in the development process.
Emmanuelle Saillard, Patrick Carribault, Denis Barthou
EuroMPI3
2012 Automatic efficient data layout for multithreaded stencil codes on CPU sand GPUs
abstract
Stencil based computation on structured grids is a kernel at the heart of a large number of scientific applications. The variety of stencil kernels used in practice make this computation pattern difficult to assemble into a high performance computing library. With the multiplication of cores on a single chip, answering architectural alignment requirements became an even more important key to high performance. Along with vector accesses, data layout optimization must also consider concurrent parallel accesses. In this paper, we develop a strategy to automatically generate stencil codes for multicore vector architectures, searching for the best data layout possible to answer architectural alignment problems. We introduce a new method for aligning multidimensional data structures, called multipadding, that can be adapted to specificities of multicores and GPUs architectures. We present multiple methods with different level of complexity. We show on different stencil patterns that generated codes with multipadding display better performance than existing optimizations.
Julien Jaeger, Denis Barthou
HiPC2
2011 Introduction
Mitsuhisa Sato, Denis Barthou, Pedro C. Diniz, P. Saddayapan
Euro-Par (1)2
2010 A Multidimensional Array Slicing DSL for Stream Programming
abstract
Stream languages offer a simple multi-core programming model and achieve good performance. Yet expressing data rearrangement patterns (like a matrix block decomposition) in these languages is verbose and error prone. In this paper, we propose a high-level programming language to elegantly describe n-dimensional data reorganization patterns. We show how to compile it to stream languages.
Pablo de Oliveira Castro, Stéphane Louise, Denis Barthou
CISIS3
2010 Study of Variations of Native Program Execution Times on Multi-Core Architectures
abstract
Program performance optimisations, feedback-directed iterative compilation and auto-tuning systems all assume a fixed estimation of execution time given a fixed input data for the program. However, in practice we observe non-negligible program performance variations on hardware platforms. While these variations are insignificant for sequential applications, we show that parallel native OpenMP programs have less performance stability. This article does not try to quantify nor to qualify the factors influencing the variations of program execution times, that we let for a future work. This article demonstrates three observations: 1) The performance variations of sequential applications is insignificant. 2) OpenMP program execution times on multi-core platforms show important variations. 3) The distribution of the execution times is not a Gaussian distribution in almost all cases. We finish by a discussion explaining why considering the minimal or the mean execution time within a sample of experiments is not the best estimation of program performance.
Abdelhafid Mazouz, Sid Ahmed Ali Touati, Denis Barthou
CISIS3
2010 High Performance Architectures and Compilers
Pedro C. Diniz, Marco Danelutto, Denis Barthou, Marc Gonzales, Michael Hübner 0001
Euro-Par (1)3
2009 Computing the Transitive Closure of a Union of Affine Integer Tuple Relations
Anna Beletska, Denis Barthou, Wlodzimierz Bielecki, Albert Cohen 0001
COCOA2
2009 Compositional approach applied to loop specialization
abstract
Abstract An optimizing compiler cannot generate one best code pattern for all input data. There is no ‘one optimization fits all’ inputs. To attain high performance for a large range of inputs, it is therefore desirable to resort to some kind of specialization. Data specialization significantly improves the performance delivered by the compiler‐generated codes. Specialization is, however, limited by code expansion and introduces a time overhead for the selection of the appropriate version. We propose a new method to specialize the code at the assembly level for loop structures. Our specialization scheme focuses on different ranges of loop trip count and combines all these versions into a code that switches smoothly from one to the other while the iteration count increases. Hence, the resulting code achieves the same level of performance than each version on its specific iteration interval. We illustrate the benefit of our method on the SPEC benchmarks with detailed experimental results. Copyright © 2008 John Wiley & Sons, Ltd.
Lamia Djoudi, Jean-Thomas Acquaviva, Denis Barthou
Concurr. Comput. Pract. Exp.3
2009 Improving performance of optimized kernels through fast instantiations of templates
abstract
Abstract To fully exploit the instruction‐level parallelism offered by modern processors, compilers need the necessary information available during the execution of the program. This advocates for iterative or dynamic compilation. Unfortunately, dynamic compilation is suitable only for applications where the cost of compilation may be amortized by multiple invocations of the same code. Similarly, the cost of iterative compilation makes it impractical to be widely used for performance improvement. In this article, we suggest a novel approach for improving the performance of mathematical kernels through fast instantiations of templates. Optimized templates are generated at static compile time with a limited number of compilations. The initial instantiations of these templates are performed at static compile time, and the runtime instantiations are performed with a very small overhead through specialized data, requiring no computations at runtime. It represents an effective solution in terms of reduced overhead incurring at static compile time and dynamic compile time. The experiments have been performed on an Itanium‐II architecture using highly optimized kernels ofATLASandFFTWwithiccandgcccompilers. Copyright © 2008 John Wiley & Sons, Ltd.
Minhaj Ahmad Khan, Henri-Pierre Charles, Denis Barthou
Concurr. Comput. Pract. Exp.3
2007 Hybrid Specialization: A Trade-off Between Static and Dynamic Specialization
Minhaj Ahmad Khan, Henri-Pierre Charles, Denis Barthou
PACT3
2007 Loop Optimization using Hierarchical Compilation and Kernel Decomposition
abstract
The increasing complexity of hardware features for recent processors makes high performance code generation very challenging. In particular, several optimization targets have to be pursued simultaneously (minimizing L1/L2/L3/TLB misses and maximizing instruction level parallelism). Very often, these optimization goals impose different and contradictory constraints on the transformations to be applied. We propose a new hierarchical compilation approach for the generation of high performance code relying on the use of state-of-the-art compilers. This approach is not application-dependent and do not require any assembly hand-coding. It relies on the decomposition of the original loop nest into simpler kernels, typically 1D to 2D loops, much simpler to optimize. We successfully applied this approach to optimize dense matrix muliply primitives (not only for the square case but to the more general rectangular cases) and convolution. The performance of the optimized codes on Itanium 2 and Pentium 4 architectures outperforms ATLAS and in most cases, matches hand-tuned vendor libraries (e.g. MKL)
Denis Barthou, Sébastien Donadio, Patrick Carribault, Alexandre Duchateau, William Jalby
CGO1
2007 Compositional Approach Applied to Loop Specialization
Lamia Djoudi, Jean-Thomas Acquaviva, Denis Barthou
Euro-Par3
2005 Deciding Where to Call Performance Libraries
Christophe Alias, Denis Barthou
Euro-Par2
2005 On Domain-Specific Languages Reengineering
Christophe Alias, Denis Barthou
GPCE2
2002 On the Equivalence of Two Systems of Affine Recurrence Equations (Research Note)
Denis Barthou, Paul Feautrier, Xavier Redon
Euro-Par1
1998 Maximal Static Expansion
abstract
Memory expansions are classical means to extract parallelism from imperative programs. However, for dynamic control programs with general memory accesses, such transformations either fail or require some run-time mechanism to restore the data flow. This paper presents an expansion framework for any type of data structure in any imperative program, without the need for dynamic data flow restoration. The key idea is to group together the write operations that participate in the flow of the same datum. We show that such an expansion boils down to mapping each group to a single memory cell. We give a practical algorithm for code transformation. This algorithm, however, is valid for (possibly non-affine) loops over arrays only.
Denis Barthou, Albert Cohen 0001, Jean-Francois Collard
POPL1
1997 Automatic data mapping of signal processing applications
abstract
This paper presents a technique to map automatically a complete digital signal processing (DSP) application onto a parallel machine with distributed memory. Unlike other applications where coarse or medium grain scheduling techniques can be used, DSP applications integrate several thousand of tasks and hence necessitate fine grain considerations. Moreover finding an effective mapping imperatively require to take into account both architectural resources constraints and real time constraints. The main contribution of this paper is to show how it is possible to handle and to solve data partitioning, and fine-grain scheduling under the above operational constraints using concurrent constraints logic programming languages (CCLP). Our concurrent resolution technique undertaking linear and nonlinear constraints takes advantage of the special features of signal processing applications and provides a solution equivalent to a manual solution for the representative panoramic analysis (PA) application.
Corinne Ancourt, Denis Barthou, Christophe Guettier, François Irigoin, Bertrand Jeannet, Jean Jourdan, Juliette Mattioli
ASAP2
1997 Fuzzy Array Dataflow Analysis
Denis Barthou, Jean-Francois Collard, Paul Feautrier
J. Parallel Distributed Comput.1
1995 Fuzzy Array Dataflow Analysis
abstract
Exact array dataflow analysis can be achieved in the general case if the only control structures are do-loops and structural ifs, and if loop counter bounds and array subscripts are affine expressions of englobing loop counters and possibly some integer constants. In this paper, we begin the study of dataflow analysis of dynamic control programs, where arbitrary ifs and whiles are allowed. In the general case, this dataflow analysis can only be fuzzy.
Jean-Francois Collard, Denis Barthou, Paul Feautrier
PPoPP2