Julien Bigot

dblp:05/2495 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-0015-4304ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 An Optimization Technique for Hiding Communication Costs in 3D Parallel Training of Deep Learning
abstract
In large-scale model training, distributing the workload is essential to address memory usage and computation time constraints. 3D parallelism, which combines data parallelism, tensor parallelism, and pipeline parallelism, has emerged as a state-of-the-art method for distributed training. However, exploring optimal parallel strategies and modifying code places a significant burden on users. While some machine learning compilers support automatic parallelization, they are generally limited to SPMD parallelism and do not support the automation of 3D parallelism, including MPMD parallelism for pipeline parallelism. This study identifies key issues in implementing 3D parallelism with the XLA machine learning compiler, particularly communication delays caused by the limitations of sharding notation in XLA's intermediate representation. To address this, we propose a commshift optimization, which relocates specific communication instructions to different computation phases, reducing the impact of communication waiting times. This optimization reduces training time, achieving up to a 27% throughput improvement in GPT-J training.
Ryubu Hosoki, Kento Sato, Toshio Endo, Julien Bigot, Edouard Audit
CCGrid4
2024 Extreme-scale workflows: A perspective from the JLESC international community
Orcun Yildiz, Amal Gueroudji, Julien Bigot, Bruno Raffin, Rosa M. Badia, Tom Peterka
Future Gener. Comput. Syst.3
2021 DEISA: Dask-Enabled In Situ Analytics
abstract
A widening performance gap is separating CPU performance and IO bandwidth on large scale systems. In some fields such as weather forecast and nuclear fusion, numerical models generate such amounts of data that classical post hoc processing is not feasible anymore due to the limits in both storage capacity and IO performance. In situ approaches are attractive to bypass disk accesses in these cases and fully leverage the HPC platform. They are however often complex to set up and can require to re-develop parallel versions of the analysis from scratch. In this paper we propose a hybrid model that is well suited for in situ workflows that combine regular simulations and irregular analytics. Our model couples the bulk synchronous parallel paradigm for simulation with a distributed task-based one for analysis. This reduces complexity and leverages the best of each of these two powerful paradigms. We validate the model with a prototype, called DEISA, that supports coupling MPI parallel codes with analyses written using Dask. This implementation requires minimal modifications of both the simulation and analysis codes compared to their post hoc counterpart. It give access to an already existing rich ecosystem to be used in situ such as the parallel versions of Numpy, Pandas and scikit-learn. Experiments in configurations up to 1024 cores show that DEISA can improve the simulation wallclock time (excluding analysis) by a factor up to 3 and the total experiment (including analysis) hour. core cost by a factor of up to 5 compared to parallel post hoc with plain Dask while requiring the modification of only two lines of python code, three of YAML, and none at all in a C simulation code already instrumented with PDI Data Interface.
Amal Gueroudji, Julien Bigot, Bruno Raffin
HiPC2
2020 Overlapping communications in gyrokinetic codes on accelerator-based platforms
abstract
Summary Communication and computation overlapping techniques have been introduced in the five‐dimensional gyrokinetic codes GYSELA and GKV. In order to anticipate some of the exa‐scale requirements, these codes were ported to the modern accelerators, Xeon Phi KNL and Tesla P 100 GPU. On accelerators, a serial version of GYSELA on KNL and GKV on GPU are respectively 1.3× and 7.4× faster than those on a single Skylake processor (a single socket). For the scalability, we have measured GYSELA performance on Xeon Phi KNL from 16 to 512 KNLs (1024 to 32k cores) and GKV performance on Tesla P 100 GPU from 32 to 256 GPUs. In their parallel versions, transpose communication in semi‐Lagrangian solver in GYSELA or Convolution kernel in GKV turned out to be a main bottleneck. This indicates that in the exa‐scale, the network constraints would be critical. In order to mitigate the communication costs, the pipeline and task‐based overlapping techniques have been implemented in these codes. The GYSELA 2D advection solver has achieved a 33% to 92% speed up, and the GKV 2D convolution kernel has achieved a factor of 2 speed up with pipelining. The task‐based approach gives 11% to 82% performance gain in the derivative computation of the electrostatic potential in GYSELA. We have shown that the pipeline‐based approach is applicable with the presence of symmetry, while the task‐based approach can be applicable to more general situations.
Yuuichi Asahi, Guillaume Latu, Julien Bigot, Shinya Maeyama, Virginie Grandgirard, Yasuhiro Idomura
Concurr. Comput. Pract. Exp.3
2020 InKS: a programming model to decouple algorithm from optimization in HPC codes
Ksander Ejjaaouani, Olivier Aumage, Julien Bigot, Michel Mehrenberger, Hitoshi Murai, Masahiro Nakao, Mitsuhisa Sato
J. Supercomput.3
2019 Fine-Grained MPI+OpenMP Plasma Simulations: Communication Overlap with Dependent Tasks
Jérôme Richard, Guillaume Latu, Julien Bigot
Euro-Par3
2018 Scaling and Optimizing the Gysela Code on a Cluster of Many-Core Processors
abstract
The current generation of the Xeon Phi Knights Landing (KNL) processor provides a highly multi-threaded environment on which regular programming models such as MPIjopenMP can be used. Many factors impact the performance achieved by applications on these devices: one of the key points is the efficient exploitation of SIMD vector units, and one another is the memory access pattern. Works have been conducted to adapt a plasma turbulence application, namely Gysela, for this architecture. A set of different techniques have been used: standard vectorization techniques, auto-tuning of one computation kernel, switching to high-order scheme. As a result, KNL execution times have been reduced by up to a factor 3. This effort has also permitted to gain a speedup of 2x on Broadwell architecture and 3x on Skylake. Nice scalability curves up to a few thousands cores have been obtained on a strong scaling experiment. Incremental work meant a large payoff without resorting to using low-level intrinsics.
Guillaume Latu, Yuuichi Asahi, Julien Bigot, Tamas B. Fehér, Virginie Grandgirard
SBAC-PAD3
2017 Combining Both a Component Model and a Task-based Model for HPC Applications: a Feasibility Study on Gysela
abstract
This paper studies the feasibility of efficiently combining both a software component model and a task-based model. Task based models are known to enable efficient executions on recent HPC computing nodes while component models ease the separation of concerns of application and thus improve their modularity and adaptability. This paper describes a prototype version of the COMET programming model combining concepts of task-based and component models, and a preliminary version of the COMET runtime built on top of StarPU and L2C. Evaluations of the approach have been conducted on a real-world use-case analysis of a subpart of the production application GYSELA. Results show that the approach is feasible and that it enables easy composition of independent software codes without introducing overheads. Performance results are equivalent to those obtained with a plain OpenMP based implementation.
Olivier Aumage, Julien Bigot, Hélène Coullon, Christian Pérez, Jérôme Richard
CCGrid2
2009 Increasing Reuse in Component Models through Genericity
Julien Bigot, Christian Pérez
ICSR1