Thibaut Tachon

dblp:196/8012 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-3264-5535ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2025 BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN Training
abstract
Pipeline parallelism and activation recomputation are widely adopted optimization techniques, among others, to scale DNN training on large accelerator clusters. However, as DNNs grow in complexity and heterogeneity, it becomes increasingly difficult to determine the optimal combination of pipeline partitioning and recomputation strategies. Existing solutions either propose manual optimization approaches that do not scale or automated approaches that explore only a subset of optimization possibilities due to an explosion of search space. In this paper, we present BMPipe, a bubble-memory co-optimization planner that holistically optimizes computation imbalance, memory under utilization, redundant computation, and schedulinginduced preparation time. At its core, BMPipe uses symbolic representations that unify computation, memory, and bubbles into a single model that is solved by using an ILP-based planner. Using BMPipe, we perform a thorough experimental evaluation where we train several large, state-of-the-art DNN models on a 16K-NPU cluster. We show that BMPipe achieves up to$1.36 \times$speedup compared to the state-of-the-art solution Megatron. Against automatic planners PipeDream, Merak and AdaPipe,, it yields as$1.27 \times$speed-up. In addition, BMPipe boosts peak device-memory utilization by$\mathbf{1. 4 2} \times$compared with Megatron.
Ruiwen Wang, Chong Li 0003, Thibaut Tachon, Raja Appuswamy, Teng Su
CLUSTER3
2022 Parallelizing Neural Network Models Effectively on GPU by Implementing Reductions Atomically
abstract
Due to the missing of a good orchestration of loop transformations, existing optimizing compilers for deploying neural networks on GPU either parallelize reductions ineffectively or miss the fusion opportunities with other operators. Neural network models thus exhibit sub-optimal performance on GPU. We present a practical approach called Panamera for the effective parallelization of reductions in neural networks on GPU. Panamera first leverages loop coalescing to flatten the loop dimensions of reductions, converting all reduction operators into canonical forms eligible for the polyhedral model. Next, Panamera uses polyhedral transformations to reduce the data movements caused by unfused reductions and perform multi-block hardware binding not considered by many compilers. Finally, Panamera embeds a highly optimized routine implemented using GPU atomic instructions, further improving the performance of neural network models while guaranteeing the correctness of parallel reductions. The experimental results demonstrate the effectiveness of our approach: for single operators our code obtains a mean speedup of 33.7×, 3.5×, 5.4× and 9.6× over cuDNN, CUB, TVM and Ansor, for sub-graphs our approach outperforms cuDNN, TVM and Ansor by 9.5×, 2.6× and 2.7×, and for end-to-end workloads, a tensor compiler integrated with our approach outperforms them by 122.5%, 19.3% and 15.2%.
Jie Zhao 0002, Cédric Bastoul, Yanzhi Yi, Wang Nie, Renwei Zhang, Zhen Geng, Chong Li 0003, Thibaut Tachon, Zhiliang Gan
PACT9
2021 Efficient and Systematic Partitioning of Large and Deep Neural Networks for Parallelization
Chong Li 0003, Thibaut Tachon, Sébastien Limet, Sophie Robert 0001
Euro-Par3