Eric Petit 0002

dblp:85/197-2 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-5047-1407ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 since 2021Theory of computation · 8 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Harnessing MPI mutations for AI error detection
abstract
MPI errors are challenging to identify despite the significant number of expert verification tools. Dynamic tools (i.e., requiring profiling) are computationally expensive and accurate in error detection, whereas static analysis (i.e., operating at source code or compilation) is computationally cheap but less accurate. Interestingly, the recent success of AI and LLMs offers an alternative to increase static analysis accuracy while preserving its low overhead. Yet current methods remain difficult to benchmark, too general, and poorly adapted to the specific challenges of high-performance computing.
Asia Auville, Tim Jammer, Eric Petit 0002, Pablo de Oliveira Castro, Emmanuelle Saillard, Mihail Popov
ICS3
2025 Productively Generating a High-Performance Linear Algebra Library on FPGAs
abstract
Linear algebra computations can be greatly accelerated using spatial accelerators on FPGAs. As a standard building block of linear algebra applications, BLAS covers a wide range of compute patterns that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. However, existing implementations of BLAS routines on FPGAs are stuck in the dilemma of productivity and performance. They either require extensive human effort or fail to leverage the properties of routines for acceleration. We introduce Lasa, a framework composed of a programming model and a compiler, designed to address the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. The programming model realizes systolic arrays using uniform recurrence equations and space-time transforms. Streaming tensors, an intuitive dataflow-style abstraction, is proposed to uniformly describe the movement, storage, and transpose of input and output data across the spatial components. According to streaming tensors, a customized memory hierarchy is automatically built on an FPGA by our compiler. The compiler further specializes the architecture with transparent optimizations on FPGAs. Using this framework, we develop a complete BLAS library, demonstrating performance in parity with expert-written HLS code for BLAS level 3 routines, 76%–94% machine peak for level 1 and 2 routines, and 1.6X–13X speedup by leveraging the matrix properties such as symmetry, triangularity, and bandness.
Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001
ACM Trans. Reconfigurable Technol. Syst.8
2024 Deconstructing HPL-MxP Benchmark: A Numerical Perspective
Greg Henry, Eric Petit 0002, Alexander Lyashevsky, Peter Caday
Euro-Par (1)2
2023 Lasa: Abstraction and Specialization for Productive and Performant Linear Algebra on FPGAs
abstract
Linear algebra can often be significantly expedited by spatial accelerators on FPGAs. As a broadly-adopted linear algebra library, BLAS requires extensive optimizations for routines that vary vastly in data reuse, bottleneck resources, matrix storage layouts, and data types. Existing solutions are stuck in the dilemma of productivity and performance. We introduce Lasa, a framework composed of a programming model and a compiler, that addresses the dilemma by abstracting (for productivity) and specializing (for performance) the architecture of a spatial accelerator. Lasa abstracts a compute and its I/O as two dataflow graphs. A compiler maps the graphs onto systolic arrays and a customized memory heirarchy. The compiler further specializes the architecture transparently. In this framework, we develop 14 key BLAS routines, and demonstrate performance in parity with expert-written HLS code for BLAS level 3 routines, >=80% machine peak performance for level 2 and 1 routines, and 1.6X-7X speed up by taking advantage of matrix properties of symmetry, triangularity and bandness.
Xiaochen Hao, Mingzhe Zhang 0002, Ce Sun 0001, Zhuofu Tao, Hongbo Rong, Yu Zhang 0086, Lei He 0001, Eric Petit 0002, Yun Liang 0001
FCCM8
2022 The Positive Effects of Stochastic Rounding in Numerical Algorithms
abstract
Recently, stochastic rounding (SR) has been implemented in specialized hardware but most current computing nodes do not yet support this rounding mode. Several works empirically illustrate the benefit of stochastic rounding in various fields such as neural networks and ordinary differential equations. For some algorithms, such as summation, inner product or matrix-vector multiplication, it has been proved that SR provides probabilistic error bounds better than the traditional deterministic bounds. In this paper, we extend this theoretical ground for a wider adoption of SR in computer architecture. First, we analyze the biases of the two SR modes: SR-nearness and SR-up-or-down. We demonstrate on a case-study of Euler's forward method that IEEE-754 default rounding modes and SR-up-or-down accumulate rounding errors across iterations and that SR-nearness, being unbiased, does not. Second, we prove a $O(\sqrt{n})$ probabilistic bound on the forward error of Horner's polynomial evaluation method with SR, improving on the known deterministic O(n) bound.
El-Mehdi El Arar, Devan Sohier, Pablo de Oliveira Castro, Eric Petit 0002
ARITH4
2022 A BF16 FMA is All You Need for DNN Training
abstract
Multiply-Add (FMA) functional units constitute a fundamental hardware component to train Deep Neural Networks (DNNs).Its silicon area grows quadratically with the mantissa bit count of the computer number format, which has motivated the adoption of the BrainFloat16 format (BF16).BF16 features 1 sign, 8 exponent and 7 explicit mantissa bits.Some approaches to train DNNs achieve significant performance benefits by using the BF16 format.However, these approaches must combine BF16 with the standard IEEE 754 Floating-Point 32-bit (FP32) format to achieve state-of-the-art training accuracy, which limits the impact of adopting BF16.This paper proposes the first approach able to train complex DNNs entirely using the BF16 format.We propose a new class of FMA operators, FMA bf16 n m , that entirely rely on BF16 FMA hardware instructions and deliver the same accuracy as FP32.FMA bf16 n m operators achieve performance improvements within the 1.28-1.35×range on ResNet101 with respect to FP32.FMA bf16 n m enables training complex DNNs on simple low-end hardware devices without requiring expensive FP32 FMA functional units.
John Osorio Ríos, Adrià Armejach, Eric Petit 0002, Greg Henry, Marc Casas
ARITH3
2022 FASE: A Fast, Accurate and Seamless Emulator for Custom Numerical Formats
abstract
Deep Neural Networks (DNNs) have become ubiquitous in a wide range of application domains. Despite their success, training DNNs is an expensive task that has motivated the use of reduced numerical precision formats to improve performance and reduce power consumption. Emulation techniques are a good fit to understand the properties of new numerical formats on a particular workload. However, current SoA techniques are not able to perform these tasks quickly and accurately on a wide variety of workloads.We propose FASE, a Fast, Accurate, and Seamless Emulator that leverages dynamic binary translation to enable emulation of custom numerical formats. FASE is fast: allowing emulation of large unmodified workloads; accurate: emulating at the instruction operand level; and seamless: as it does not require any code modifications and works on any application or DNN framework without any language, compiler, or source code access restrictions.
John Osorio Ríos, Adrià Armejach, Eric Petit 0002, Greg Henry, Marc Casas
ISPASS3
2022 FASE: A Fast, Accurate and Seamless Emulator for Custom Numerical Formats
John Osorio Ríos, Adrià Armejach, Eric Petit 0002, Greg Henry, Marc Casas
ECML/PKDD (5)3
2021 A Study of the Effects and Benefits of Custom-Precision Mathematical Libraries for HPC Codes
abstract
Published in "IEEE Transactions on Emerging Topics in Computing, Volume: 9, Issue: 3, JulySeptember 2021" and orally presented at ARITH 2021.
Emeric Brun, David Defour, Pablo de Oliveira Castro, Matei Istoan, Davide Mancusi, Eric Petit 0002, Alan Vaquet
ARITH6
2021 Shadow computation with BFloat16 to estimate the numerical accuracy of summations
abstract
In this article, we propose to exploit the new computational capability offered by the Bfloat16 representation format to perform shadow computations and compute estimations of the relative error. We demonstrate and evaluate the assumptions under which shadow computation is valid for the summation problem.
David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002
ARITH4
2021 Dynamically Adapting Floating-Point Precision to Accelerate Deep Neural Network Training
abstract
Mixed-precision (MP) arithmetic combining both single- and half-precision operands has been successfully applied to train deep neural networks. Despite its advantages in terms of reducing the need for key resources like memory bandwidth or register file size, it has a limited capacity for diminishing further computing costs, as it requires 32-bits to represent its output. On the other hand, full half-precision arithmetic fails to deliver state-of-the-art training accuracy. We design a binary tool SERP based on Intel Pin which allows us to characterize and analyze computer arithmetic usage in machine learning frameworks (Pytorch, Caffe, Tensorflow) and to emulate different floating point formats. Based on empirical observations about precision needs on representative deep neural networks, this paper proposes a seamless approach to dynamically adapt floating point arithmetic. Our dynamically adaptive methodology enables the use of full half-precision arithmetic for up to 96.4% of the computations when training state-of-the-art neural networks; while delivering comparable accuracy to 32-bit floating point arithmetic. Microarchitectural simulations indicate that our Dynamic approach accelerates training deep convolutional and recurrent networks with respect to FP32 by 1.39 × and 1.26 ×, respectively.
John Osorio Ríos, Adrià Armejach, Eric Petit 0002, Greg Henry, Marc Casas
ICMLA3
2021 Confidence Intervals for Stochastic Arithmetic
abstract
Quantifying errors and losses due to the use of Floating-point (FP) calculations in industrial scientific computing codes is an important part of the Verification, Validation, and Uncertainty Quantification process. Stochastic Arithmetic is one way to model and estimate FP losses of accuracy, which scales well to large, industrial codes. It exists in different flavors, such as CESTAC or MCA, implemented in various tools such as CADNA, Verificarlo, or Verrou. These methodologies and tools are based on the idea that FP losses of accuracy can be modeled via randomness. Therefore, they share the same need to perform a statistical analysis of programs results to estimate the significance of the results. In this article, we propose a framework to perform a solid statistical analysis of Stochastic Arithmetic. This framework unifies all existing definitions of the number of significant digits (CESTAC and MCA), and also proposes a new quantity of interest: the number of digits contributing to the accuracy of the results. Sound confidence intervals are provided for all estimators, both in the case of normally distributed results, and in the general case. The use of this framework is demonstrated by two case studies of industrial codes: Europlexus and code_aster.
Devan Sohier, Pablo de Oliveira Castro, François Févotte, Bruno Lathuilière, Eric Petit 0002, Olivier Jamond
ACM Trans. Math. Softw.5
2020 Custom-Precision Mathematical Library Explorations for Code Profiling and Optimization
abstract
The typical processors used for scientific computing have fixed-width data-paths. This implies that mathematical libraries were specifically developed to target each of these fixed precisions (binary16, binary32, binary64). However, to address the increasing energy consumption and throughput requirements of scientific applications, library and hardware designers are moving beyond this one-size-fits-all approach. In this article we propose to study the effects and benefits of using user-defined floating-point formats and target accuracies in calculations involving mathematical functions. Our tool collects input-data profiles and iteratively explores lower precisions for each call-site of a mathematical function in user applications. This profiling data will be a valuable asset for specializing and fine-tuning mathematical function implementations for a given application. We demonstrate the tool's capabilities on SGP4, a satellite tracking application. The profile data shows the potential for specialization and provides insight into answering where it is useful to provide variable-precision designs for elementary function evaluation.
David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002
ARITH4
2020 Evaluating Mixed-Precision Arithmetic for 3D Generative Adversarial Networks to Simulate High Energy Physics Detectors
abstract
Several hardware companies are proposing native Brain Float 16-bit (BF16) support for neural network training. The usage of Mixed Precision (MP) arithmetic with floating-point 32-bit (FP32) and 16-bit half-precision aims at improving memory and floating-point operations throughput, allowing faster training of bigger models. This paper proposes a binary analysis tool enabling the emulation of lower precision numerical formats in Neural Network implementation without the need for hardware support. This tool is used to analyze BF16 usage in the training phase of a 3D Generative Adversarial Network (3DGAN) simulating High Energy Physics detectors. The binary tool allows us to confirm that BF16 can provide results with similar accuracy as the full-precision 3DGAN version and the costly reference numerical simulation using double precision arithmetic.
John Osorio Ríos, Adrià Armejach, Gul Rukhkhattak, Eric Petit 0002, Sofia Vallecorsa, Marc Casas
ICMLA4
2019 Automatic Exploration of Reduced Floating-Point Representations in Iterative Methods
Yohan Chatelain, Eric Petit 0002, Pablo de Oliveira Castro, Ghislain Lartigue, David Defour
Euro-Par2
2018 VeriTracer: Context-enriched tracer for floating-point arithmetic analysis
abstract
VeriTracer automatically instruments a code and traces the accuracy of floating-point variables over time. VeriTracer enriches the visual traces with contextual information such as the call site path in which a value was modified. Contextual information is important to understand how the floating-point errors propagate in complex codes. VeriTracer is implemented as an LLVM compiler tool on top of Verificarlo. We demonstrate how VeriTracer can detect accuracy loss and quantify the impact of using a compensated algorithm on ABINIT, an industrial HPC application for Ab Initio quantum computation.
Yohan Chatelain, Pablo de Oliveira Castro, Eric Petit 0002, David Defour, Jordan Bieder, Marc Torrent 0001
ARITH3
2018 Asynchronous and multithreaded communications on irregular applications using vectorized divide and conquer approach
Loïc Thébault, Eric Petit 0002
J. Parallel Distributed Comput.2
2016 Verificarlo: Checking Floating Point Accuracy through Monte Carlo Arithmetic
abstract
Numerical accuracy of floating point computation is a well studied topic which has not made its way to the end-user in scientific computing. Yet, it has become a critical issue with the recent requirements for code modernization to harness new highly parallel hardware and perform higher resolution computation. To democratize numerical accuracy analysis, it is important to propose tools and methodologies to study large use cases in a reliable and automatic way. In this paper, we propose verificarlo, an extension to the LLVM compiler to automatically use Monte Carlo Arithmetic in a transparent way for the end-user. It supports all the major languages including C, C++, and Fortran. Unlike source-to-source approaches, our implementation captures the influence of compiler optimizations on the numerical accuracy. We illustrate how Monte Carlo Arithmetic using the verificarlo tool outperforms the existing approaches on various use cases and is a step toward automatic numerical analysis.
Christophe Denis, Pablo de Oliveira Castro, Eric Petit 0002
ARITH3
2016 A software scheduling solution to avoid corrupted units on GPUs
David Defour, Eric Petit 0002
J. Parallel Distributed Comput.2
2015 CERE: LLVM-Based Codelet Extractor and REplayer for Piecewise Benchmarking and Optimization
abstract
This article presents Codelet Extractor and REplayer (CERE), an open-source framework for code isolation. CERE finds and extracts the hotspots of an application as isolated fragments of code, called codelets . Codelets can be modified, compiled, run, and measured independently from the original application. Code isolation reduces benchmarking cost and allows piecewise optimization of an application. Unlike previous approaches, CERE isolates codes at the compiler Intermediate Representation (IR) level. Therefore CERE is language agnostic and supports many input languages such as C, C++, Fortran, and D. CERE automatically detects codelets invocations that have the same performance behavior. Then, it selects a reduced set of representative codelets and invocations, much faster to replay, which still captures accurately the original application. In addition, CERE supports recompiling and retargeting the extracted codelets. Therefore, CERE can be used for cross-architecture performance prediction or piecewise code optimization. On the SPEC 2006 FP benchmarks, CERE codelets cover 90.9% and accurately replay 66.3% of the execution time. We use CERE codelets in a realistic study to evaluate three different architectures on the NAS benchmarks. CERE accurately estimates each architecture performance and is 7.3 × to 46.6 × cheaper than running the full benchmark.
Pablo de Oliveira Castro, Chadi Akel, Eric Petit 0002, Mihail Popov, William Jalby
ACM Trans. Archit. Code Optim.3
2013 Adaptive sampling for performance characterization of application kernels
abstract
SUMMARY Characterizing performance is essential to optimize programs and architectures. The open source Adaptive Sampling Kit (ASK) measures the performance trade‐off in large design spaces. Exhaustively sampling all sets of parameters is computationally intractable. Therefore, ASK concentrates exploration in the most irregular regions of the design space through multiple adaptive sampling strategies. The paper presents the ASK architecture and a set of adaptive sampling strategies, including a new approach called Hierarchical Variance Sampling. ASK's usage is demonstrated on three performance characterization problems: memory stride accesses, Jacobian stencil code, and an industrial seismic application using 3D stencils. ASK builds accurate models of performance with a small number of measures. It considerably reduces the cost of performance exploration. For instance, the Jacobian stencil code design space, which has more than 31 × 108 combinations of parameters, is accurately predicted using only 1500 combinations. Copyright © 2013 John Wiley & Sons, Ltd.
Pablo de Oliveira Castro, Eric Petit 0002, Asma Farjallah, William Jalby
Concurr. Comput. Pract. Exp.2
2012 ASK: Adaptive Sampling Kit for Performance Characterization
Pablo de Oliveira Castro, Eric Petit 0002, Jean Christophe Beyler, William Jalby
Euro-Par2