David Defour

dblp:08/1837 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
3since 2021 · last 2023
0000-0001-9923-2394ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-authorTheory of computation · 8 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Chromatic Analysis of Numerical Programs
abstract
This paper introduces the concept of chromatic numbers, which allows to tint a scalar or a set of scalars to estimate the relations between input and output variables under additive property. This consists in proposing a decomposition of a resulting value as a sum of tinted values. We illustrate how this concept can be used on a deep neural networks example by tinting group of value at once allowing to track a large number of values together.
David Defour, Franck Védrine
ARITH1
2021 A Study of the Effects and Benefits of Custom-Precision Mathematical Libraries for HPC Codes
abstract
Published in "IEEE Transactions on Emerging Topics in Computing, Volume: 9, Issue: 3, JulySeptember 2021" and orally presented at ARITH 2021.
Emeric Brun, David Defour, Pablo de Oliveira Castro, Matei Istoan, Davide Mancusi, Eric Petit 0002, Alan Vaquet
ARITH2
2021 Shadow computation with BFloat16 to estimate the numerical accuracy of summations
abstract
In this article, we propose to exploit the new computational capability offered by the Bfloat16 representation format to perform shadow computations and compute estimations of the relative error. We demonstrate and evaluate the assumptions under which shadow computation is valid for the summation problem.
David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002
ARITH1
2020 Custom-Precision Mathematical Library Explorations for Code Profiling and Optimization
abstract
The typical processors used for scientific computing have fixed-width data-paths. This implies that mathematical libraries were specifically developed to target each of these fixed precisions (binary16, binary32, binary64). However, to address the increasing energy consumption and throughput requirements of scientific applications, library and hardware designers are moving beyond this one-size-fits-all approach. In this article we propose to study the effects and benefits of using user-defined floating-point formats and target accuracies in calculations involving mathematical functions. Our tool collects input-data profiles and iteratively explores lower precisions for each call-site of a mathematical function in user applications. This profiling data will be a valuable asset for specializing and fine-tuning mathematical function implementations for a given application. We demonstrate the tool's capabilities on SGP4, a satellite tracking application. The profile data shows the potential for specialization and provides insight into answering where it is useful to provide variable-precision designs for elementary function evaluation.
David Defour, Pablo de Oliveira Castro, Matei Istoan, Eric Petit 0002
ARITH1
2019 Automatic Exploration of Reduced Floating-Point Representations in Iterative Methods
Yohan Chatelain, Eric Petit 0002, Pablo de Oliveira Castro, Ghislain Lartigue, David Defour
Euro-Par5
2018 VeriTracer: Context-enriched tracer for floating-point arithmetic analysis
abstract
VeriTracer automatically instruments a code and traces the accuracy of floating-point variables over time. VeriTracer enriches the visual traces with contextual information such as the call site path in which a value was modified. Contextual information is important to understand how the floating-point errors propagate in complex codes. VeriTracer is implemented as an LLVM compiler tool on top of Verificarlo. We demonstrate how VeriTracer can detect accuracy loss and quantify the impact of using a compensated algorithm on ABINIT, an industrial HPC application for Ab Initio quantum computation.
Yohan Chatelain, Pablo de Oliveira Castro, Eric Petit 0002, David Defour, Jordan Bieder, Marc Torrent 0001
ARITH4
2018 FP-ANR: A representation format to handle floating-point cancellation at run-time
abstract
When dealing with floating-point numbers, there are several sources of error which can drastically reduce the numerical quality of computed results. One of those error sources is the loss of significance or cancellation, which occurs during for example, the subtraction of two nearly equal numbers. In this article, we propose a representation format named Floating-Point Adaptive Noise Reduction (FP-ANR). This format embeds cancellation information directly into the floating-point representation format thanks to a dedicated pattern. With this format, insignificant trailing bits lost during cancellation are removed from every manipulated floating-point number. The immediate consequence is that it increases the numerical confidence of computed values. The proposed representation format corresponds to a simple and efficient implementation of significance arithmetic based and compatible with the IEEE Standard 754 standard.
David Defour
ARITH1
2017 Asynchronous Power Flow on Graphic Processing Units
abstract
Asynchronous iterations can be used to implement fixed-point methods such as Jacobi and Gauss-Seidel on parallel computers with high synchronization costs. However, they are rarely considered in practice due to the slow convergence rate. This paper describes an implementation on GPUs of a novel Power Flow analysis model using asynchronous iterations. We present our model for the solution of the Power Flow analysis problem, prove its convergence and evaluate its performance for a GPU execution.
Manuel Marin, David Defour, Federico Milano 0001
PDP2
2017 Exact Lookup Tables for the Evaluation of Trigonometric and Hyperbolic Functions
abstract
Elementary mathematical functions are pervasively used in many applications such as electronic calculators, computer simulations, or critical embedded systems. Their evaluation is always an approximation, which usually makes use of mathematical properties, precomputed tabulated values, and polynomial approximations. Each step generally combines error of approximation and error of evaluation on finite-precision arithmetic. When they are used, tabulated values generally embed rounding error inherent to the transcendence of elementary functions. In this article, we propose a general method to use error-free values that is worthy when two or more terms have to be tabulated in each table row. For the trigonometric and hyperbolic functions, we show that Pythagorean triples can lead to such tables in little time and memory usage. When targeting correct rounding in double precision for the same functions, we also show that this method saves memory and floating-point operations by up to 29 and 42 percent, respectively.
Hugues de Lassus Saint-Genies, David Defour, Guillaume Revy
IEEE Trans. Computers2
2017 An Efficient Representation Format for Fuzzy Intervals Based on Symmetric Membership Functions
abstract
This article addresses the execution cost of arithmetic operations with a focus on fuzzy arithmetic. Thanks to an appropriate representation format for fuzzy intervals, we show that it is possible to halve the number of operations and divide by 2 to 8 the memory requirements compared to conventional solutions. In addition, we demonstrate the benefit of some hardware features encountered in today’s accelerators (GPU) such as static rounding, memory usage, instruction-level parallelism (ILP), and thread-level parallelism (TLP). We then describe a library of fuzzy arithmetic operations written in CUDA and C++. The library is evaluated against traditional approaches using compute-bound and memory-bound benchmarks on Nvidia GPUs, with an observed performance gain of 2 to 20.
Manuel Marin, David Defour, Federico Milano 0001
ACM Trans. Math. Softw.2
2016 A software scheduling solution to avoid corrupted units on GPUs
David Defour, Eric Petit 0002
J. Parallel Distributed Comput.1
2015 Range reduction based on Pythagorean triples for trigonometric function evaluation
abstract
Software evaluation of elementary functions usually requires three steps: a range reduction, a polynomial evaluation, and a reconstruction step. These evaluation schemes are designed to give the best performance for a given accuracy, which requires a fine control of errors. One of the main issues is to minimize the number of sources of error and/or their influence on the final result. The work presented in this article addresses this problem as it removes one source of error for the evaluation of trigonometric functions. We propose a method that eliminates rounding errors from tabulated values used in the second range reduction for the sine and cosine evaluation. When targeting correct rounding, we show that such tables are smaller and make the reconstruction step less expensive than existing methods. This approach relies on Pythagorean triples generators. Finally, we show how to generate tables indexed by up to 10 bits in a reasonable time and with little memory consumption.
Hugues de Lassus Saint-Genies, David Defour, Guillaume Revy
ASAP2
2015 Reproducible floating-point atomic addition in data-parallel environment
abstract
Floating-point additions in concurrent execution environment are known to be hazardous, as the result depends on the order in which operations are performed.This problem is encountered in data parallel execution environments such as GPUs, where reproducibility involving floating-point atomic addition is challenging.This problem is due to the rounding error or cancellation that appears for each operation, combined with the lack of control over execution order.In this article we propose two solutions to address this problem: work reassignment and fixed-point accumulation.Work reassignment consists in enforcing an execution order that leads to weak reproducibility.Fixed-point accumulation consists in avoiding rounding errors altogether thanks to a long accumulator and enables strong reproducibility.
David Defour, Caroline Collange
FedCSIS1
2015 Numerical reproducibility for the parallel reduction on multi- and many-core architectures
Caroline Collange, David Defour, Stef Graillat, Roman Iakymchuk
Parallel Comput.2
2014 FuzzyGPU: A Fuzzy Arithmetic Library for GPU
abstract
Data are traditionally represented using native format such as integer or floating-point numbers in various flavor. However, some applications rely on more complex representation format. This is the case when uncertainty needs to be apprehended. Fuzzy arithmetic is one of the major tools to address this problem, but the execution time of basic operations such as addition or multiplication makes its usage prohibitive. In this article, thanks to a new representation format and modern GPU characteristics we show that it is possible to greatly reduce the execution time of those operations. These techniques have been implemented in fuzzyGPU, a freely distributed library of common operations over fuzzy number.
David Defour, Manuel Marin
PDP1
2014 A Pseudo-Random Bit Generator Based on Three Chaotic Logistic Maps and IEEE 754-2008 Floating-Point Arithmetic
Michael François, David Defour, Pascal Berthomé
TAMC2
2010 Implementing LNS using filtering units of GPUs
abstract
Current GPUs offer specialized graphics hardware in addition to generic floating-point processing units. We propose a method which reuses specialized texture filtering units to perform piecewise polynomial evaluations, which helps accelerate LNS computations and can be used in combination with hardware-based transcendental functions.
Mark G. Arnold, Caroline Collange, David Defour
ICASSP3
2010 Barra: A Parallel Functional Simulator for GPGPU
abstract
We present Barra, a simulator of Graphics Processing Units (GPU) tuned for general purpose processing (GPGPU). It is based on the UNISIM framework and it simulates the native instruction set of the Tesla architecture at the functional level. The inputs are CUDA executables produced by NVIDIA tools. No alterations are needed to perform simulations. As it uses parallelism, Barra generates detailed statistics on executions in about the time needed by CUDA to operate in emulation mode. We use it to understand and explore the micro-architecture design spaces of GPUs.
Caroline Collange, Marc Daumas, David Defour, David Parello
MASCOTS3
2007 Graphic processors to speed-up simulations for the design of high performance solar receptors
abstract
Graphics processing units (GPUs) are now powerful and flexible systems adapted and used for other purposes than graphics calculations (general purpose computation on GPU — GPGPU). We present here a prototype to be integrated into simulation codes that estimate temperature, velocity and pressure to design next generations of solar receptors. Such codes will delegate to our contribution on GPUs the computation of heat transfers due to radiations. We use Monte-Carlo line-by-line ray-tracing through finite volumes. This means data-parallel arithmetic transformations on large data structures. Our prototype is inspired on the source code of GPUBench. Our performances on two recent graphics cards (Nvidia 7800GTX and ATI RX1800XL) show some speed-up higher than 400 compared to CPU implementations leaving most of CPU computing resources available. As there were some questions pending about the accuracy of the operators implemented in GPUs, we start this report with a survey and some contributed tests on the various floating point units available on GPUs.
Caroline Collange, Marc Daumas, David Defour
ASAP3
2005 The instruction register file micro-architecture
Bernard Goossens, David Defour
Future Gener. Comput. Syst.2
2005 A New Range-Reduction Algorithm
abstract
Range-reduction is a key point for getting accurate elementary function routines. We introduce a new algorithm that is fast for input arguments belonging to the most common domains, yet accurate over the full double-precision range.
Nicolas Brisebarre, David Defour, Peter Kornerup, Jean-Michel Muller, Nathalie Revol
IEEE Trans. Computers2