Yves Durand

dblp:27/5990 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0002-3146-8461ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7Software engineering, systems software and programming languages · 3 · 1 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 SpDCache: Region-Based Reduction Cache for Outer-Product Sparse Matrix Kernels
abstract
Improvements in computer performance depend increasingly on specialized accelerators and recently, numerous architectures optimized for sparse matrix kernels have been proposed, however, they do not exploit the structural properties of the matrices. SpDCache is a cache for outer-product Sparse Matrix-Vector Multiplication (SpMV) which has storage strategies optimized for both dense and sparse regions and which performs reductions locally in this cache. Real world matrices typically have a dense band which benefits from being blocked in the dense region of our cache, while the sparse regions benefit from fine-grained storage and a shift of the computation close to the main memory. We present the architectural principals of SpDCache and show that it reduces main memory traffic by -8x and increases the cache utilization by - 2x for banded matrices.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ASAP3
2024 Dedicated Hardware Accelerators for Processing of Sparse Matrices and Vectors: A Survey
abstract
Performance in scientific and engineering applications such as computational physics, algebraic graph problems or Convolutional Neural Networks (CNN), is dominated by the manipulation of large sparse matrices—matrices with a large number of zero elements. Specialized software using data formats for sparse matrices has been optimized for the main kernels of interest: SpMV and SpMSpM matrix multiplications, but due to the indirect memory accesses, the performance is still limited by the memory hierarchy of conventional computers. Recent work shows that specific hardware accelerators can reduce memory traffic and improve the execution time of sparse matrix multiplication, compared to the best software implementations. The performance of these sparse hardware accelerators depends on the choice of the sparse format, COO , CSR , etc, the algorithm, inner-product , outer-product , Gustavson , and many hardware design choices. In this article, we propose a systematic survey which identifies the design choices of state-of-the-art accelerators for sparse matrix multiplication kernels. We introduce the necessary concepts and then present, compare, and classify the main sparse accelerators in the literature, using consistent notations. Finally, we propose a taxonomy for these accelerators to help future designers make the best choices depending on their objectives.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ACM Trans. Archit. Code Optim.3
2024 Xvpfloat: RISC-V ISA Extension for Variable Extended Precision Floating Point Computation
abstract
A key concern in the field of scientific computation is the convergence of numerical solvers when applied to large problems. The numerical workarounds used to improve convergence are often problem specific, time consuming and require skilled numerical analysts. An alternative is to simply increase the working precision of the computation, but this is difficult due to the lack of efficient hardware support for extended precision. We proposeXvpfloat, a RISC-V ISA extension for dynamically variable and extended precision computation, a hardware implementation and a full software stack. Our architecture provides a comprehensive implementation of this ISA, with up to 512 bits of significand, including full support for common rounding modes and heterogeneous precision arithmetic operations. The memory subsystem handles IEEE 754 extendable formats, and features specialized indexed loads and stores with hardware-assisted prefetching. This processor can either operate standalone or as an accelerator for a general purpose host. We demonstrate that the number of solver iterations can be reduced up to 5× and, for certain, difficult problems, convergence is only possible with very high precision (≥384 bits). This accelerator provides a new approach to accelerate large scale scientific computing.
Eric Guthmuller, César Fuguet Tortolero, Andrea Bocco, Jérôme Fereyre, Riccardo Alidori, Ihsane Tahir, Yves Durand
IEEE Trans. Computers7
2022 Accelerating Variants of the Conjugate Gradient with the Variable Precision Processor
abstract
Linear algebra kernels such as linear solvers, eigen-solvers are the actual working engine underneath many scientific applications. The growing scale of these applications has led researchers to rely on high-precision computing for improving their efficiency and their stability. In this work, we investigate the impact of arbitrary extended precision on multiple variants of the Conjugate Gradient method (CG). We show how our VRP processor improves the convergence and the efficiency of these kernels. We also illustrate how our set of tools (library, software environment) enables to migrate legacy applications in a fast and intuitive way while preserving high-performance. We observe up to an 8X improvements on kernel iteration count, and up to a 40 % improvement on latency. Nevertheless, the main benefit is the stability gained with the precision. It makes it possible to resolve larger and ill-conditioned systems without costly compensating techniques.
Yves Durand, Eric Guthmuller, César Fuguet Tortolero, Jérôme Fereyre, Andrea Bocco, Riccardo Alidori
ARITH1
2021 Seamless Compiler Integration of Variable Precision Floating-Point Arithmetic
abstract
Floating-Point (FP) units in processors are generally limited to supporting a subset of formats defined by the IEEE 754 standard. As a result, high-efficiency languages and optimizing compilers for high-performance computing only support IEEE standard types and applications needing higher precision involve cumbersome memory management and calls to external libraries, resulting in code bloat and making the intent of the program unclear. We present an extension of the C type system that can represent generic FP operations and formats, supporting both static precision and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and low-level FP instructions or software libraries. The effectiveness of our solution is demonstrated through an LLVM-based implementation, leveraging aggressive optimizations in LLVM including the Polly loop nest optimizer, which targets two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision floating-point library. Our optimizing compilation flow targeting MPFR outperforms the Boost programming interface for the MPFR library by a factor of 1.80 × and 1.67 × in sequential execution of the Poly Bench and RAJAPerf suites, respectively, and by a factor of 7.62 x on an 8-core (and 16-thread) machine for RAJAPerf in OpenMP.
Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot
CGO2
2020 VP Float: First Class Treatment for Variable Precision Floating Point Arithmetic
abstract
Optimizing compilers for high performance computing only support IEEE~754 floating-point (FP) types and applications needing higher precision involve cumbersome memory management and calls to external libraries. We introduce an extension of the C type system to represent variable-precision FP arithmetic, supporting both static and dynamically variable precision. We design and implement a compilation flow bridging the abstraction gap between this type system and hardware FP instructions or software libraries. We demonstrate the effectiveness of our solution by enabling the full range of LLVM optimizations and leveraging two backend code generators: one for the ISA of a variable precision FP arithmetic coprocessor, and one for the MPFR multi-precision FP library. Both targets support the static and dynamically adaptable precision of our type system. On the PolyBench suite, our optimizing compilation flow targeting MPFR is shown to outperform the Boost programming interface for the MPFR library.
Tiago T. Jost, Yves Durand, Christian Fabre, Albert Cohen 0001, Frédéric Pétrot
PACT2
2019 Dynamic Precision Numerics Using a Variable-Precision UNUM Type I HW Coprocessor
abstract
A very large internal accumulation register has been proposed to increase the accuracy of scientific code. However, there is a general class of iterative kernels where a vector of high-precision data must be saved from one iteration to the next. Saving the large internal accumulator to memory is impractical in such cases. This work proposes a Variable Precision (VP) Floating Point (FP) arithmetic co-processor architecture based on RISC-V, which 1/ supports legacy IEEE formats for input and output variables, 2/ uses variable length internal registers (up to 512 bits of mantissa) for inner loop multiply-add and 3/ supports loads and stores of intermediate results to cache memory with a dynamically adjustable precision (up to 256 bits of mantissa). It exploits the UNUM type I floating point format, proposing solutions to address some of its pitfalls such as the variable latency of the internal operation, and the variable memory footprint of the intermediate variables. This work is integrated on FPGA and demonstrated on a representative example.
Andrea Bocco, Yves Durand, Florent de Dinechin
ARITH2
2019 Error Analysis of the Square Root Operation for the Purpose of Precision Tuning: A Case Study on K-means
abstract
In this paper, we propose an analytical approach to study the impact of floating point (FLP) precision variation on the square root operation, in terms of computational accuracy and performance gain. We estimate the round-off error resulting from reduced precision. We also inspect the Newton Raphson algorithm used to approximate the square root in order to bound the error caused by algorithmic deviation. Consequently, the implementation of the square root can be optimized by fittingly adjusting its number of iterations with respect to any given FLP precision specification, without the need for long simulation times. We evaluate our error analysis of the square root operation as part of approximating a classic data clustering algorithm known as K-means, for the purpose of reducing its energy footprint. We compare the resulting inexact K-means to its exact counterpart, in the context of color quantization, in terms of energy gain and quality of the output. The experimental results show that energy savings could be achieved without penalizing the quality of the output (e.g., up to 41.87% of energy gain for an output quality, measured using structural similarity, within a range of [0.95,1]).
Oumaima Matoussi, Yves Durand, Olivier Sentieys, Anca Mariana Molnos
ASAP2
2019 Evaluation of variable bit-width units in a RISC-V processor for approximate computing
abstract
Among various power reduction methods, variable bit-width arithmetic units have been proposed in approximate computing literature. In this paper, we add a variable bit-width memory unit in a RISC-V processor. Integrating both computation and memory units with variable bit-width leads to a power reduction: from 7% to 29% for Sobel filter application and from 13% to 24% for an application that computes the position of a robotic arm (forwardk2j). We also propose a global energy model for a RISC-V processor with variable bit-width units (for computation and memory). This model allows us to evaluate the impact of various parameters in both the software application (e.g., the amount of instructions that can be executed with a reduced bit-width) and the hardware architecture (e.g., impact of potential reduction for each unit).
Geneviève Ndour, Tiago T. Jost, Anca Mariana Molnos, Yves Durand, Arnaud Tisserand
CF4
2019 Byte-Aware Floating-point Operations through a UNUM Computing Unit
abstract
Most floating-point (FP) hardware support the IEEE 754 format, which defines fixed-size data types from 16 to 128 bits. However, a range of applications benefit from different formats, implementing different tradeoffs. This paper proposes a Variable Precision (VP) computing unit offering a finer granularity of high precision FP operations. The chosen memory format is derived from UNUM type I, where the size of a number is stored within the representation itself. The unit implements a fully pipelined architecture, and it supports up to 512 bits of precision for both interval and scalar computing. The user can conFigure the storage format up to 8-bit granularity, and the internal computing precision at 64-bit granularity. The system is integrated as a RISC-V coprocessor. Dedicated compiler support exposes the unit through a high level programming abstraction, covering all the operating features of UNUM type I. FPGA-based measurements show that the latency and the computation accuracy of this system scale linearly with the memory format length set by the user. Compared with a highly optimized software implementation, the proposed unit achieves speedups between 3.5 × and 18 ×, with comparable accuracy.
Andrea Bocco, Tiago T. Jost, Albert Cohen 0001, Florent de Dinechin, Yves Durand, Christian Fabre
VLSI-SoC5
2019 Fine-Grain Back Biasing for the Design of Energy-Quality Scalable Operators
abstract
Energy-quality scalable systems are a promising solution to cope with the small energy budgets and high processing demands of mobile and Internet of Things applications. These systems leverage the error resilience of applications to obtain high energy efficiency, at the expense of tolerable reductions in the output quality. Hardware datapath operators able to reconfigure their precision and power consumption at runtime are key components of such systems. However, most implementations of these operators require manual, architecture-specific modifications and tend to have large power overheads compared to standard designs, when working at maximum precision. One promising design-independent alternative is dynamic voltage and accuracy scaling, whose adoption, however, is hindered by incompatibilities with standard design flows. In this paper, we propose a new methodology for the design of energy-quality scalable operators; our solution leverages runtime tuning of transistors threshold voltages to obtain a fine-grain control of the speed and power consumption of standard-cells within an operator. Thanks to the additional flexibility provided by this fine-grain knob, our method overcomes the main limitations of previous solutions, at the cost of a small area overhead. We demonstrate our approach on a 28 nm FDSOI technology; by exploiting the strong effect of back-gate biasing on threshold voltage, we achieve a power consumption reduction of more than 40% compared to the state-of-the-art, for the same precision.
Daniele Jahier Pagliari, Yves Durand, David Coriat, Edith Beigné, Enrico Macii, Massimo Poncino
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 A methodology for the design of dynamic accuracy operators by runtime back bias
abstract
Mobile and IoT applications must balance increasing processing demands with limited power and cost budgets. Approximate computing achieves this goal leveraging the error tolerance features common in many emerging applications to reduce power consumption. In particular, adequate (i.e., energy/quality-configurable) hardware operators are key components in an error tolerant system. Existing implementations of these operators require significant architectural modifications, hence they are often design-specific and tend to have large overheads compared to accurate units. In this paper, we propose a methodology to design adequate data-path operators in an automatic way, which uses threshold voltage scaling as a knob to dynamically control the power/accuracy tradeoff. The method overcomes the limitations of previous solutions based on supply voltage scaling, in that it introduces lower overheads and it allows fine-grain regulation of this tradeoff. We demonstrate our approach on a state-of-the-art 28nm FDSOI technology, exploiting the strong effect of back biasing on threshold voltage. Results show a power consumption reduction of as much as 39% compared to solutions based only on supply voltage scaling, at iso-accuracy.
Daniele Jahier Pagliari, Yves Durand, David Coriat, Anca Mariana Molnos, Edith Beigné, Enrico Macii, Massimo Poncino
DATE2
2017 A Programmable Inbound Transfer Processor for Active Messages in Embedded Multicore Systems
abstract
The "Internet of Things" requires new multicore computing devices with very high energy-efficiency. We propose an improved architecture of these embedded devices with emphasis on the efficiency of data transfers. By performing data re-organization at transport layer within the NoC infrastructure, we avoid the need for intermediate buffers for data distribution and organization. To complement "smart DMA” that structure the traffic at source side, we use a simple programmable processor to reorganize incoming data at target side. By doing this, only useful data is transported on the network, and unpacking at destination restores their structure in the most suitable way for the application, without the need of duplication. We have prototyped an inbound data processor in a MIPS-based multicore architecture. Applied on an image compression application, we save up to 33% memory footprint and divide the processing latency by a factor of 2.
Yves Durand, Christian Bernard, Romain Lemaire, César Fuguet Tortolero, Emilie Garat
DSD1
2016 EUROSERVER: Share-anything scale-out micro-server design
Manolis Marazakis, John Goodacre, Didier Fuin, Paul M. Carpenter, John Thomson, Emil Matús, Antimo Bruno, Per Stenström, Jérôme Martin, Yves Durand, Isabelle Dor
DATE10
2014 EUROSERVER: Energy Efficient Node for European Micro-Servers
abstract
EUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers.
Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson
DSD1
2014 Dry snow analysis in alpine regions using RADARSAT-2 full polarimetry data. Comparison with in situ measurements
abstract
In this paper we describe the benefits of RADARSAT-2 in the analysis of temporal changes in polarimetric parameters linked to the snow cover evolution during the winter season. The presented study took place over an instrumented area in the region of French Alps. The focus is set on the dry snow depth retrieval, using an original method based on principal component statistical analysis (PCA) of the polarimetric parameters values. The results obtained by this mean are compared with the network of simultaneous snow in situ measurements. The most thought-provoking result is the strong inverse correlation between the snow depth above the ice crust and the entropy, reflected through the very high coefficient of determination R2= 0.8439. In order to justify this observation, we propose an appropriate physical hypothesis.
Jean-Pierre Dedieu, Nikola Besic, Gabriel Vasile, Julien Mathieu, Yves Durand, Frederic Gottardi
IGARSS5
2014 Comparison between DMRT simulations for multilayer snowpack and data from NoSREx report
abstract
This paper presents a multilayer snowpack Electromagnetic Backscattering Model (EBM), based on Dense Media Radiative Transfer (DMRT). This model is capable of simulating the interaction of electromagnetic waves (EMW) at X-band and Ku-band frequencies with multilayer snowpack. The air-snow interface and snow-ground backscattering components are calculated using the Integral Equation Model (IEM), Fung et al. [1], whereas the volume backscattering component is calculated by the solution of Vector Radiative Transfer (VRT) equation at order 1. We have applied these models using measurement data from NoSREx report [2], which includes SnowScat data in X-band and Ku-band, TerraSAR-X acquisitions and snowpack stratigraphic profiles. The results of model simulations show consistency with the radar observations, and therefore allow the EBM to be used in various applications, such as data assimilation [3].
Xuan-Vu Phan, Laurent Ferro-Famil, Michel Gay, Yves Durand, Marie Dumont
IGARSS4
2014 Assimilation of TerraSAR-X data into a snowpack model
abstract
This paper presents an approach using data assimilation to take into account X-band Synthetic Aperture Radar (SAR) satellite observations in a detailed snowpack model. The SURFEX/Crocus snow model, developed by MeteoFrance, is used to simulate the detail stratigraphy of multilayer snow-pack from meteorological conditions. The Dense Media Radiative Transfer (DMRT) model allows the simulation of SAR backscattering coefficient using the physical parameters of snowpack (density and optical grain diameter of each layer). The development of an adjoint model of the DMRT and the implementation of three-dimensional variational (3D-Var) data assimilation algorithm enable us to reduce the discrepancy between the backscattering coefficients simulated using DMRT and measured using SAR satellite, through modifying the previously said physical parameters of snowpack. This approach provides the ability to constrain the detailed snowpack model SURFEX/Crocus using SAR observations and allows the retrieval of snowpack properties such as Snow Water Equivalent (SWE). Case study has been carried out using a time series of TerraSAR-X acquisitions on Argentière glacier (Chamonix Mont Blanc, France) in winter 2008-2009.
Xuan-Vu Phan, Michel Gay, Laurent Ferro-Famil, Yves Durand, Marie Dumont
IGARSS4
2012 Multi-temporal wet snow mapping in alpine context using polarimetric Radarsat-2 time-series
abstract
The temporal monitoring of snow cover in mountainous areas is a very important challenge in order to predict snow melting. A way to solve this problematic is the snow detection using polarimetric SAR data. By this way, a set of eight Radarsat-2 images have been acquired over the French Alps during 2009 and 2010. This project proposes to apply snow detection algorithms with some improvements to map dry/no snow cover over the study area under different snow conditions.
Audrey Lessard-Fontaine, Sophie Allain-Bailhache, Jean-Pierre Dedieu, Yves Durand
IGARSS4
2012 Multilayer snowpack backscattering model and assimilation of TerraSAR-X satellite data
abstract
The advantages of the new generation of radar systems with high resolution image, short revisit time provide the possibility of characterization and monitoring the evolution of the cryoshpere. In this paper, we propose an adaptation of the multilayer snow backscattering model based on radiative transfer theory in order to estimate the total backscattering coefficient of high frequency (X-band) electromagnetic wave on snowcover area. Next, from the physical model, we develop the adjoint operator and implement a variational assimilation scheme in order to constrain the snow stratigraphy profiles calculated by CROCUS, a snow metamorphism model used by MeteoFrance. Some tests are carried out with TerraSAR-X image data. The results show that the snow stratigraphy profiles obtained after the data assimilation process have good agreement with the measured profiles, and therefore show the high potential of this method in constraining the snowpack profiles of CROCUS.
Xuan-Vu Phan, Laurent Ferro-Famil, Michel Gay, Yves Durand, Marie Dumont, Guy D'Urso
IGARSS4
2009 The Radio Virtual Machine: A solution for SDR portability and platform reconfigurability
abstract
Instead of a single circuit dedicated to a particular physical (PHY) layer standard, a Software Defined Radio (SDR) platform embeds several hardware accelerators which enable it to support different modulation schemes. In this study we propose an architecture for a SDR PHY layer based on the Virtual Machine (VM) concept. Once a program is compiled in a portable byte-code, the VM can then execute it to manage the desired PHY layer. We demonstrate the feasibility of the proposed architecture through a case study and a proof-of-concept implementation.
Riadh Ben Abdallah, Tanguy Risset, Antoine Fraboulet, Yves Durand
IPDPS4
2009 Snowpack Characterization in Mountainous Regions Using C-Band SAR Data and a Meteorological Model
abstract
This paper presents a method to characterize snow cover in mountainous regions using dual-polarization C-band synthetic aperture radar (SAR) data. It is demonstrated that an accurate modeling of the liquid water distribution inside the snowpack, using a multilayer meteorological snow model, is required to characterize snow with precision. A multilayer-snow electromagnetic (EM) backscattering model is developed based on the vector radiative transfer, the strong fluctuation theory, and physical parameters supplied by the meteorological model. However, the limited resolution of the meteorological snow model is insufficient for predicting a refined EM backscattering at a massif scale. An adequate spatial reorganization of these snow profiles, based on a comparison between simulated and measured dual-polarization SAR data, leads to a better estimation of some snowpack parameters. In particular, the monitoring of snow liquid water content is presented improving the capacity of wet snow mapping as compared to a classical SAR-based method. This methodology shows good capacities both for qualitative and quantitative snow assessments, opening the way for a new operational method.
Nicolas Longépé, Sophie Allain-Bailhache, Laurent Ferro-Famil, Eric Pottier, Yves Durand
IEEE Trans. Geosci. Remote. Sens.5
2009 Methods for power optimization in SOC-based data flow systems
abstract
Whereas the computing power of DSP or general-purpose processors was sufficient for 3G baseband telecommunication algorithms, stringent timing constraints of 4G wireless telecommunication systems require computing-intensive data-driven architectures. Managing the complexity of these systems within the energy constraints of a mobile terminal is becoming a major challenge for designers. System-level low-power policies have been widely explored for generic software-based systems, but data-flow architectures used for high data-rate telecommunication systems feature heterogeneous components that require specific configurations for power management. In this study, we propose an innovative power optimization scheme tailored to self-synchronized data-flow systems. Our technique, based on the synchronous data-flow modeling approach, takes advantage of the latest low-power techniques available for digital architectures. We illustrate our optimization method on a complete 4G telecommunication baseband modem and show the energy savings expected by this technique considering present and future silicon technologies.
Philippe Grosse, Yves Durand, Paul Feautrier
ACM Trans. Design Autom. Electr. Syst.2
2003 Radiometric and geometric correction of RADARSAT-1 images acquired in alpine regions for mapping the snow water equivalent (SWE)
abstract
In this paper, introduced is an application of two radiometric slope correction methods on standard RADARSAT images in a mountainous environment like the Alps. Because of the highly varying topography, such corrections are needed to reduce the distortions on the backscattering coefficients when trying to monitor the snow characteristics from SAR data in alpine regions. This paper discusses the results obtained by the two different methods over dry and wet snow cover; both algorithms significantly reduced the effect of local slope facing the radar, but may not compensate enough for the steep slope over 30/spl deg/.
Jean-Pierre Dedieu, Yves Gauthier, Monique Bernier, Stéphane Hardy, Pierre Vincent, Yves Durand
IGARSS6