Eric Guthmuller

dblp:119/3710 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2024
0009-0002-0678-7599ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A Scalable Low-Latency FPGA Architecture for Spin Qubit Control Through Direct Digital Synthesis
abstract
Scaling qubit control is a key issue for Large Scale Quantum (LSQ) computing and hardware control systems are increasingly costly in logic and memory resources. We present a newly developed compact Direct Digital Synthesis (DDS) architecture for signal generation for spin qubits that is scalable in terms of waveform accuracy and the number of synchronized channels. Fine control of gate voltages is achieved by on-the-fly generation of very precise ramps. Embedded memory requirements are reduced by orders of magnitude compared to current Arbitrary Waveform Generator (AWG) architectures, removing a major scalability barrier for quantum computing.
Mathieu Toubeix, Eric Guthmuller, Adrian Evans, Tristan Meunier
DATE2
2024 Xvpfloat: RISC-V ISA Extension for Variable Extended Precision Floating Point Computation
abstract
A key concern in the field of scientific computation is the convergence of numerical solvers when applied to large problems. The numerical workarounds used to improve convergence are often problem specific, time consuming and require skilled numerical analysts. An alternative is to simply increase the working precision of the computation, but this is difficult due to the lack of efficient hardware support for extended precision. We proposeXvpfloat, a RISC-V ISA extension for dynamically variable and extended precision computation, a hardware implementation and a full software stack. Our architecture provides a comprehensive implementation of this ISA, with up to 512 bits of significand, including full support for common rounding modes and heterogeneous precision arithmetic operations. The memory subsystem handles IEEE 754 extendable formats, and features specialized indexed loads and stores with hardware-assisted prefetching. This processor can either operate standalone or as an accelerator for a general purpose host. We demonstrate that the number of solver iterations can be reduced up to 5× and, for certain, difficult problems, convergence is only possible with very high precision (≥384 bits). This accelerator provides a new approach to accelerate large scale scientific computing.
Eric Guthmuller, César Fuguet Tortolero, Andrea Bocco, Jérôme Fereyre, Riccardo Alidori, Ihsane Tahir, Yves Durand
IEEE Trans. Computers1
2022 Accelerating Variants of the Conjugate Gradient with the Variable Precision Processor
abstract
Linear algebra kernels such as linear solvers, eigen-solvers are the actual working engine underneath many scientific applications. The growing scale of these applications has led researchers to rely on high-precision computing for improving their efficiency and their stability. In this work, we investigate the impact of arbitrary extended precision on multiple variants of the Conjugate Gradient method (CG). We show how our VRP processor improves the convergence and the efficiency of these kernels. We also illustrate how our set of tools (library, software environment) enables to migrate legacy applications in a fast and intuitive way while preserving high-performance. We observe up to an 8X improvements on kernel iteration count, and up to a 40 % improvement on latency. Nevertheless, the main benefit is the stability gained with the precision. It makes it possible to resolve larger and ill-conditioned systems without costly compensating techniques.
Yves Durand, Eric Guthmuller, César Fuguet Tortolero, Jérôme Fereyre, Andrea Bocco, Riccardo Alidori
ARITH2
2021 Storage Class Memory with Computing Row Buffer: A Design Space Exploration
abstract
Today computing centric von Neumann architectures face strong limitations in the data-intensive context of numerous applications, such as deep learning. One of these limitations corresponds to the well known von Neumann bottleneck. To overcome this bottleneck, the concepts of In-Memory Computing (IMC) and Near-Memory Computing (NMC) have been proposed. IMC solutions based on volatile memories, such as SRAM and DRAM, with nearly infinite endurance, solve only partially the data transfer problem from the Storage Class Memory (SCM). Computing in SCM is extremely limited by the intrinsic poor endurance of the Non-Volatile Memory (NVM) technologies. In this paper, we propose to take the best of both solutions, by introducing a Computing Row Buffer (C-RB), using a Computing SRAM (C-SRAM) model, in place of the standard Row Buffer (RB) in the SCM. The principle is to keep operations on large vectors in the C-RB of the SCM, minimizing data movement to and from the CPU, thus drastically reducing energy consumption of the overall system. To evaluate the proposed architecture, we use an instruction accurate platform based on Intel Pin software. Pin instruments run time binaries in order to get applications' full memory traces of our solution. We achieve energy reduction up to 7.9x on average and up to 45x for the best case and speedup up to 3.8x on average and up to 13x for the best case, and a reduction of write accesses in the SCM up to 18 %, compared to SIMD 512-bit architecture.
Valentin Egloff, Jean-Philippe Noël, Maha Kooli, Bastien Giraud, Lorenzo Ciampolini, Roman Gauchi, César Fuguet Tortolero, Eric Guthmuller, Mathieu Moreau, Jean-Michel Portal
DATE8
2018 Dynamic Coherent Cluster: A Scalable Sharing Set Management Approach
abstract
The most widely used programming models expect hardware to guarantee coherent shared memory accesses. However, with the increasing number of integrated cores on chip, resource and performance efficient scalable cache coherence protocols are needed. To address the scalability issues due to the size of the sharing set, we propose to encode, on a fixed size bit-vector, a rectangular cluster whose goal is to cover most of the sharers. The cluster size is fixed but its height, width and position are determined for each cache block and can change during execution. We use a fixed size linked list for the first few outliers, and resort to broadcast when the list overflows. We compare our solution to snoop, directory-based full bit-vector, and Ackwise. It leads to similar mean latency and 10% less traffic than Ackwise, and only a few percent more than the complete sharing set on these metrics. More importantly, it generates ten times less broadcasts than Ackwise while using similar hardware resources for a 64 cores architecture.
Julie Dumas, Eric Guthmuller, Frédéric Pétrot
ASAP2
2013 3D integration for power-efficient computing
abstract
3D stacking is currently seen as a breakthrough technology for improving bandwidth and energy efficiency in multi-core architectures. The expectation is to solve major issues such as external memory pressure and latency while maintaining reasonable power consumption. In this paper, we show some advances in this field of research, starting with memory interface solutions as WIDEIO experience on a real chip for solving DRAM accesses issue. We explain the integration of a 512-bit memory interface in a Network-on-Chip multi-core framework and we show the performance we can achieve, these results being based on a 65nm prototype integrating 10µm diameter Through Silicon Vias. We then present the potentiality of new fine grain 3D stacking technology for power-efficient memory hierarchy. We expose an innovative 3D stacked multi-cache strategy aimed at lowering memory latency and external memory bandwidth requirements and thus demonstrating the efficiency of 3D stacking to rethink architectures for obtaining unequalled performances in power efficiency.
Denis Dutoit, Eric Guthmuller, Ivan Miro Panades
DATE2
2013 3D stacking for multi-core architectures: From WIDEIO to distributed caches
abstract
3D stacking has been viewed as a breakthrough solution for increasing performance in multi-core architectures. The hope is to solve some of the main issues in current multi-core architectures: external memory pressure and latency; I/O bottleneck; communication power consumption. In this paper, some advances of this field of research are shown, starting with a WIDEIO experience on a real chip for solving DRAM accesses issue. The integration of a 512 bit-width bus is demonstrated in a Network-on-Chip (NoC) multi-core framework and the resulting performance based on a 65nm prototype with 10μm diameter Through Silicon Vias (TSV). The potentiality of 3D scaling thanks to 3D asynchronous Network-on-Chip implementation is then shown. Finally, an innovative 3D stacked distributed cache strategy aimed at lowering memory latency and external memory bandwidth requirements is presented. This new memory partitioning demonstrates the efficiency of 3D stacking to rethink architectures for addressing multi-core scaling challenges.
Fabien Clermidy, Denis Dutoit, Eric Guthmuller, Ivan Miro Panades, Pascal Vivet
ISCAS3
2013 Architectural exploration of a fine-grained 3D cache for high performance in a manycore context
abstract
New fine-grained 3D cache architectures have been recently proposed to embed more memory on-chip and thus reduce off-chip memory accesses. These 3D architectures provide a high access bandwidth thanks to wide vertical links. In this paper, we analyze the performances of such caches in a manycore context. We first propose to improve the microarchitecture of an existing 3D non uniform cache architecture. Then we evaluate the impact of the granularity (number of tiles) of this 3D cache on an existing multicore architecture executing high performance computing workloads. We show that the granularity of the 3D cache can affect the performances by a factor of 300%. We also evaluate the impact of the vertical links granularity (number of vertical 3D NoC links) on performances and show that a high number of these links is necessary to achieve the best performanes. Finally, we compare this fine-grained architecture to a memory using a Wide IO interface and show that the latter is less efficient in a manycore context.
Eric Guthmuller, Ivan Miro Panades, Alain Greiner
VLSI-SoC1