VLDB 2026 Research / reviewers in the wild / expert
César Fuguet Tortolero
dblp:182/3984 · also César Fuguet
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-0656-2023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order ExecutionabstractOpen-source RISC-V cores are increasingly demanded in domains like automotive and space, where achieving high instructions per cycle (IPC) through superscalar and out-of-order (OoO) execution is crucial.However, high-performance open-source RISC-V cores face adoption challenges: some (e.g.BOOM, Xiangshan) are developed in Chisel with limited support from industrial electronic design automation (EDA) tools.Others, like the XuanTie C910 core, use proprietary interfaces and protocols, including non-standard AXI protocol extensions, interrupts, and debug support.In this work, we present a modified version of the OoO C910 core to achieve full RISC-V standard compliance in its debug, interrupt, and memory interfaces.We also introduce CVA6S+, an enhanced version of the dual-issue, industry-supported open-source CVA6 core.CVA6S+ achieves 34.4% performance improvement compared to the scalar configuration.We conduct a detailed performance, area, power, and energy analysis on the superscalar out-of-order C910, superscalar in-order CVA6S+ and vanilla, single-issue in-order CVA6, all implemented in GF22FDX technology and integrated into Cheshire, an open-source modular SoC platform.We examine the performance and efficiency of different microarchitectures using the same ISA, SoC, and implementation with identical technology, tools, and methodologies.The area and performance rankings of CVA6, CVA6S+, and C910 follow expected trends: compared to the scalar CVA6, CVA6S+ shows an area increase of 6% and an IPC improvement of 34.4%, while C910 exhibits a 75% increase in area and a 119.5% improvement in IPC.However, efficiency analysis reveals that CVA6S+ leads in area efficiency (GOPS/mm2), while the C910 is highly competitive in energy efficiency (GOPS/W).This challenges the common belief that high performance in superscalar and out-of-order cores inherently comes at a significant cost in terms of area and energy efficiency. Zexin Fu, Riccardo Tedeschi, Gianmarco Ottavi, Nils Wistoff, César Fuguet Tortolero, Davide Rossi 0001, Luca Benini |
CF | 5 |
| 2025 | FetchFlare: An Open-Source Strided Data Prefetcher for High-Performance Cache HierarchiesabstractIn recent years, the rise of open-source hardware has transformed the landscape of technology development. In particular, RISC-V has offered hardware designers the possibility of designing processors in a much cheaper way by leveraging a rich ecosystem of open-source designs that can be easily reused, extended, and customized. Although the RISC-V ecosystem is rapidly growing and open-source processors are becoming increasingly sophisticated, some advanced architectural techniques typically employed in commercial high-performance processors are still not prevalent in RISC-V open-source architectures. Among them, hardware prefetchers have been ubiquitous in highend processors for many years, but they are not as commonly found in open-source RISC-V processors. To bridge this gap, this work presents FetchFlare, a stride prefetcher for highperformance cache hierarchies. FetchFlare is able to capture the memory access patterns of applications, predict future memory accesses, and issue prefetch requests for them. We provide an open-source RTL implementation of FetchFlare and integrate it into a complete open-source setup formed by the OpenPiton framework, the Sargantana core, and the High-Performance Data Cache (HPDCache). Compared to a baseline system without prefetching, FetchFlare achieves an average speedup of $63 \%$, avoids cache misses in the L1D and the L2 caches, and presents an average accuracy, coverage, and timeliness of $86 \%, 39 \%$, and 99%, respectively. Golnaz Korkian, Neiel Leyva, Arnau Bigas, Noelia Oliete-Escuín, Abbas Haghi, Alireza Monemi, César Fuguet Tortolero, Lluc Alvarez |
DSD | 7 |
| 2024 | Breaking the Memory Wall with a Flexible Open-Source L1 Data-CacheabstractThe lack of concurrency and pipelining in the memory sub-system of recent open-source RISC-V processors, such as the CVA611https://github.com/openhwgroup/cva6, is increasingly becoming the performance bottleneck. Recent updates to the new RISC-V High Performance Ll Data-cache (HPDcache), now fully integrated with the CVA6, bring significant (up to +234%) speedups in key benchmarks with a negligible 5.92% area impact. In this short paper, we detail these improvements, compare performance with existing caches and highlight the benefits of this new, open-source data-cache. Davy Million, Noelia Oliete-Escuín, César Fuguet Tortolero |
DATE | 3 |
| 2024 | OpenSource Heterogeneous Chiplet-based Computing ArchitecturesabstractLeading edge processors, such as AMD's MI300 and Intel's Ponte Vecchio, rely on 3D integration of heterogeneous architectures including CPUs and GPUs or vector processors to provide the highest performance. There are many models for memory coherency within such chips and there is a need for research on the best way to map various kernels to these architectures including studying how best to share data. Unfortunately, the hardware in commercial processors is closed, which limits research opportunities, particularly for hardware/software co-design. Recent projects such as OpenPiton, and numerous follow-on projects, have made it possible for the research community to develop coherent, multi-core systems based on RISC-V. The next step is to enhance the support in open source, multi-core platforms for heterogeneous computing (GPUs, FPGA accelerators) and to explore 3D partitioning of such systems. In this paper, we present the state-of-the art of such open source platforms and sketch a roadmap which we hope will enable the research community to continue to contribute to the development of today's advanced heterogeneous architectures. Adrian Evans, César Fuguet Tortolero, Davy Million |
ICCAD | 2 |
| 2024 | Page size exploration for RISC-V systems: the case for HPCabstractThe page size used for virtual to physical address translation has globally not changed since the late 1960’s: the IBM 360, circa 1964, already had 4 KiB pages. This 4 KiB page size has proven to be incredibly robust given the changes in processor architectures, workloads behavior, memory size, and access patterns. However, with 64-bit registers, 57-bit virtual addresses, and increasingly bigger physical memories, we have to ask ourselves whether 4 KiB is still an adequate page size for modern workloads on modern machines. Inherently, the page size has an influence on (a) the miss rate of the translation lookaside buffer, the cache that contains the recently used virtual to physical translations, and (b) the memory allocated by the system versus the memory actually used by a process. The page size also constraints some microarchitectural choices, such as cache design, which impacts the overall performance and energy efficiency. We focus more particularly on High Performance Computing (HPC) applications because they are extremely demanding in terms of memory, and are indicative of future general-purpose needs.In this paper, we empirically study the evolution of the miss rate and memory occupancy with respect to the page size, and conclude that a page size of 32 KiB is better suited for current HPC systems. We also propose a page table scheme for RISC-V-based HPC systems based on our observations and discuss its benefits. Eduardo Tomasi, César Fuguet Tortolero, Christian Fabre, Frédéric Pétrot |
RSP | 2 |
| 2024 | Xvpfloat: RISC-V ISA Extension for Variable Extended Precision Floating Point ComputationabstractA key concern in the field of scientific computation is the convergence of numerical solvers when applied to large problems. The numerical workarounds used to improve convergence are often problem specific, time consuming and require skilled numerical analysts. An alternative is to simply increase the working precision of the computation, but this is difficult due to the lack of efficient hardware support for extended precision. We proposeXvpfloat, a RISC-V ISA extension for dynamically variable and extended precision computation, a hardware implementation and a full software stack. Our architecture provides a comprehensive implementation of this ISA, with up to 512 bits of significand, including full support for common rounding modes and heterogeneous precision arithmetic operations. The memory subsystem handles IEEE 754 extendable formats, and features specialized indexed loads and stores with hardware-assisted prefetching. This processor can either operate standalone or as an accelerator for a general purpose host. We demonstrate that the number of solver iterations can be reduced up to 5× and, for certain, difficult problems, convergence is only possible with very high precision (≥384 bits). This accelerator provides a new approach to accelerate large scale scientific computing. Eric Guthmuller, César Fuguet Tortolero, Andrea Bocco, Jérôme Fereyre, Riccardo Alidori, Ihsane Tahir, Yves Durand |
IEEE Trans. Computers | 2 |
| 2023 | HPDcache: Open-Source High-Performance L1 Data Cache for RISC-V CoresabstractFor many compute applications the performance bottleneck is the memory bandwidth and latency. This is particularly true in the domain of High-Performance Computing (e.g. scientific applications). Cache memories are an essential component of modern processors, which further pushes the "Memory Wall". Caches in the domain of HPC must enable both high memory throughput and energy efficiency César Fuguet Tortolero |
CF | 1 |
| 2022 | Accelerating Variants of the Conjugate Gradient with the Variable Precision ProcessorabstractLinear algebra kernels such as linear solvers, eigen-solvers are the actual working engine underneath many scientific applications. The growing scale of these applications has led researchers to rely on high-precision computing for improving their efficiency and their stability. In this work, we investigate the impact of arbitrary extended precision on multiple variants of the Conjugate Gradient method (CG). We show how our VRP processor improves the convergence and the efficiency of these kernels. We also illustrate how our set of tools (library, software environment) enables to migrate legacy applications in a fast and intuitive way while preserving high-performance. We observe up to an 8X improvements on kernel iteration count, and up to a 40 % improvement on latency. Nevertheless, the main benefit is the stability gained with the precision. It makes it possible to resolve larger and ill-conditioned systems without costly compensating techniques. Yves Durand, Eric Guthmuller, César Fuguet Tortolero, Jérôme Fereyre, Andrea Bocco, Riccardo Alidori |
ARITH | 3 |
| 2021 | Storage Class Memory with Computing Row Buffer: A Design Space ExplorationabstractToday computing centric von Neumann architectures face strong limitations in the data-intensive context of numerous applications, such as deep learning. One of these limitations corresponds to the well known von Neumann bottleneck. To overcome this bottleneck, the concepts of In-Memory Computing (IMC) and Near-Memory Computing (NMC) have been proposed. IMC solutions based on volatile memories, such as SRAM and DRAM, with nearly infinite endurance, solve only partially the data transfer problem from the Storage Class Memory (SCM). Computing in SCM is extremely limited by the intrinsic poor endurance of the Non-Volatile Memory (NVM) technologies. In this paper, we propose to take the best of both solutions, by introducing a Computing Row Buffer (C-RB), using a Computing SRAM (C-SRAM) model, in place of the standard Row Buffer (RB) in the SCM. The principle is to keep operations on large vectors in the C-RB of the SCM, minimizing data movement to and from the CPU, thus drastically reducing energy consumption of the overall system. To evaluate the proposed architecture, we use an instruction accurate platform based on Intel Pin software. Pin instruments run time binaries in order to get applications' full memory traces of our solution. We achieve energy reduction up to 7.9x on average and up to 45x for the best case and speedup up to 3.8x on average and up to 13x for the best case, and a reduction of write accesses in the SCM up to 18 %, compared to SIMD 512-bit architecture. Valentin Egloff, Jean-Philippe Noël, Maha Kooli, Bastien Giraud, Lorenzo Ciampolini, Roman Gauchi, César Fuguet Tortolero, Eric Guthmuller, Mathieu Moreau, Jean-Michel Portal |
DATE | 7 |
| 2020 | POPSTAR: a Robust Modular Optical NoC Architecture for Chiplet-based 3D Integrated SystemsabstractSilicon photonics technology is now gaining maturity with increasing levels of design complexity from devices to large photonic integrated circuits. Close integration of control electronics with 3D assembly of photonics and CMOS opens the way to high-performance computing architectures partitioned in chiplets connected by optical NoC on silicon photonic interposers. In this paper, we give an overview of our works on optical links and NoC for manycore systems, from low-level control of photonic devices to high-level system optimization of the optical communications. We detail the POPSTAR optical NoC topology and architecture (Processors On Photonic Silicon interposer Terascale ARchitecture) with electro-optical interface chiplets, the corresponding nested spiral topology for single-writer multiple- reader links and the associated control electronics, in charge of high-speed drivers, thermal stabilization and handling of the protocol stack, from data integrity to flow-control, routing and arbitration of the optical communications. The strengths and opportunities for this architecture will be discussed, with a shift in system & implementation constraints with respect to previous optical NoC proposals, and new challenges to be addressed. Yvain Thonnart, Stéphane Bernabé, Jean Charbonnier, Christian Bernard, David Coriat, César Fuguet Tortolero, Pierre Tissier, Benoît Charbonnier, Stephane Malhouitre, Damien Saint-Patrice, Myriam Assous, Aditya Narayan, Ayse K. Coskun, Denis Dutoit, Pascal Vivet |
DATE | 6 |
| 2019 | WAVES: Wavelength Selection for Power-Efficient 2.5D-Integrated Photonic NoCsabstractPhotonic Network-on-Chips (PNoCs) offer promising benefits over Electrical Network-on-Chips (ENoCs) in many-core systems owing to their lower latencies, higher bandwidth, and lower energy-per-bit communication with negligible data-dependent power. These benefits, however, are limited by a number of challenges. Microring resonators (MRRs) that are used for photonic communication have high sensitivity to process variations and on-chip thermal variations, giving rise to possible resonant wavelength mismatches. State-of-the-art microheaters, which are used to tune the resonant wavelength of MRRs, have poor efficiency resulting in high thermal tuning power. In addition, laser power and high static power consumption of drivers, serializers, comparators, and arbitration logic partially negate the benefits of the sub-pJ operating regime that can be obtained with PNoCs. To reduce PNoC power consumption, this paper introduces WAVES, a wavelength selection technique to identify and activate the minimum number of laser wavelengths needed, depending on an application's bandwidth requirement. Our results on a simulated 2.5D manycore system with PNoC demonstrate an average of 23% (resp. 38%) reduction in PNoC power with only <;1% (resp. <;5%) loss in system performance. Aditya Narayan, Yvain Thonnart, Pascal Vivet, César Fuguet Tortolero, Ayse K. Coskun |
DATE | 4 |
| 2017 | A Programmable Inbound Transfer Processor for Active Messages in Embedded Multicore SystemsabstractThe "Internet of Things" requires new multicore computing devices with very high energy-efficiency. We propose an improved architecture of these embedded devices with emphasis on the efficiency of data transfers. By performing data re-organization at transport layer within the NoC infrastructure, we avoid the need for intermediate buffers for data distribution and organization. To complement "smart DMA” that structure the traffic at source side, we use a simple programmable processor to reorganize incoming data at target side. By doing this, only useful data is transported on the network, and unpacking at destination restores their structure in the most suitable way for the application, without the need of duplication. We have prototyped an inbound data processor in a MIPS-based multicore architecture. Applied on an image compression application, we save up to 33% memory footprint and divide the processing latency by a factor of 2. Yves Durand, Christian Bernard, Romain Lemaire, César Fuguet Tortolero, Emilie Garat |
DSD | 4 |