EDBT 2026 Demo / reviewers in the wild / expert
Gianmarco Ottavi
dblp:272/2498
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-0041-7917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ramping Up Open-Source RISC-V Cores: Assessing the Energy Efficiency of Superscalar, Out-of-Order ExecutionabstractOpen-source RISC-V cores are increasingly demanded in domains like automotive and space, where achieving high instructions per cycle (IPC) through superscalar and out-of-order (OoO) execution is crucial.However, high-performance open-source RISC-V cores face adoption challenges: some (e.g.BOOM, Xiangshan) are developed in Chisel with limited support from industrial electronic design automation (EDA) tools.Others, like the XuanTie C910 core, use proprietary interfaces and protocols, including non-standard AXI protocol extensions, interrupts, and debug support.In this work, we present a modified version of the OoO C910 core to achieve full RISC-V standard compliance in its debug, interrupt, and memory interfaces.We also introduce CVA6S+, an enhanced version of the dual-issue, industry-supported open-source CVA6 core.CVA6S+ achieves 34.4% performance improvement compared to the scalar configuration.We conduct a detailed performance, area, power, and energy analysis on the superscalar out-of-order C910, superscalar in-order CVA6S+ and vanilla, single-issue in-order CVA6, all implemented in GF22FDX technology and integrated into Cheshire, an open-source modular SoC platform.We examine the performance and efficiency of different microarchitectures using the same ISA, SoC, and implementation with identical technology, tools, and methodologies.The area and performance rankings of CVA6, CVA6S+, and C910 follow expected trends: compared to the scalar CVA6, CVA6S+ shows an area increase of 6% and an IPC improvement of 34.4%, while C910 exhibits a 75% increase in area and a 119.5% improvement in IPC.However, efficiency analysis reveals that CVA6S+ leads in area efficiency (GOPS/mm2), while the C910 is highly competitive in energy efficiency (GOPS/W).This challenges the common belief that high performance in superscalar and out-of-order cores inherently comes at a significant cost in terms of area and energy efficiency. Zexin Fu, Riccardo Tedeschi, Gianmarco Ottavi, Nils Wistoff, César Fuguet Tortolero, Davide Rossi 0001, Luca Benini |
CF | 3 |
| 2023 | Reducing Load-Use Dependency-Induced Performance Penalty in the Open-Source RISC-V CVA6 CPUabstractEmbedded CPUs play a critical role in many modern electronic devices and are commonly used in a range of applications, from IoT edge, to automotive and industrial systems. In particular, application class processors are designed to run operating systems such as Linux, providing a platform for running a broad range of software ecosystems. As such, the performance of these processors is critical for ensuring that these systems can operate efficiently and reliably. However, performance enhancements should be area and power neutral to avoid significant impacts on cost and energy efficiency. In this work, we aim to enhance the performance of CVA6, an Open-Source application class RISC-V core. CVA6's performance has been analyzed with the Embench-IoT benchmark suite, which revealed that load-use dependencies were a key cause of stalls on which CVA6 could be improved. To improve load-use dependency handling, we propose an optimization to the micro-architecture of the processor's backend. Specifically, the backend was redesigned by replacing the scoreboard mechanism of CVA6 with a deeper pipeline that includes a second ALU dedicated to executing instructions with load-use dependencies. The new implementation resulted in a 6.5% improvement in IPC on average and a peak of 29% in applications that suffer heavily from load-use dependencies in Embench-IoT. Additionally, the proposed micro-architecture reduces area and power by 2.5% and increases the clock speed by 4%, leading to an overall improvement of the performance of 11% in instruction throughput and 6.5% more efficiency. Gianmarco Ottavi, Florian Zaruba, Luca Benini, Davide Rossi 0001 |
DSD | 1 |
| 2023 | Dustin: A 16-Cores Parallel Ultra-Low-Power Cluster With 2b-to-32b Fully Flexible Bit-Precision and Vector Lockstep Execution ModeabstractComputationally intensive algorithms such as Deep Neural Networks (DNNs) are becoming killer applications for edge devices. Porting heavily data-parallel algorithms on resource-constrained and battery-powered devices while retaining the flexibility granted by instruction processor-based architectures poses several challenges related to memory footprint, computational throughput, and energy efficiency. Low-bitwidth and mixed-precision arithmetic have been proven to be valid strategies for tackling these problems. We present Dustin, a fully programmable compute cluster integrating 16 RISC-V cores capable of 2- to 32-bit arithmetic and all possible mixed-precision combinations. In addition to a conventional Multiple-Instruction Multiple-Data (MIMD) processing paradigm, Dustin introduces a Vector Lockstep Execution Mode (VLEM) to minimize power consumption in highly data-parallel kernels. In VLEM, a single leader core fetches instructions and broadcasts them to the 15 follower cores. Clock gating Instruction Fetch (IF) stages and private caches of the follower cores leads to 38% power reduction. The cluster, implemented in 65 nm CMOS technology, achieves a peak performance of 58 GOPS and a peak efficiency of 1.15 TOPS/W. Gianmarco Ottavi, Angelo Garofalo, Giuseppe Tagliavini, Francesco Conti 0001, Alfio Di Mauro, Luca Benini, Davide Rossi 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | A Low-Power Transprecision Floating-Point Cluster for Efficient Near-Sensor Data AnalyticsabstractRecent applications in low-power (1-20 mW) near-sensor computing require the adoption of floating-point arithmetic to reconcile high precision results with a wide dynamic range. In this article, we propose a low-power multi-core computing cluster that leverages the fined-grained tunable principles of transprecision computing to provide support to near-sensor applications at a minimum power budget. Our solution – based on the open-source RISC-V architecture – combines parallelization and sub-word vectorization with a dedicated interconnect design capable of sharing floating-point units (FPUs) among the cores. On top of this architecture, we provide a full-fledged software stack support, including a parallel low-level runtime, a compilation toolchain, and a high-level programming model, with the aim to support the development of end-to-end applications. We performed an exhaustive exploration of the design space of the transprecision cluster on a cycle-accurate FPGA emulator, varying the number of cores and FPUs to maximize performance. Orthogonally, we performed a vertical exploration to identify the most efficient solutions in terms of non-functional requirements (operating frequency, power, and area). We conducted an experimental assessment on a set of benchmarks representative of the near-sensor processing domain, complementing the timing results with a post place-&-route analysis of the power consumption. A comparison with the state-of-the-art shows that our solution outperforms the competitors in energy efficiency, reaching a peak of 97 Gflop/s/W on single-precision scalars and 162 Gflop/s/W on half-precision vectors. Finally, a real-life use case demonstrates the effectiveness of our approach in fulfilling accuracy constraints. Fabio Montagna, Stefan Mach, Simone Benatti, Angelo Garofalo, Gianmarco Ottavi, Luca Benini, Davide Rossi 0001, Giuseppe Tagliavini |
IEEE Trans. Parallel Distributed Syst. | 5 |