Gabriele Montanaro

dblp:319/6596 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-1119-2629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Omega: A Hardware-Software Framework for Complete Design Space Exploration of FPGA-Based Heterogeneous Multi-Core SoCs
abstract
The design space exploration (DSE) of heterogeneous multi-core systems-on-chip (SoCs) presents a massive challenge due to the vast and complex configuration space, demanding simultaneous optimization of performance, energy efficiency, and resource utilization under diverse constraints. Traditional approaches leveraging analytical models, heuristics, and machine learning (ML) techniques fail to comprehensively cover this space, often yielding suboptimal solutions. The Omega framework, designed for exhaustive DSE of FPGA-based heterogeneous multi-core SoCs, addresses the former limitations by fully exploring the design space and guaranteeing the identification of globally optimal configurations. Omega leverages FPGAs’ dynamic partial reconfiguration to accelerate the DSE drastically and can serve as a golden model for evaluating novel DSE heuristics and ML methods. This manuscript demonstrates the proposed framework’s effectiveness through an extensive experimental campaign on 16-core SoCs with accelerators for up to five different applications, achieving a substantial speedup, of 29 times on average, compared to traditional techniques while ensuring solution optimality. Omega is released as a comprehensive open-source ecosystem, compatible with commercially available FPGA platforms, to facilitate future research and practical adoption, setting a new benchmark for DSE methodologies and providing a robust tool for optimizing next-generation computing platforms.
Gabriele Montanaro, Andrea Galimberti, Davide Zoni
IEEE Trans. Computers1
2026 FARMER: Online-Learning-Based Workload Consolidation on Large FPGAs Accelerated With Dynamic Partial Reconfiguration
abstract
As the demand for performance and scalability in cloud applications continues to grow, high-performance computing (HPC) facilities increasingly integrate FPGAs to accelerate computational workloads. To fully utilize the extensive resources available on modern high-end FPGAs, it is essential to optimize the allocation of multiple applications on a single device. This article introduces FARMER, a novel online learning methodology that leverages machine learning (ML) to model the throughput of different applications running concurrently on the same FPGA. It combines this with a sequential decision-making strategy and an in-circuit exploration flow based on dynamic partial reconfiguration (DPR) to drastically speed up the exploration of large design spaces. Experimental evaluations across a wide range of representative scenarios, conducted on a real prototyping platform using an AMD Alveo U55C FPGA board, demonstrate that FARMER consistently identifies a feasible solution while exploring less than$\mathbf {0.012\%}$of the total design space.
Gabriele Montanaro, Francesco Trovò, Davide Zoni
IEEE Trans. Very Large Scale Integr. Syst.1
2025 A Benchmarking Platform for DDR4 Memory Performance in Data-Center-Class FPGAs
abstract
FPGAs are increasingly utilized in data centers due to their ability to exploit parallelism in computationally intensive workloads. Modern workloads demand the transfer of vast amounts of information, making it essential to optimize communication between FPGAs and memory. This paper introduces a novel benchmarking platform for evaluating DDR4 memory performance in data-center-class FPGAs. The proposed solution features highly configurable traffic generation with complex memory access patterns defined at run time and can be flexibly instantiated on the target FPGA to support multiple memory channels and varying data rates. An extensive experimental campaign targets the AMD Kintex UltraScale 115 FPGA, encompassing up to three memory channels with data rates ranging from 1600 to 2400 MT/s. The results demonstrate the benchmaking platform’s capability to effectively evaluate DDR4 performance across various memory traffic configurations.
Andrea Galimberti, Gabriele Montanaro, Andrea Motta, Federico Proverbio, Davide Zoni
ISCAS2
2025 FARMER: An Online-Learning Driven Methodology for Workload Consolidation on Large FPGAs
abstract
With the ever-increasing demand for performance and scalability in cloud applications, high-performance computing (HPC) facilities are starting to include FPGAs for workload acceleration. To efficiently exploit the massive amount of resources of high-end FPGAs, it is paramount to optimize the allocation of multiple applications on a single device. This paper proposes FARMER, a novel online learning methodology harnessing the power of Gaussian Process regression to model the throughput of different applications running on the same FPGA, and a sequential decision-making approach to explore the available configurations efficiently. Experimental results considering a large variety of representative scenarios tested on a real prototyping platform featuring an AMD Virtex-7 FPGA show that FARMER always finds a feasible solution with an exploration of less than 0.1% of the whole design space.
Gabriele Montanaro, Francesco Trovò, Davide Zoni
ISCAS1
2024 A Prototype-Based Framework to Design Scalable Heterogeneous SoCs with Fine-Grained DFS
abstract
Frameworks for the agile development of modern system-on-chips are crucial to dealing with the complexity of de-signing such architectures. The open-source Vespa framework for designing large, FPGA-based, multi-core heterogeneous system-on-chips enables a faster and more flexible design space exploration of such architectures and their run-time optimization. Vespa, built on ESP, introduces the capabilities to instantiate multiple replicas of the same accelerator in a single network-on-chip node and to partition the system-on-chips into frequency islands with independent dynamic frequency scaling actuators, as well as a dedicated run-time monitoring infrastructure. Experiments on 4-by-4 tile-based system-on-chips demonstrate the possibility of effectively exploring a multitude of solutions that differ in the replication of accelerators, the clock frequencies of the frequency islands, and the tiles' placement, as well as monitoring a variety of statistics related to the traffic on the interconnect and the accelerators' performance at run time.
Gabriele Montanaro, Andrea Galimberti, Davide Zoni
ICCD1
2022 On the use of hardware accelerators in QC-MDPC code-based cryptography
abstract
Public-key cryptography (PKC) allows exchanging keys over an insecure channel without sharing a secret key. However, quantum computers threaten to break traditional PKC, thus, to mitigate such risk, post-quantum cryptography (PQC) aims to develop cryptosystems that are secure against attacks from quantum and classical computers. BIKE [1] is a key encapsulation mechanism (KEM) based on quasi-cyclic moderate-density parity-check (QC-MDPC) codes that is a candidate within the NIST standardization process to identify a set of PQC algorithms [4]. Figure 1 depicts the key exchange between two client and server nodes, which requires the sequential execution of the key generation, encapsulation, and decapsulation KEM primitives. Key generation and decapsulation are performed on the client side, while encapsulation is carried out by the server. Despite the vast literature targeting efficient hardware support for BIKE, each proposal delivered computing platforms meant either to maximize performance or minimize resource utilization.
Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni
CF3
2022 FPGA implementation of BIKE for quantum-resistant TLS
abstract
The recent advances in quantum computers impose the adoption of post-quantum cryptosystems into secure communication protocols. This work proposes two FPGA-based, client- and server-side hardware architectures to support the integration of the BIKE post-quantum KEM within TLS. Thanks to the parametric hardware design, the paper explores the best option between hardware and software implementations, given a set of available hardware resources and a realistic use-case scenario. The experimental evaluation comparing our client and server designs against the reference AVX2 and hardware implementations of BIKE highlighted two aspects. First, the proposed client and server architectures outperform the reference hardware implementation of BIKE by eight and four times, respectively. Second, the performance comparison between our client and server designs against the reference AVX2 implementation strongly depends on the available resource. Our solution is almost twice as fast as the AVX2 implementation while implemented on the Artix-7 200 FPGA, while it is up to six times slower when targeting smaller FPGAs, thus motivating a careful analysis of the available hardware resources and the optimization of the design's parallelism before opting for hardware support.
Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni
DSD3
2022 Efficient and Scalable FPGA Design of GF($2^m$2m) Inversion for Post-Quantum Cryptosystems
abstract
Post-quantum cryptosystems based on QC-MDPC codes are designed to mitigate the security threat posed by quantum computers to traditional public-key cryptography. The polynomial inversion is the core operation of key generation in such cryptosystems and the adoption of ephemeral keys imposes the execution of key generation for each session. To this end, there is a need for efficient and scalable hardware implementations of the binary polynomial inversion operation to support the key generation primitive across a wide range of computational platforms. This manuscript proposes an efficient and scalable architecture implementing the binary polynomial inversion at the hardware level. Our solution can deliver a performance-optimized implementation for the large polynomials used in post-quantum code-based cryptosystems and for each FPGA of the mid-range Xilinx Artix-7 family. The effectiveness of the proposed solution was validated by means of the BIKE and LEDAcrypt post-quantum QC-MDPC cryptosystems as representative use cases. Compared to the C11- and the optimized AVX2-based software implementations of LEDAcrypt, instances of the proposed architecture targeting the Artix-7 200 FPGA show an average performance improvement of 31.7 and 2.2 times, respectively. Moreover, the proposed architecture delivers a performance improvement up to 18.1 and 21.5 times for AES-128 and AES-192 security levels, respectively, compared to the BIKE hardware implementation.
Andrea Galimberti, Gabriele Montanaro, Davide Zoni
IEEE Trans. Computers2