EDBT 2026 Demo / reviewers in the wild / expert
Davide Zoni
dblp:53/10927
· DBLP profile ↗
38ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0002-9951-062XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 15 first-author · 22 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Omega: A Hardware-Software Framework for Complete Design Space Exploration of FPGA-Based Heterogeneous Multi-Core SoCsabstractThe design space exploration (DSE) of heterogeneous multi-core systems-on-chip (SoCs) presents a massive challenge due to the vast and complex configuration space, demanding simultaneous optimization of performance, energy efficiency, and resource utilization under diverse constraints. Traditional approaches leveraging analytical models, heuristics, and machine learning (ML) techniques fail to comprehensively cover this space, often yielding suboptimal solutions. The Omega framework, designed for exhaustive DSE of FPGA-based heterogeneous multi-core SoCs, addresses the former limitations by fully exploring the design space and guaranteeing the identification of globally optimal configurations. Omega leverages FPGAs’ dynamic partial reconfiguration to accelerate the DSE drastically and can serve as a golden model for evaluating novel DSE heuristics and ML methods. This manuscript demonstrates the proposed framework’s effectiveness through an extensive experimental campaign on 16-core SoCs with accelerators for up to five different applications, achieving a substantial speedup, of 29 times on average, compared to traditional techniques while ensuring solution optimality. Omega is released as a comprehensive open-source ecosystem, compatible with commercially available FPGA platforms, to facilitate future research and practical adoption, setting a new benchmark for DSE methodologies and providing a robust tool for optimizing next-generation computing platforms. Gabriele Montanaro, Andrea Galimberti, Davide Zoni |
IEEE Trans. Computers | 3 |
| 2026 | FARMER: Online-Learning-Based Workload Consolidation on Large FPGAs Accelerated With Dynamic Partial ReconfigurationabstractAs the demand for performance and scalability in cloud applications continues to grow, high-performance computing (HPC) facilities increasingly integrate FPGAs to accelerate computational workloads. To fully utilize the extensive resources available on modern high-end FPGAs, it is essential to optimize the allocation of multiple applications on a single device. This article introduces FARMER, a novel online learning methodology that leverages machine learning (ML) to model the throughput of different applications running concurrently on the same FPGA. It combines this with a sequential decision-making strategy and an in-circuit exploration flow based on dynamic partial reconfiguration (DPR) to drastically speed up the exploration of large design spaces. Experimental evaluations across a wide range of representative scenarios, conducted on a real prototyping platform using an AMD Alveo U55C FPGA board, demonstrate that FARMER consistently identifies a feasible solution while exploring less than$\mathbf {0.012\%}$of the total design space. Gabriele Montanaro, Francesco Trovò, Davide Zoni |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and ReliabilityabstractModern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners. Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo |
DSD | 24 |
| 2025 | Rhea: a Framework for Fast Design and Validation of RTL Cache-Coherent Memory SubsystemsabstractDesigning and validating efficient cache-coherent memory subsystems is a critical yet complex task in the development of modern multi-core system-on-chip architectures. Rhea is a unified framework that streamlines the design and system-level validation of RTL cache-coherent memory subsystems. On the design side, Rhea generates synthesizable, highly configurable RTL supporting various architectural parameters. On the validation side, Rhea integrates Verilator’s cycle-accurate RTL simulation with gem5’s full-system simulation, allowing realistic workloads and operating systems to run alongside the actual RTL under test. We apply Rhea to design MSI-based RTL memory subsystems with one and two levels of private caches and scaling up to sixteen cores. Their evaluation with 22 applications from state-of-the-art benchmark suites shows intermediate performance relative to gem5 Ruby’s MI and MOESI models. The hybrid gem5-Verilator co-simulation flow incurs a moderate simulation overhead, up to 2.7 times compared to gem5 MI, but achieves higher fidelity by simulating real RTL hardware. This overhead decreases with scale, down to 1.6 times in sixteen-core scenarios. These results demonstrate Rhea’s effectiveness and scalability in enabling fast development of RTL cache-coherent memory subsystem designs. Davide Zoni, Andrea Galimberti, Adriano Guarisco |
ICCAD | 1 |
| 2025 | A Benchmarking Platform for DDR4 Memory Performance in Data-Center-Class FPGAsabstractFPGAs are increasingly utilized in data centers due to their ability to exploit parallelism in computationally intensive workloads. Modern workloads demand the transfer of vast amounts of information, making it essential to optimize communication between FPGAs and memory. This paper introduces a novel benchmarking platform for evaluating DDR4 memory performance in data-center-class FPGAs. The proposed solution features highly configurable traffic generation with complex memory access patterns defined at run time and can be flexibly instantiated on the target FPGA to support multiple memory channels and varying data rates. An extensive experimental campaign targets the AMD Kintex UltraScale 115 FPGA, encompassing up to three memory channels with data rates ranging from 1600 to 2400 MT/s. The results demonstrate the benchmaking platform’s capability to effectively evaluate DDR4 performance across various memory traffic configurations. Andrea Galimberti, Gabriele Montanaro, Andrea Motta, Federico Proverbio, Davide Zoni |
ISCAS | 5 |
| 2025 | Rabbit: Dynamic Clock Randomization to Protect against Side-Channel AttacksabstractThe continuous evolution of side-channel analysis motivates a continuous investigation to deliver novel countermeasures. This work presents a hiding countermeasure leveraging a randomized Dynamic Frequency Scaling (DFS) actuator built on top of the clocking resources available in modern FPGAs. In contrast to state-of-the-art DFS-based solutions, our approach is meant to optimize security and performance metrics with a modest increase in power consumption. We experimentally validated our countermeasure on real hardware by comparing it against recently proposed hiding methods employing clock desynchronization. To strengthen our security assessment, we also considered a large variety of state-of-the-art side-channel attacks, including recent deep-learning ones. The experimental results confirm that none of the evaluated attack techniques can breach our protected target, and TLVA shows no information leakage with 10 million traces. The performance overhead is zero, while the power overhead is limited to 1.55×. Davide Galli, Matteo Matteucci, Davide Zoni |
ISCAS | 3 |
| 2025 | FARMER: An Online-Learning Driven Methodology for Workload Consolidation on Large FPGAsabstractWith the ever-increasing demand for performance and scalability in cloud applications, high-performance computing (HPC) facilities are starting to include FPGAs for workload acceleration. To efficiently exploit the massive amount of resources of high-end FPGAs, it is paramount to optimize the allocation of multiple applications on a single device. This paper proposes FARMER, a novel online learning methodology harnessing the power of Gaussian Process regression to model the throughput of different applications running on the same FPGA, and a sequential decision-making approach to explore the available configurations efficiently. Experimental results considering a large variety of representative scenarios tested on a real prototyping platform featuring an AMD Virtex-7 FPGA show that FARMER always finds a feasible solution with an exploration of less than 0.1% of the whole design space. Gabriele Montanaro, Francesco Trovò, Davide Zoni |
ISCAS | 3 |
| 2025 | A Deep Learning-Assisted Template Attack Against Dynamic Frequency Scaling CountermeasuresabstractIn the last decades, machine learning techniques have been extensively used in place of classical template attacks to implement profiled side-channel analysis. This manuscript focuses on the application of machine learning to counteract Dynamic Frequency Scaling defenses. While state-of-the-art attacks have shown promising results against desynchronization countermeasures, a robust attack strategy has yet to be realized. Motivated by the simplicity and effectiveness of template attacks for devices lacking desynchronization countermeasures, this work presents a Deep Learning-assisted Template Attack (DLaTA) methodology specifically designed to target highly desynchronized traces through Dynamic Frequency Scaling. A deep learning-based pre-processing step recovers information obscured by desynchronization, followed by a template attack for key extraction. Specifically, we developed a three-stage deep learning pipeline to resynchronize traces to a uniform reference clock frequency. The experimental results on the AES cryptosystem executed on a RISC-V System-on-Chip reported a Guessing Entropy equal to 1 and a Guessing Distance greater than 0.25. Results demonstrate the method's ability to successfully retrieve secret keys even in the presence of high desynchronization. As an additional contribution, we publicly release ourDFS_DESYNCHdatabase11https://github.com/hardware-fab/DLaTAcontaining the first set of real-world highly desynchronized power traces from the execution of a software AES cryptosystem. Davide Galli, Francesco Lattari, Matteo Matteucci, Davide Zoni |
IEEE Trans. Computers | 4 |
| 2025 | An FPGA-Based Open-Source Hardware-Software Framework for Side-Channel Security ResearchabstractAttacks based on side-channel analysis (SCA) pose a severe security threat to modern computing platforms, further exacerbated on IoT devices by their pervasiveness and handling of private and critical data. Designing SCA-resistant computing platforms requires a significant additional effort in the early stages of the IoT devices’ life cycle, which is severely constrained by strict time-to-market deadlines and tight budgets. This manuscript introduces a hardware-software framework meant for SCA research on FPGA targets. It delivers an IoT-class system-on-chip (SoC) that includes a RISC-V CPU, provides observability and controllability through an ad-hoc debug infrastructure to facilitate SCA attacks and evaluate the platform's security, and streamlines the deployment of SCA countermeasures through dedicated hardware and software features such as a DFS actuator and FreeRTOS support. The open-source release of the framework includes the SoC, the scripts to configure the computing platform, compile a target application, and assess the SCA security, as well as a suite of state-of-the-art attacks and countermeasures. The goal is to foster its adoption and novel developments in the field, empowering designers and researchers to focus on studying SCA countermeasures and attacks while relying on a sound and stable hardware-software platform as the foundation for their research. Davide Zoni, Andrea Galimberti, Davide Galli |
IEEE Trans. Computers | 1 |
| 2024 | A Deep- Learning Technique to Locate Cryptographic Operations in Side-Channel TracesabstractSide-channel attacks allow extracting secret infor-mation from the execution of cryptographic primitives by cor-relating the partially known computed data and the measured side-channel signal. However, to set up a successful side-channel attack, the attacker has to perform i) the challenging task of locating the time instant in which the target cryptographic primitive is executed inside a side-channel trace and then ii) the time-alignment of the measured data on that time instant. This paper presents a novel deep-learning technique to locate the time instant in which the target computed cryptographic operations are executed in the side-channel trace. In contrast to state-of-the-art solutions, the proposed methodology works even in the presence of trace deformations obtained through random delay insertion techniques. We validated our proposal through a successful attack against a variety of unprotected and protected cryptographic primitives that have been executed on an FPGA-implemented system-on-chip featuring a RISC- V CPU. Giuseppe Chiari, Davide Galli, Francesco Lattari, Matteo Matteucci, Davide Zoni |
DATE | 5 |
| 2024 | The TEXTAROSSA Project: Cool all the Way Down to the HardwareabstractThe TEXTAROSSA project aims to bridge the technology gaps that exascale computing systems will face in the near future in order to overcome their performance and energy efficiency challenges. This project provides solutions for improved energy efficiency and thermal control, seamless integration of heterogeneous accelerators in HPC multi-node platforms, and new arithmetic methods. Challenges are tacked through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models, and tools derived from European research. Antonio Filgueras, Giovanni Agosta, Marco Aldinucci, Carlos Álvarez 0001, Pasqua D'Ambra, Massimo Bernaschi, Andrea Biagioni, Daniele Cattaneo 0002, Alessandro Celestini, Massimo Celino, Carlotta Chiarini, Francesca Lo Cicero, Paolo Cretaro, William Fornaciari, Ottorino Frezza, Andrea Galimberti, Francesco Giacomini, Juan Miguel De Haro Ruiz, Francesco Iannone, Daniel Jaschke, Daniel Jiménez-González, Michal Kulczewski, Alberto Leva, Alessandro Lonardo, Michele Martinelli, Xavier Martorell, Simone Montangero, Lucas Morais, Ariel Oleksiak, Paolo Palazzari, Luca Pontisso, Federico Reghenzani, Cristian Rossi, Sergio Saponara, Carlo Saverio Lodi, Francesco Simula, Federico Terraneo, Piero Vicini, Miquel Vidal, Davide Zoni, Giuseppe Zummo |
DSD | 40 |
| 2024 | Rethinking the Switch Architecture for Stateful In-network ComputingabstractProgrammable switches are a disruptive technology that has seen increasing adoption in the past decade. Since their inception, however, there has been tension regarding how to design these switches. Classic programmable switches operate at line rate but impose significant limitations on the expressiveness of their programming models. In contrast, alternative designs relax the strict line rate requirement but are more easily programmable. The common belief is that a switch's performance and its programmability are at odds. Alberto Lerner, Davide Zoni, Paolo Costa, Gianni Antichi |
HotNets | 2 |
| 2024 | Hound: Locating Cryptographic Primitives in Desynchronized Side-Channel Traces using Deep-LearningabstractSide-channel attacks allow the extraction of sensitive information from cryptographic primitives by correlating the partially known computed data and the measured side-channel signal. Starting from the raw side-channel trace, the preprocessing of the side-channel trace to pinpoint the time at which each cryptographic primitive is executed, and, then, to re-align all the collected data to this specific time represent a critical step to setup a successful side-channel attack. The use of hiding techniques has been widely adopted as a low-cost solution to hinder the preprocessing of side-channel traces, thus limiting side-channel attacks in real scenarios. This work introduces Hound, a novel deep-learning-based pipeline to locate the execution of cryptographic primitives within the side-channel trace even in the presence of trace deformations introduced by the use of dynamic frequency scaling actuators. Hound has been validated through successful attacks on various cryptographic primitives executed on an FPGA-based system-on-chip incorporating a RISC- V CPU while dynamic frequency scaling is active. Experimental results demonstrate the possibility of identifying the cryptographic primitives in DFS-deformed side-channel traces. Davide Galli, Giuseppe Chiari, Davide Zoni |
ICCD | 3 |
| 2024 | A Prototype-Based Framework to Design Scalable Heterogeneous SoCs with Fine-Grained DFSabstractFrameworks for the agile development of modern system-on-chips are crucial to dealing with the complexity of de-signing such architectures. The open-source Vespa framework for designing large, FPGA-based, multi-core heterogeneous system-on-chips enables a faster and more flexible design space exploration of such architectures and their run-time optimization. Vespa, built on ESP, introduces the capabilities to instantiate multiple replicas of the same accelerator in a single network-on-chip node and to partition the system-on-chips into frequency islands with independent dynamic frequency scaling actuators, as well as a dedicated run-time monitoring infrastructure. Experiments on 4-by-4 tile-based system-on-chips demonstrate the possibility of effectively exploring a multitude of solutions that differ in the replication of accelerators, the clock frequencies of the frequency islands, and the tiles' placement, as well as monitoring a variety of statistics related to the traffic on the interconnect and the accelerators' performance at run time. Gabriele Montanaro, Andrea Galimberti, Davide Zoni |
ICCD | 3 |
| 2024 | Design-time methodology for optimizing mixed-precision CPU architectures on FPGAabstractApproximate computing can significantly reduce the energy consumption of computing systems. Mixed-precision hardware architectures and precision-tuning tools for software provide the ability to introduce approximations, but when applied separately, they do not give complete control over the accuracy-energy trade-off. The co-optimization of approximations in hardware and software is a complex task, but it promises considerable benefits. We present a methodology for the fast design-time selection of mixed-precision hardware-software combinations that minimize the energy consumption and the area of the target FPGA-based softcore CPUs with configurable support for floating-point and fixed-point arithmetic. Our approach can evaluate configurations more than 2000 times faster than the alternative approach of using gate-level simulation. On benchmarks from the PolyBench suite the identified hardware-software configurations showed improvement of the energy-to-solution metric ranging from 20% to 95%. Lev Denisov, Andrea Galimberti, Daniele Cattaneo 0002, Giovanni Agosta, Davide Zoni |
J. Syst. Archit. | 5 |
| 2023 | Hardware and Software Support for Mixed Precision Computing: a Roadmap for Embedded and HPC SystemsabstractMixed precision is an approximate computing technique that can be used to trade-off computation accuracy for performance and/or energy. It can be applied to many error-tolerant applications, but manual precision tuning is both tedious and error-prone. Furthermore, the effectiveness of the technique heavily depends on hardware characteristics. Therefore, a hardware/software co-design approach is necessary for an effective exploitation of precision tuning opportunities offered by the applications. In this paper, we propose, based on the state of the art of precision tuning software and mixed precision hardware, a roadmap for the evolution of hardware designs and compiler-based precision tuning support, which is ongoing in the context of the European projects TEXTAROSSA and APROPOS. William Fornaciari, Giovanni Agosta, Daniele Cattaneo 0002, Lev Denisov, Andrea Galimberti, Gabriele Magnani, Davide Zoni |
DATE | 7 |
| 2022 | On the use of hardware accelerators in QC-MDPC code-based cryptographyabstractPublic-key cryptography (PKC) allows exchanging keys over an insecure channel without sharing a secret key. However, quantum computers threaten to break traditional PKC, thus, to mitigate such risk, post-quantum cryptography (PQC) aims to develop cryptosystems that are secure against attacks from quantum and classical computers. BIKE [1] is a key encapsulation mechanism (KEM) based on quasi-cyclic moderate-density parity-check (QC-MDPC) codes that is a candidate within the NIST standardization process to identify a set of PQC algorithms [4]. Figure 1 depicts the key exchange between two client and server nodes, which requires the sequential execution of the key generation, encapsulation, and decapsulation KEM primitives. Key generation and decapsulation are performed on the client side, while encapsulation is carried out by the server. Despite the vast literature targeting efficient hardware support for BIKE, each proposal delivered computing platforms meant either to maximize performance or minimize resource utilization. Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni |
CF | 5 |
| 2022 | FPGA implementation of BIKE for quantum-resistant TLSabstractThe recent advances in quantum computers impose the adoption of post-quantum cryptosystems into secure communication protocols. This work proposes two FPGA-based, client- and server-side hardware architectures to support the integration of the BIKE post-quantum KEM within TLS. Thanks to the parametric hardware design, the paper explores the best option between hardware and software implementations, given a set of available hardware resources and a realistic use-case scenario. The experimental evaluation comparing our client and server designs against the reference AVX2 and hardware implementations of BIKE highlighted two aspects. First, the proposed client and server architectures outperform the reference hardware implementation of BIKE by eight and four times, respectively. Second, the performance comparison between our client and server designs against the reference AVX2 implementation strongly depends on the available resource. Our solution is almost twice as fast as the AVX2 implementation while implemented on the Artix-7 200 FPGA, while it is up to six times slower when targeting smaller FPGAs, thus motivating a careful analysis of the available hardware resources and the optimization of the design's parallelism before opting for hardware support. Andrea Galimberti, Davide Galli, Gabriele Montanaro, William Fornaciari, Davide Zoni |
DSD | 5 |
| 2022 | Gated-CNN: Combating NBTI and HCI aging effects in on-chip activation memories of Convolutional Neural Network acceleratorsabstractNegative Bias Temperature Instability (NBTI) and Hot Carrier Injection (HCI) are two of the main reliability threats in current technology nodes. These aging phenomena degrade the transistor’s threshold voltage (Vth) over the lifetime of a digital circuit, resulting in slower transistors that eventually lead to a faulty operation when the critical paths become longer than the processor cycle time. Among all the transistors on a chip, the most vulnerable transistors to such wearout effects are those used to implement SRAM storage, since memory cells are continuously degrading. In particular, NBTI ages PMOS cell transistors when a given logic value is stored for a long period (i.e., a long duty cycle), whereas HCI ages NMOS cell transistors not only when the stored value flips but also when it is accessed. This work focuses on mitigating aging in the on-chip SRAM memories of Convolutional Neural Network (CNN) accelerators storing activations. This paper makes two main contributions. At the software level, we quantify the aging induced by current CNN benchmarks with a characterization study of duty cycle, flip, and access patterns in every activation memory cell. Based on the insights from this study, this work proposes a novel microarchitectural technique, Gated-CNN, that ensures a uniform aging degradation of every memory cell. To do so, Gated-CNN exploits power-gating and address rotation techniques tailored to the memory demands and temporal/spatial localities exhibited by CNN applications, as well as the memory organization and management of CNN accelerators. Experimental results show that, compared to a conventional design, the average Vth degradation savings are at least as much as 49% depending on the type of transistor. Nicolás Landeros Muñoz, Alejandro Valero, Ruben Gran Tejero, Davide Zoni |
J. Syst. Archit. | 4 |
| 2022 | Cost-effective fixed-point hardware support for RISC-V embedded systems
Davide Zoni, Andrea Galimberti |
J. Syst. Archit. | 1 |
| 2022 | Efficient and Scalable FPGA Design of GF($2^m$2m) Inversion for Post-Quantum CryptosystemsabstractPost-quantum cryptosystems based on QC-MDPC codes are designed to mitigate the security threat posed by quantum computers to traditional public-key cryptography. The polynomial inversion is the core operation of key generation in such cryptosystems and the adoption of ephemeral keys imposes the execution of key generation for each session. To this end, there is a need for efficient and scalable hardware implementations of the binary polynomial inversion operation to support the key generation primitive across a wide range of computational platforms. This manuscript proposes an efficient and scalable architecture implementing the binary polynomial inversion at the hardware level. Our solution can deliver a performance-optimized implementation for the large polynomials used in post-quantum code-based cryptosystems and for each FPGA of the mid-range Xilinx Artix-7 family. The effectiveness of the proposed solution was validated by means of the BIKE and LEDAcrypt post-quantum QC-MDPC cryptosystems as representative use cases. Compared to the C11- and the optimized AVX2-based software implementations of LEDAcrypt, instances of the proposed architecture targeting the Artix-7 200 FPGA show an average performance improvement of 31.7 and 2.2 times, respectively. Moreover, the proposed architecture delivers a performance improvement up to 18.1 and 21.5 times for AES-128 and AES-192 security levels, respectively, compared to the BIKE hardware implementation. Andrea Galimberti, Gabriele Montanaro, Davide Zoni |
IEEE Trans. Computers | 3 |
| 2022 | Design of Side-Channel-Resistant Power MonitorsabstractIn modern computing platforms, power monitors (PwrMons) are employed to deliver online power estimates to support different runtime power-performance optimization methodologies. However, the possibility of setting up a successful side-channel attack by analyzing the power estimates imposes the use of a suitable and systematic approach in the design of such PwrMons. This article proposes a design methodology to automatically identify and implement side-channel-resistant PwrMons at the hardware level, for generic computing platforms. The methodology works by designing a PwrMon for which the switching activity of the signals used to compute the power estimates is not a function of both the secret key and the plaintext/ciphertext values processed by the computing platform. According to the most recent standardized methodologies to assess the side-channel security, our experimental validation leverages both correlation power analysis and$t$-test analysis considering a general purpose System on Chip executing different cryptographic primitives and an application-specific accelerator implementing the AES-128 algorithm. Our results confirm the impossibility of retrieving the secret key from the power estimates provided by our side-channel-resistant PwrMon. Considering several temporal resolutions, we highlight an accuracy error of the power estimates limited to less than 2.7%, as well as an average area and power overheads for the protected PwrMons lower than 6% and 5%, respectively. To this end, the proposed methodology is able to deliver a side-channel-resistant PwrMon within state-of-the-art accuracy error and overheads. Davide Zoni, Luca Cremona, William Fornaciari |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascaleabstractTo achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research. Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola |
DSD | 8 |
| 2020 | All-Digital Control-Theoretic Scheme to Optimize Energy Budget and Allocation in Multi-CoresabstractThe Internet-of-Things (IoT) revolution fueled new challenges and opportunities to achieve computational efficiency goals. Embedded devices are required to execute multiple applications for which a suitable distribution of the computing power must be adapted at run-time. Such complex hardware platforms have to sustain the continuous acquisition and processing of data under severe energy budget constraints, since most of them are battery powered. The state-of-the-art offers several ad-hoc contributions to selectively optimize the performance considering aspects like energy, power, thermal, or reliability. However, there is a need for a generic coordinated management strategy able to cope with all of these dimensions, while allowing the Operating System (OS) and the applications to “suggest” or constrain the actuation. This article proposes a unified control-theoretic scheme to coordinate the design of energy-budget and energy allocation solutions for multi-cores. The proposed controller can work with any actuator and it can interact, at run-time, with both the applications and the OS to optimize the actuation signals steering the computing platform. Such control scheme offers the possibility to integrate any performance related policy in the form of an energy-allocation strategy, still ensuring the theoretic exponential stability of the overall controller if the actuation of the policy, coming from the OS and the applications, “is not too fast.” To demonstrate the feasibility of our solution, we have implemented the controller into a RISC multi-core running on the Xilinx Artix 100t FPGA device, available in the the Digilent Nexys4-DDR board. Results considering two actuators and both the quadand the eight-core version of the considered computing platform, highlight the scalability of the proposed solution as well as an area overhead for the -all digital, on chip-controller limited to 0.86 percent (FFs) and 5.3 percent (LUTs) of the FPGA chip. We also considered a dynamic scenario validating the speed of the controller, where our framework has to face with modifications to the energy-allocation control policy carried out by the OS and the applications. The obtained results are collected by executing a huge mix of benchmarks and the statistical significance is accounted by executing each scenario 30 times. Such results are analyzed considering three quality metrics. First, the efficiency in exploiting the imposed budget (EFF9) that is on average 98.27 percent. Second, the overflow of the actual average power consumption with respect to the assigned budget (OνF9), which is limited to 1.43 mW on average. Last, the performance utility loss due to the control scheme that is limited to 1.87 percent on average. Davide Zoni, Luca Cremona, William Fornaciari |
IEEE Trans. Computers | 1 |
| 2020 | Scramble Suit: A Profile Differentiation Countermeasure to Prevent Template AttacksabstractEnsuring protection against side channel attacks (SCAs) is a crucial requirement in the design of modern secure embedded systems. Profiled SCAs, the class to which template attacks and machine learning attacks belong, derive a model of the side channel behavior of a device identical to the target one, and exploit the said model to extract the key from the target, under the hypothesis that the side channel behaviors of the two devices match. We propose an architectural countermeasure against cross-device profiled attacks which differentiates the side channel behavior of different instances of the same hardware design, preventing the reuse of a model derived on a device other than the target one. In particular, we describe an instance of our solution providing a protected hardware implementation of the advanced encryption standard (AES) block cipher and experimentally validate its resistance against both Bayesian templates and machine learning approaches based on support vector machines also considering different state-of-the-art feature reduction techniques to increase the effectiveness of the profiled attacks. Results show that our countermeasure foils the key retrieval attempts via profiled attacks ensuring a key derivation accuracy equivalent to a random guess. Alessandro Barenghi, William Fornaciari, Gerardo Pelosi, Davide Zoni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Partial Packet Forwarding to Improve Performance in Fully Adaptive Routing for Cache-Coherent NoCsabstractIn the contest of cache-coherent Networks-on-Chip (NoCs), fully adaptive routing algorithms guarantee maximum flexibility to implement power-performance, fault tolerant, thermal and Quality of Service (QoS) management policies. However, to get rid of deadlock at both protocol and network level, their implementation imposes a relevant resource increase. Moreover, their performance are inferior to the one of deterministic and partially adaptive schemes mainly due to the additional constraints imposed to the virtual channel (VC) re-use policy. This work proposes a novel flow control scheme to improve the performance of fully adaptive routing algorithms by allowing an aggressive reuse of VCs in presence of both long and short packets. Our proposal works by splitting long packets in multiple chunks and by reallocating the VCs to the chunks rather that to the entire packet. By carefully sizing each chunk to fit the available space in the reallocated, eventually not empty, VC, we are avoiding deadlocks while increasing the NoC utilization and performance. Experimental results show that our solution offers a 23.8% increase, on average, in the saturation point when compared to the best state of the art flow control scheme for fully adaptive routing algorithms. Moreover, our flow control scheme offers similar or better performance than the XY routing algorithm with the same number of resources, and we also ensure superior flexibility in the definition of the routing function. Tamer Eltaras, William Fornaciari, Davide Zoni |
PDP | 3 |
| 2018 | PowerProbe: Run-time power modeling through automatic RTL instrumentationabstractOnline power monitoring represents a de-facto solution to enable energyand power-aware run-time optimizations for current and future computing architectures. Traditionally, the performance counters of the target architecture are used to feed a software-based, power model that is continuously updated to deliver the required run-time power estimates. The solution introduces a non-negligible performance and energy overhead. Moreover, itis limited to the availability of such performance counters that, however, are not primarily intended for online power monitoring. This paper introduces PowerProbe, a run-time power monitoring methodology that automatically extracts and implements a power model from the RTL description of the target architecture. The solution does not leverage any performance counter to ensure wide applicability. Moreover, the use of ad-hoc hardware that continuously updates the power estimate minimizes both the performance and the power overheads. We employ a fully compliant OpenRisc 1000 implementation to validate PowerProbe. The results highlight an average prediction error within 9% (standard deviation less than 2%), with a power and area overheads limited to 6.89% and 4.71%, respectively. Davide Zoni, Luca Cremona, William Fornaciari |
DATE | 1 |
| 2018 | DarkCache: Energy-Performance Optimization of Tiled Multi-Cores by Adaptively Power-Gating LLC BanksabstractThe Last Level Cache (LLC) is a key element to improve application performance in multi-cores. To handle the worst case, the main design trend employs tiled architectures with a large LLC organized in banks, which goes underutilized in several realistic scenarios. Our proposal, named DarkCache , aims at properly powering off such unused banks to optimize the Energy-Delay Product (EDP) through an adaptive cache reconfiguration, thus aggressively reducing the leakage energy. The implemented solution is general and it can recognize and skip the activation of the DarkCache policy for the few strong memory intensive applications that actually require the use of the entire LLC. The validation has been carried out on 16- and 64-core architectures also accounting for two state-of-the-art methodologies. Compared to the baseline solution, DarkCache exhibits a performance overhead within 2% and an average EDP improvement of 32.58% and 36.41% considering 16 and 64 cores, respectively. Moreover, DarkCache shows an average EDP gain between 16.15% (16 cores) and 21.05% (64 cores) compared to the best state-of-the-art we evaluated, and it confirms a good scalability since the gain improves with the size of the architecture. Davide Zoni, William Fornaciari |
ACM Trans. Archit. Code Optim. | 1 |
| 2018 | A Comprehensive Side-Channel Information Leakage Analysis of an In-Order RISC CPU MicroarchitectureabstractSide-channel attacks are a prominent threat to the security of embedded systems. To perform them, an adversary evaluates the goodness of fit of a set of key-dependent power consumption models to a collection of side-channel measurements taken from an actual device, identifying the secret key value as the one yielding the best-fitting model. In this work, we analyze for the first time the microarchitectural components of a 32-bit in-order RISC CPU, showing which one of them is accountable for unexpected side-channel information leakage. We classify the leakage sources, identifying the data serialization points in the microarchitecture and providing a set of hints that can be fruitfully exploited to generate implementations resistant against side-channel attacks, either writing or generating proper assembly code. Davide Zoni, Alessandro Barenghi, Gerardo Pelosi, William Fornaciari |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2017 | MANGO: Exploring Manycore Architectures for Next-GeneratiOn HPC SystemsabstractThe Horizon 2020 MANGO project aims at exploring deeply heterogeneous accelerators for use in High-Performance Computing systems running multiple applications with different Quality of Service (QoS) levels. The main goal of the project is to exploit customization to adapt computing resources to reach the desired QoS. For this purpose, it explores different but interrelated mechanisms across the architecture and system software. In particular, in this paper we focus on the runtime resource management, the thermal management, and support provided for parallel programming, as well as introducing three applications on which the project foreground will be validated. José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Etienne Cappe, Alessandro Cilardo, Leon Dragic, Alexandre Dray, Alen Duspara, William Fornaciari, Gerald Guillaume, Ynse Hoornenborg, Arman Iranfar, Mario Kovac, Simone Libutti, Bruno Maitre, José Maria Martínez, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Tomás Picornell, Igor Piljic, Anna Pupykina, Federico Reghenzani, Isabelle Staub, Rafael Tornero, Marina Zapater, Davide Zoni |
DSD | 29 |
| 2017 | BlackOut: Enabling fine-grained power gating of buffers in Network-on-Chip routers
Davide Zoni, Andrea Canidio, William Fornaciari, Panayiotis Englezakis, Chrysostomos Nicopoulos, Yiannakis Sazeides |
J. Parallel Distributed Comput. | 1 |
| 2016 | Enabling HPC for QoS-sensitive applications: The MANGO approach
José Flich, Giovanni Agosta, Philipp Ampletzer, David Atienza 0001, Carlo Brandolese, Alessandro Cilardo, William Fornaciari, Ynse Hoornenborg, Mario Kovac, Bruno Maitre, Giuseppe Massari, Hrvoje Mlinaric, Ermis Papastefanakis, Fabrice Roudet, Rafael Tornero, Davide Zoni |
DATE | 16 |
| 2016 | CUTBUF: Buffer Management and Router Design for Traffic Mixing in VNET-Based NoCsabstractRouter's buffer design and management strongly influence energy, area and performance of on-chip networks, hence it is crucial to encompass all of these aspects in the design process. At the same time, the NoC design cannot disregard preventing network-level and protocol-level deadlocks by devoting ad-hoc buffer resources to that purpose. In chip multiprocessor systems the coherence protocol usually requires different virtual networks (VNETs) to avoid deadlocks. Moreover, VNET utilization is highly unbalanced and there is no way to share buffers between them due to the need to isolate different traffic types. This paper proposes CUTBUF, a novel NoC router architecture to dynamically assign virtual channels (VCs) to VNETs depending on the actual VNETs load to significantly reduce the number of physical buffers in routers, thus saving area and power without decreasing NoC performance. Moreover, CUTBUF allows to reuse the same buffer for different traffic types while ensuring that the optimized NoC is deadlock-free both at network and protocol level. In this perspective, all the VCs are considered spare queues not statically assigned to a specific VNET and the coherence protocol only imposes a minimum number of queues to be implemented. Synthetic applications as well as real benchmarks have been used to validate CUTBUF, considering architectures ranging from 16 up to 48 cores. Moreover, a complete RTL router has been designed to explore area and power overheads. Results highlight how CUTBUF can reduce router buffers up to 33 percent with 2 percent of performance degradation, a 5 percent of operating frequency decrease and area and power saving up to 30.6 and 30.7 percent, respectively. Conversely, the flexibility of the proposed architecture improves by 23.8 percent the performance of the baseline NoC router when the same number of buffers is used. Davide Zoni, José Flich, William Fornaciari |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | TEST: Assessing NoC Policies Facing Aging and Leakage PowerabstractThe trend to increase the number of cores integrated on a single die makes Networks-on-Chip (NoCs) a key component from the interconnection viewpoint. Unfortunately, continuous scaling of CMOS technology poses severe concerns regarding failure mechanisms, such as NBTI, that are crucial in achieving a reasonable component lifetime. Furthermore, the leakage power became more and more a critical issues as the technology scales up. Finally, Process Variation (PV) makes harder the scenario, decreasing device lifetime and performance predictability during chip fabrication. Several techniques have been presented in literature facing the NBTI and or the static power consumption. This paper proposes a methodology to analyze such techniques from the feasibility viewpoint. It is explored their effectiveness in contrasting NBTI and saving static power in the NoC as well as the associated overheads and drawbacks. For the two considered policies, it is achieved a NBTI mitigation up to 55% and a power saving up to 51% with performance and area overheads less than 10% and 5%, respectively. Davide Zoni, Luca Borghese, Giuseppe Massari, Simone Libutti, William Fornaciari |
DSD | 1 |
| 2015 | Modeling DVFS and Power-Gating Actuators for Cycle-Accurate NoC-Based SimulatorsabstractNetworks-on-chip (NoCs) are a widely recognized viable interconnection paradigm to support the multi-core revolution. One of the major design issues of multicore architectures is still the power, which can no longer be considered mainly due to the cores, since the NoC contribution to the overall energy budget is relevant. To face both static and dynamic power while balancing NoC performance, different actuators have been exploited in literature, mainly dynamic voltage frequency scaling (DVFS) and power gating. Typically, simulation-based tools are employed to explore the huge design space by adopting simplified models of the components. As a consequence, the majority of state-of-the-art on NoC power-performance optimization do not accurately consider timing and power overheads of actuators, or (even worse) do not consider them at all, with the risk of overestimating the benefits of the proposed methodologies. This article presents a simulation framework for power-performance analysis of multicore architectures with specific focus on the NoC. It integrates accurate power gating and DVFS models encompassing also their timing and power overheads. The value added of our proposal is manyfold: (i) DVFS and power gating actuators are modeled starting from SPICE-level simulations; (ii) such models have been integrated in the simulation environment; (iii) policy analysis support is plugged into the framework to enable assessment of different policies; (iv) a flexible GALS ( globally asynchronous locally synchronous ) support is provided, covering both handshake and FIFO re-synchronization schemas. To demonstrate both the flexibility and extensibility of our proposal, two simple policies exploiting the modeled actuators are discussed in the article. Davide Zoni, William Fornaciari |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2015 | A control-based methodology for power-performance optimization in NoCs exploiting DVFS
Davide Zoni, Federico Terraneo, William Fornaciari |
J. Syst. Archit. | 1 |
| 2013 | Sensor-wise methodology to face NBTI stress of NoC buffersabstractNetworks-on-Chip (NoCs) are a key component for the new many-core architectures, from the performance and reliability stand-points. Unfortunately, continuous scaling of CMOS technology poses severe concerns regarding failure mechanisms such as NBTI and stress-migration. Process variation makes harder the scenario, decreasing device lifetime and performance predictability during chip fabrication. This paper presents a novel cooperative sensor-wise methodology to reduce the NBTI degradation in the network on-chip (NoC) virtual channel (VC) buffers, considering process variation effects as well. The changes introduced to the reference NoC model exhibit an area overhead below 4%. Experimental validation is obtained using a cycle accurate simulator considering both real and synthetic traffic patterns. We compare our methodology to the best sensor-less round-robin approach used as reference model. The proposed sensor-wise strategy achieves up to 26.6% and 18.9% activity factor improvement over the reference policy on synthetic and real traffic patterns respectively. Moreover a net NBTI Vthsaving up to 54.2% is shown against the baseline NoC that does not account for NBTI. Davide Zoni, William Fornaciari |
DATE | 1 |
| 2012 | HANDS: heterogeneous architectures and networks-on-chip design and simulationabstractIn current multi-core scenario, Networks-on-Chip (NoC) represent a suitable choice to face the increasing communication and performance requirements, however introducing additional design challenges to already complex architectures. In this perspective, there is a need for flexible and configurable virtual platforms for early-stage design exploration. We present the Heterogeneous Architectures and Networks-on-Chip Design and Simulation framework for large-scale high-performance computer simulation, integrating performance, power, thermal and reliability metrics under a unique methodology. Moreover, NoC exploration is possible from a reliability/performance and thermal/performance trade-offs. Davide Zoni, Simone Corbetta, William Fornaciari |
ISLPED | 1 |