Stefano Di Matteo

dblp:257/5265 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-5711-432XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Evaluating Communication and Architectural Overheads in NTT Accelerators for ML-KEM
abstract
The standardization of post-quantum cryptography (PQC) has led to the adoption of ML-KEM, a lattice-based key encapsulation mechanism derived from CRYSTALS-Kyber. ML-KEM heavily relies on polynomial multiplication over the Module Learning With Errors (MLWE) problem, efficiently implemented using the Number Theoretic Transform (NTT). While the NTT significantly reduces computational complexity, its irregular memory access patterns and high arithmetic intensity pose challenges for efficient software execution, especially in constrained environments. Consequently, hardware acceleration has become essential for enabling practical ML-KEM deployment in embedded and edge systems. Recent research has proposed several hardware accelerators for the NTT, but often without demonstrating their effectiveness in accelerating full matrix-vector products within ML-KEM. Designing such accelerators requires careful architectural trade-offs involving memory organization, data movement, parallelism, and system integration. The main contribution of this paper is a comprehensive assessment of communication, memory arrangement, and architectural overheads in NTT accelerators. We analyze state-of-the-art NTT acceleration techniques for ML-KEM and we propose a hardware architecture optimized for efficient matrixvector multiplication. Our NTT-based accelerator is integrated into a RISC-V System-on-Chip (SoC), and we evaluate the impact of key architectural choices (e.g., parallelism on the butterfly units, latency of communication and memory arrangement) on area and performance.
Stefano Di Matteo, Emanuele Valea
DDECS1
2026 KEM-22: An Efficient Post-Quantum ML-KEM Hardware Accelerator on 22-nm ASIC
abstract
This paper presents a performant and compact hardware accelerator for the ML-KEM algorithm, compliant with the NIST FIPS 203 specification and implemented in a 22nm ASIC technology. The proposed design supports all three standardized security levels (ML-KEM-512, -768, and -1024) using a unified, parameter-agnostic architecture that avoids logic duplication. The design relies exclusively on SRAM blocks for storage, completely eliminating FIFOs and intermediate buffers, while flip-flops are used only in timing-critical paths to achieve high-frequency operation with minimal area overhead. Among all known ASIC implementations, our architecture achieves the best normalized area-time product across all ML-KEM parameter sets, demonstrating efficiency and scalability, while also delivering the lowest power and energy per operation among state-of-the-art solutions. These results make the proposed design a strong candidate for real-world post-quantum cryptographic deployments on constrained hardware and IoT platforms.
Stefano Di Matteo, Ivan Sarno, Emanuele Valea, Sergio Saponara
IEEE Internet Things J.1
2025 TYRCA: A RISC-V Tightly-Coupled Accelerator for Code-Based Cryptography
abstract
Post-quantum cryptography (PQC) has garnered significant attention across various communities, particularly with the National Institute of Standards and Technology (NIST) advancing to the fourth round of PQC standardization. One of the leading candidates is Hamming Quasi-Cyclic (HQC), which received a significant update on February 23, 2024. This update, which introduces a classical dense-dense multiplication approach, has no known dedicated hardware implementations yet. The innovative Core-V eXtension InterFace (CV-X-IF) is a communication interface for RISC-V processors that significantly facilitates the integration of new instructions to the Instruction Set Architecture (ISA), through tightly connected accelerators. In this paper, we present a TightlY-coupled accelerator for RISC-V for Code-based cryptogrAphy (TYRCA), proposing the first fully tightly-coupled hardware implementation of the HQC-PQC algorithm, leveraging the CV-X-IF. The proposed architecture is implemented on the Xilinx Kintex-7 FPGA. Experimental results demonstrate that TYRCA reduces the execution time by 94% to 96% for HQC-128, HQC-192, and HQC-256, showcasing its potential for efficient HQC code-based cryptography.
Alessandra Dolmeta, Stefano Di Matteo, Emanuele Valea, Mikael Carmona, Antoine Loiseau, Maurizio Martina, Guido Masera
DATE2
2025 Hardware Design of an Advanced-Feature Cryptographic Tile Within the European Processor Initiative
abstract
This work describes the hardware implementation of a cryptographic accelerators suite, named Crypto-Tile, in the framework of the European Processor Initiative (EPI) project. The EPI project traced the roadmap to develop the first family of low-power processors with the design fully made in Europe, for Big Data, supercomputers and automotive. Each of the coprocessors of Crypto-Tile is dedicated to a specific family of cryptographic algorithms, offering functions for symmetric and public-key cryptography, computation of digests, generation of random numbers, and Post-Quantum cryptography. The performances of each coprocessor outperform other available solutions, offering innovative hardware-native services, such as key management, clock randomisation and access privilege mechanisms. The system has been synthesised on a 7 nm standard-cell technology, being the first Cryptoprocessor to be characterised in such an advanced silicon technology. The post-synthesis netlist has been employed to assess the resistance of Crypto-Tile to power analysis side-channel attacks. Finally, a demoboard has been implemented, integrating a RISC-V softcore processor and the Crypto-Tile module, and drivers for hardware abstraction layer, bare-metal applications and drivers for Linux kernel in C language have been developed. Finally, we exploited them to compare in terms of execution speed the hardware-accelerated algorithms against software-only solutions.
Pietro Nannipieri, Luca Crocetti, Stefano Di Matteo, Luca Fanucci, Sergio Saponara
IEEE Trans. Computers3
2022 VLSI Design of Advanced-Features AES Cryptoprocessor in the Framework of the European Processor Initiative
abstract
This article presents a cryptographic hardware (HW) accelerator supporting multiple advanced encryption standard (AES)-based block cipher modes, including the more advanced cipher-based MAC (CMAC), counter with CBC-MAC (CCM), Galois counter mode (GCM), and XOR-encrypt-XOR-based tweaked-codebook mode with ciphertext stealing (XTS) modes. The proposed design implements advanced and innovative features in HW, such as AES key secure management, on-chip clock randomization, and access privilege mechanisms. The system has been tested in a RISC-V-based system-on-chip (SoC), specifically designed for this purpose, on an Ultrascale + Xilinx FPGA, analyzing resource and power consumption, together with system performances. The cryptoprocessor has been then synthesized on a 7-nm CMOS standard-cells technology; performances, complexity, and power consumption information are analyzed and compared with the state of the art. The proposed cryptoprocessor is ready to be embedded within the innovative European Processor Initiative (EPI) chip.
Pietro Nannipieri, Stefano Di Matteo, Luca Baldanzi, Luca Crocetti, Luca Zulberti, Sergio Saponara, Luca Fanucci
IEEE Trans. Very Large Scale Integr. Syst.2