Giuseppe Sorrentino

dblp:243/4736 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 7 first-author · 12 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring a Resource-Efficient NTT FPGA Accelerator for Fully Homomorphic Encryption
abstract
CKKS encryption scheme stands as one of the most valuable solutions for Fully Homomorphic Encryption (FHE), enabling privacy-preserving computation on encrypted data, at the cost of high computational bottlenecks. In such a scheme, the Number Theoretic Transform (NTT) consumes most of the computational resources due to irregular memory access patterns. Thus, literature accelerates this step on specific hardware devices, such as FPGAs, often exhausting device resources while gaining performance and energy efficiency improvements. However, this prevents further utilization of the FPGA to accelerate other compute-intensive stages. As an alternative, we perform the HW/SW co-design of resource-efficient solutions by integrating them into well-known software libraries implementing CKKS encryption scheme. In particular, we deploy on a Kria KV260 SoC a resource-efficient NTT accelerator with state-of-the-art security parameters (logN ∈ 12,…,16 and logQ ∈ [32,64]), and integrate it into the full-RNS HEANN library – the reference implementation for CKKS scheme. By doing so, we obtain up to 4.47× and 3.63× speedup in the encoding and encryption steps, respectively, while minimizing hardware consumption. These results show the end-to-end improvements achievable without fully utilizing the FPGA resources, leaving headroom for accelerating additional stages of the encryption pipeline.
Valentino Guerrini, Giuseppe Sorrentino, Davide Conficconi
DATE2
2026 RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAs
abstract
Modern datacenters integrate heterogeneous accelerators, such as GPUs and FPGAs, to speed up different stages of compute-intensive pipelines. GPUs are best suited for massively parallel workloads (e.g., deep learning), while FPGAs excel at task-level parallelism, stream-oriented processing, and in-network acceleration. Since these architectures must exchange data efficiently, literature introduced Peer-To-Peer (P2P) communication across PCI Express (PCIe) devices, to reduce CPU-driven orchestration and avoid intermediate, redundant buffer copies that degrade performance. However, current solutions are either closed-source or tied to proprietary frameworks, limiting P2P communication across most PCIe-based devices and requiring significant technical effort to enable P2P capabilities on supported hardware. For this reason, we propose RoPeerTo, a fully open-source, datacenter-scale architecture for P2P DMA communication over PCIe, validated on both GPUs and FPGAs. The goal is to provide a general, open alternative that ensures flexibility, efficiency, and usability. To this end, we design a complete HW/SW stack operating across different layers, supporting standard protocols for DMA-based memory sharing, advanced tools for device virtualization, memory address translation, and access protection. The result is a unified framework exposing a high-level API to end users, that enables direct communication between accelerators such as FPGAs and GPUs, and abstracts away the underlying hardware setup and management. We validate the system across different scenarios. First, we isolate the communication layer, observing a 5.61× speedup and a 37.99% reduction in GPU power consumption during data transfer. Next, we leverage the system for a compute-intensive workload where communication is only a partial bottleneck, achieving a 6.77% speedup without any compute-side modifications. Finally, we evaluate communication-heavy distributed computing workloads, demonstrating up to a 21.79× speedup in network-bound data scattering.
Marco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer, Lucian Petrica, Dario Korolija, Marco D. Santambrogio, Davide Conficconi, Gustavo Alonso, Kenneth O'Brien
EuroSys2
2026 ReFHE-NTT: Resource-Driven NTT FPGA Architecture for Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables privacy-preserving computation on encrypted data at a high computational cost. Among the existing schemes, CKKS is gaining traction thanks to its support for approximate arithmetic over real numbers. In such a scheme, the Number Theoretic Transform (NTT) is dominant, involving intensive modular arithmetic, twiddle-factor handling, and nontrivial memory access patterns. As a result, NTT acceleration has gained significant interest, with many FPGA designs achieving high throughput by aggressively exploiting device resources. While effective for NTT-centric workloads, this approach is ill-suited to full FHE pipelines, where the NTT must coexist with other compute-intensive kernels. Thus, we propose ReFHE-NTT, a resource-efficient NTT accelerator tailored to such settings. Our design combines on-the-fly twiddle-factor fusion with specialized modular arithmetic for pseudo-Mersenne primes, reducing memory footprint and arithmetic cost without storing precomputed tables. We co-design and validate the accelerator on the KV260 MPSoC, supporting polynomial degrees log N ∈ 12…16 and CKKS parameter sets with moduli up to 64 bits per prime. To the best of our knowledge, ReFHE-NTT is the first solution targeting embedded platform to support full-scale CKKS parameter sets. Compared to prior FPGA designs, it achieves up to a 20.2× improvement in slice-equivalent efficiency over the fastest open-source accelerator and a 1.98× improvement over the most resource-efficient one. Integrated into HEAAN CKKS library, ReFHE-NTT delivers top end-to-end speedup of 15× for encoding and 7.9× for encryption, demonstrating how resource-driven NTT design can substantially improve FHE performance on embedded platforms.
Valentino Guerrini, Giuseppe Sorrentino, Alessandro Barenghi, Davide Conficconi
FCCM2
2026 Unleashing Heterogeneous Systems Capabilities to Enhance Compute-Intensive Workloads
abstract
Heterogeneous systems are a key enabler for meeting the tight performance and energy-efficiency demands of modern artificial intelligence workloads. By integrating diverse compute units, such as scalar cores, vector and matrix processors, and Programmable Logic (PL), these systems can exploit multiple, complementary forms of parallelism. However, current design methodologies rarely exploit this heterogeneity effectively. In practice, developers face fragmented toolchains, limited abstractions, and a lack of systematic design guidelines, leading to suboptimal resource utilization, longer development time, and inefficient application mapping across heterogeneous components. To address these challenges, this research proposes a unified approach that combines analytical modelling and practical design methodologies for heterogeneous systems. The objective is to provide designers with structured guidelines to effectively leverage heterogeneity, with validation on modern AMD Versal platforms that combine PL with hardened VLIW processors, namely AI Engines. The proposed approach introduces a methodology paired with an open-source development framework, proving a reduction in development effort while enabling the systematic design of efficient heterogeneous accelerators.
Giuseppe Sorrentino, Davide Conficconi
FCCM1
2026 Adaptive AIE-PL Systems for Efficient End-to-End Pyramidal 3D Image Registration
abstract
Modern accelerators maximize throughput through aggressive specialization. However, in many real-world applications, workloads often vary at runtime, requiring multiple bitstreams to handle such changes. As a result, frequent reconfigurations introduce substantial overhead that can dominate end-to-end execution time. This issue is particularly evident in AIE–PL systems, where statically scheduled AI Engines (AIEs) achieve high performance through compile-time optimization and are therefore typically tailored to fixed workloads. Although AIEs support Runtime Parameters (RTPs) under Processing System (PS) orchestration, RTPs are impractical for discrete hosts. For this reason, we present a structured approach to designing single-bitstream, runtime-adaptable AIE–PL accelerators that does not rely on RTPs, suitable for discrete hosts. We exploit the Programmable Logic (PL) to generate and stream a compact metadata packet that distributes workload configuration across a directed AIE graph before computation. By doing so, we deliberately trade a fraction of fixed-instance efficiency for flexibility. We validate our approach by devising PeterPan, a software-programmable AIE–PL accelerator for 3D image registration. PeterPan supports runtime-varying problem sizes and integrates seamlessly into multi-stage pipelines, such as pyramidal (coarse-to-fine) registration. To maximize PeterPan utilization, we couple it with an ad-hoc software module that employs a novel heuristic to rapidly select informative sub-volumes, keeping the accelerator continuously fed and preventing input-side stalls. On a VCK5000, PeterPan matches state-of-the-art accelerator performance while retaining software programmability. In the end-to-end task, instead, PeterPan delivers a 3.06× speedup and a 2.74× higher energy efficiency than the state-of-the-art AIE-PL accelerator.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Claudio Di Salvo, Eleonora D'Arnese, Davide Conficconi
FCCM1
2025 Soaring with TRILLI: An HW/SW Heterogeneous Accelerator for Multi-Modal Image Registration
abstract
3D rigid image registration is a pivotal procedure in computer vision that aligns a floating volume with a reference one to correct positional and rotational distortions. It serves either as a stand-alone process or as a pre-processing step for non-rigid registration, where the rigid part dominates the computational cost. Various hardware accelerators have been proposed to optimize its compute-intensive components: geometric transformation with interpolation and similarity metric computation. However, existing solutions fail to address both components effectively, as GPUs excel at image transformation, while FPGAs in similarity metric computation. To close this gap, we propose TRILLI, a novel Versal-based accelerator for image transformation and interpolation. TRILLI optimally maps each computational step on the proper heterogeneous hardware component. TRILLI achieves speedup of 5.32× against the top performing GPU-based solution, and an energy efficiency improvement of 36.75 × against the most efficient one. Moreover, we integrate it with an FPGA-based similarity metric from literature to complete a rigid image registration step (i.e., transformation, interpolation, and similarity metric) attaining a speedup of 18.60 × against the top performing GPU-based solution, while being 36.11 ×more efficient than the most energy efficient one.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
FCCM1
2025 Accelerating K-Means: A Vectorized Approach for AI Engines & Neural Processing Units
abstract
K-Means is a clustering technique widely employed in AI workloads, from image processing to data mining. Given its importance, researchers propose different algorithms and hardware-accelerated implementations. While algorithm suitability can depend on the target use case, there is much less doubt about the architecture: FPGAs are the de facto standard, as the design can be perfectly tailored to the target use case. Despite this, AI accelerators such as GPUs and Neural Processing Units (NPUs) are gaining traction. The former attains remarkable performance at the cost of low energy efficiency. The latter, instead, promises to maximize both, but they are strongly underutilized due to the lack of a clear approach for K-Means acceleration. Considering AMD NPU, for example, the main computing cores are AI Engines that require algorithm reshaping and code optimization to harness data parallelism effectively. Thus, this research analyzes different K-Means versions to propose a vectorized algorithm that fully uses AI Engine (AIE) features. We validate our vectorized K-Means on Versal VCK5000, using FPGAs for data movement only, as the Memory Transfer Engines and Shim Tiles of NPUs, and the AI Engine for computation. This design reflects features of modern NPUs, making the validation fair. We attain up to$59.5 \times$speedup against Torch library on GPUs while being comparable but more energy efficient than further optimized GPU solutions.
Eleonora Cabai, Giuseppe Sorrentino, Marco D. Santambrogio, Davide Conficconi
FPL2
2025 Towards Accelerated Healthcare Federated System Through Heterogeneous Accelerators
abstract
Federated Learning (FL) enables collaborative model training across multiple hosts without exposing sensitive data. Yet, privacy-preserving is critical, introducing a nonnegligible overhead. Furthermore, while GPUs perfectly suit AI workloads, FPGAs outperform them in networking and cryptographic tasks as Fully Homomorphic Encryption (FHE), widely employed in FL to safeguard privacy. In this context, applications like Image Registration (IR), requiring different computeintensive steps, each suitable for a different architecture, even worsen the hardware dichotomy. Thus, this research explores innovative federated network structures to reduce privacy risks while assigning roles to each host, guaranteeing complete hardware acceleration per each compute-intensive task by partitioning each step on different hardware layers of modern heterogeneous platforms as Neural Processing Unitss (NPUs), combining AI Engines (AIEs) and GPUs, Versal VCK5000, combining FPGA and AIEs, and multi-platform peer-to-peer systems.
Giuseppe Sorrentino, Davide Conficconi
FPL1
2025 VOTED - Versal Optimization Toolkit for Education and Heterogeneous Systems Development
abstract
Despite classic educational approaches proving their effectiveness for classic hardware acceleration, they suffer novel heterogeneous systems such as Versal, requiring a deeper system-level awareness. Versal devices merge FPGA on-field programma-bility with the performance of hardened VLIW processors, namely AI Engine, at the cost of facing novel challenges when integrating these two different layers. Given the system complexity of such devices and the low-level knowledge required to leverage them, Versal potentialities have yet to be fully exploited. Therefore, we present VOTED, a Versal Optimization Toolkit for Education and Heterogeneous Systems Development, that guides users of any expertise in discovering, using, and optimizing Versal-based applications. VOTED proved to be effective, leading five different groups of students, at their first experience with heterogeneous system design, at devising quite complex Versal-based applications capable of reaching the final stages of a design competition and even winning it with four months of work.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
ISCAS1
2024 Co-Designing a 3D Transformation Accelerator for Versal-Based Image Registration
abstract
Rigid image registration is pivotal in modern imaging for correcting distortions of images acquired with different modalities or at different time instants. Literature accelerates its main compute-intensive steps, image transformation and similarity metric computation, through GPUs or FPGAs to meet performance-efficiency constraints. However, GPUs lack energy efficiency while FPGAs lack performance for image transformation. Therefore, we adopt a single heterogeneous system through PEGASO, a methodology and its implementation to co-design the image transformation algorithm on Versal system. We maximize performance by combining custom data layout and hardware optimizations, attaining a 19x speedup over the best GPU-based transformation accelerator. When integrated with FPGA-based similarity metric, PEGASO achieves 82.73x and 20.73x speedup against FPGA- and GPU-based solutions while improving the corresponding energy efficiency of 318.18x and 52.62x.
Paolo Salvatore Galfano, Giuseppe Sorrentino, Eleonora D'Arnese, Davide Conficconi
ICCD2
2023 Hephaestus: Codesigning and Automating 3D Image Registration on Reconfigurable Architectures
abstract
Healthcare is a pivotal research field, and medical imaging is crucial in many applications. Therefore finding new architectural and algorithmic solutions would benefit highly repetitive image processing procedures. One of the most complex tasks in this sense is image registration, which finds the optimal geometric alignment among 3D image stacks and is widely employed in healthcare and robotics. Given the high computational demand of such a procedure, hardware accelerators are promising real-time and energy-efficient solutions, but they are complex to design and integrate within software pipelines. Therefore, this work presents an automation framework called Hephaestus that generates efficient 3D image registration pipelines combined with reconfigurable accelerators. Moreover, to alleviate the burden from the software, we codesign software-programmable accelerators that can adapt at run-time to the image volume dimensions. Hephaestus features a cross-platform abstraction layer that enables transparently high-performance and embedded systems deployment. However, given the computational complexity of 3D image registration, the embedded devices become a relevant and complex setting being constrained in memory; thus, they require further attention and tailoring of the accelerators and registration application to reach satisfactory results. Therefore, with Hephaestus , we also propose an approximation mechanism that enables such devices to perform the 3D image registration and even achieve, in some cases, the accuracy of the high-performance ones. Overall, Hephaestus demonstrates 1.85× of maximum speedup, 2.35× of efficiency improvement with respect to the State of the Art, a maximum speedup of 2.51× and 2.76× efficiency improvements against our software, while attaining state-of-the-art accuracy on 3D registrations.
Giuseppe Sorrentino, Marco Venere, Davide Conficconi, Eleonora D'Arnese, Marco D. Santambrogio
ACM Trans. Embed. Comput. Syst.1
2021 Understanding the Kelvin pin mitigation of the MOSFET turn-on losses by fast-switching and neutralization of the clamp diode
abstract
The super-junction MOSFETs enhanced with an additional source, called Kelvin pin (4-lead), are faster than traditional three-terminal MOSFETs (3-lead). Moreover, the former presents a lower increment in the turn-on power losses as the load increases. This paper investigates the phenomenon and explains this behavior. The drain-source voltage waveform of the 3-lead device usually presents a plateau due to the voltage drop across the parasitic inductances during the current rising and to the clamp imposed by the diode until it conducts. The higher current speed of 4-lead involves that the plateau should occur at a lower value of the drain-source voltage, then this voltage takes longer to reach this lockout value. On the other hand, the higher current speed involves a reduced time necessary for the 4-lead device to reach the boost inductor current, thus unclamping earlier the diode. Consequently, the plateau does not occur thus achieving reduced turn-on losses. The plateau extension is small in the 3-lead MOSFET at low current and it increases as the current increases, thus the advantage increases as the current increases. The formulation and argumentation supporting the understanding of this phenomenon can be applied also to faster power devices like the SiC MOSFET and the GaN HEMT.
Francesco Giorgio, Santi Agatino Rizzo, Nunzio Salerno, Giuseppe Scarcella, Alfio Scuto, Giuseppe Sorrentino
IECON6
2013 GaN HEMT devices: Experimental results on normally-on, normally-off and cascode configuration
abstract
Power electronics systems play key function in power management and motion control: power consumption and volume reduction are strongly required in new applications oriented to protect environment. Moreover, a switching frequency increase is demanded for microwave and power switching applications, thus reducing passive component and converter volume: however, increasing the switching frequency directly increases switching losses. Also, switching applications are very demanding, because semiconductor switches are required to withstand high voltage in reverse condition and handle large current in forward operation mode. New material and new devices are studied to satisfy such requirements, like SiC and GaN devices. In this work, several 600-V class GaN-on-Si HEMT prototypes are presented. These devices have been designed and developed for power switching converter applications. In order to offer a complete scenario of GaN-on-Si HEMT technology, normally-on, normally-off, and cascode-connected devices have been characterized. Achieved experimental results of static and pulsed measurements are then shown.
Giuseppe Sorrentino, Maurizio Melito, Alfonso Patti, Giovanni Parrino, Angelo Raciti
IECON1