Paolo Salvatore Galfano

dblp:395/1140 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0009-7783-5421ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Adaptive AIE-PL Systems for Efficient End-to-End Pyramidal 3D Image Registration
abstract
Modern accelerators maximize throughput through aggressive specialization. However, in many real-world applications, workloads often vary at runtime, requiring multiple bitstreams to handle such changes. As a result, frequent reconfigurations introduce substantial overhead that can dominate end-to-end execution time. This issue is particularly evident in AIE–PL systems, where statically scheduled AI Engines (AIEs) achieve high performance through compile-time optimization and are therefore typically tailored to fixed workloads. Although AIEs support Runtime Parameters (RTPs) under Processing System (PS) orchestration, RTPs are impractical for discrete hosts. For this reason, we present a structured approach to designing single-bitstream, runtime-adaptable AIE–PL accelerators that does not rely on RTPs, suitable for discrete hosts. We exploit the Programmable Logic (PL) to generate and stream a compact metadata packet that distributes workload configuration across a directed AIE graph before computation. By doing so, we deliberately trade a fraction of fixed-instance efficiency for flexibility. We validate our approach by devising PeterPan, a software-programmable AIE–PL accelerator for 3D image registration. PeterPan supports runtime-varying problem sizes and integrates seamlessly into multi-stage pipelines, such as pyramidal (coarse-to-fine) registration. To maximize PeterPan utilization, we couple it with an ad-hoc software module that employs a novel heuristic to rapidly select informative sub-volumes, keeping the accelerator continuously fed and preventing input-side stalls. On a VCK5000, PeterPan matches state-of-the-art accelerator performance while retaining software programmability. In the end-to-end task, instead, PeterPan delivers a 3.06× speedup and a 2.74× higher energy efficiency than the state-of-the-art AIE-PL accelerator.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Claudio Di Salvo, Eleonora D'Arnese, Davide Conficconi
FCCM2
2026 Tempranillo: Non-Speculative Early Register Release
abstract
Limited by the breakdown of technology scaling, CPU architects are looking for creative solutions to deliver performance improvements while unable to traditionally scale up microarchitectural structures. One promising approach is engineering a microarchitecture that uses resources more efficiently, for example, by recycling them faster to relieve pressure on critical structures and making them look bigger than they are. The physical register file (PRF) is a key structure that faces severe area and power constraints that limit its (and the whole CPU's) scalability. Based on this observation, previous work proposed solutions to reduce the pressure on the PRF by reducing the time each register remains allocated. In this paper, we corroborate earlier findings that the early release of registers is a promising approach to reduce the occupancy of the PRF and we identify novel tight conditions for safely releasing registers. Based on this analysis, we design Tempranillo: an aggressive, non-speculative microarchitecture to release registers as early as possible without requiring additional recovery mechanisms. Tempranillo delivers up to 3.3 % and 11.8 % performance improvement over conventional release on singlethreaded and a 2 -way SMT CPUs, respectively. Additionally, Tempranillo requires modest storage overheads, translating into a performance improvement per KiB of storage of up to 2.6 % and 9.3 % for single-thread and 2-way SMT, respectively. Our evaluation shows that Tempranillo improves over both the state-of-the-art non-speculative and speculative proposals.
Carlos Escuin, Paolo Salvatore Galfano, Davide B. Bartolini, Leeor Peled, Mehdi Alipour
HPCA2
2025 Soaring with TRILLI: An HW/SW Heterogeneous Accelerator for Multi-Modal Image Registration
abstract
3D rigid image registration is a pivotal procedure in computer vision that aligns a floating volume with a reference one to correct positional and rotational distortions. It serves either as a stand-alone process or as a pre-processing step for non-rigid registration, where the rigid part dominates the computational cost. Various hardware accelerators have been proposed to optimize its compute-intensive components: geometric transformation with interpolation and similarity metric computation. However, existing solutions fail to address both components effectively, as GPUs excel at image transformation, while FPGAs in similarity metric computation. To close this gap, we propose TRILLI, a novel Versal-based accelerator for image transformation and interpolation. TRILLI optimally maps each computational step on the proper heterogeneous hardware component. TRILLI achieves speedup of 5.32× against the top performing GPU-based solution, and an energy efficiency improvement of 36.75 × against the most efficient one. Moreover, we integrate it with an FPGA-based similarity metric from literature to complete a rigid image registration step (i.e., transformation, interpolation, and similarity metric) attaining a speedup of 18.60 × against the top performing GPU-based solution, while being 36.11 ×more efficient than the most energy efficient one.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
FCCM2
2025 VOTED - Versal Optimization Toolkit for Education and Heterogeneous Systems Development
abstract
Despite classic educational approaches proving their effectiveness for classic hardware acceleration, they suffer novel heterogeneous systems such as Versal, requiring a deeper system-level awareness. Versal devices merge FPGA on-field programma-bility with the performance of hardened VLIW processors, namely AI Engine, at the cost of facing novel challenges when integrating these two different layers. Given the system complexity of such devices and the low-level knowledge required to leverage them, Versal potentialities have yet to be fully exploited. Therefore, we present VOTED, a Versal Optimization Toolkit for Education and Heterogeneous Systems Development, that guides users of any expertise in discovering, using, and optimizing Versal-based applications. VOTED proved to be effective, leading five different groups of students, at their first experience with heterogeneous system design, at devising quite complex Versal-based applications capable of reaching the final stages of a design competition and even winning it with four months of work.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
ISCAS2
2024 Co-Designing a 3D Transformation Accelerator for Versal-Based Image Registration
abstract
Rigid image registration is pivotal in modern imaging for correcting distortions of images acquired with different modalities or at different time instants. Literature accelerates its main compute-intensive steps, image transformation and similarity metric computation, through GPUs or FPGAs to meet performance-efficiency constraints. However, GPUs lack energy efficiency while FPGAs lack performance for image transformation. Therefore, we adopt a single heterogeneous system through PEGASO, a methodology and its implementation to co-design the image transformation algorithm on Versal system. We maximize performance by combining custom data layout and hardware optimizations, attaining a 19x speedup over the best GPU-based transformation accelerator. When integrated with FPGA-based similarity metric, PEGASO achieves 82.73x and 20.73x speedup against FPGA- and GPU-based solutions while improving the corresponding energy efficiency of 318.18x and 52.62x.
Paolo Salvatore Galfano, Giuseppe Sorrentino, Eleonora D'Arnese, Davide Conficconi
ICCD1