Davide Conficconi

dblp:224/1533 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0002-5834-0812ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 1 first-author · 27 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring a Resource-Efficient NTT FPGA Accelerator for Fully Homomorphic Encryption
abstract
CKKS encryption scheme stands as one of the most valuable solutions for Fully Homomorphic Encryption (FHE), enabling privacy-preserving computation on encrypted data, at the cost of high computational bottlenecks. In such a scheme, the Number Theoretic Transform (NTT) consumes most of the computational resources due to irregular memory access patterns. Thus, literature accelerates this step on specific hardware devices, such as FPGAs, often exhausting device resources while gaining performance and energy efficiency improvements. However, this prevents further utilization of the FPGA to accelerate other compute-intensive stages. As an alternative, we perform the HW/SW co-design of resource-efficient solutions by integrating them into well-known software libraries implementing CKKS encryption scheme. In particular, we deploy on a Kria KV260 SoC a resource-efficient NTT accelerator with state-of-the-art security parameters (logN ∈ 12,…,16 and logQ ∈ [32,64]), and integrate it into the full-RNS HEANN library – the reference implementation for CKKS scheme. By doing so, we obtain up to 4.47× and 3.63× speedup in the encoding and encryption steps, respectively, while minimizing hardware consumption. These results show the end-to-end improvements achievable without fully utilizing the FPGA resources, leaving headroom for accelerating additional stages of the encryption pipeline.
Valentino Guerrini, Giuseppe Sorrentino, Davide Conficconi
DATE3
2026 RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAs
abstract
Modern datacenters integrate heterogeneous accelerators, such as GPUs and FPGAs, to speed up different stages of compute-intensive pipelines. GPUs are best suited for massively parallel workloads (e.g., deep learning), while FPGAs excel at task-level parallelism, stream-oriented processing, and in-network acceleration. Since these architectures must exchange data efficiently, literature introduced Peer-To-Peer (P2P) communication across PCI Express (PCIe) devices, to reduce CPU-driven orchestration and avoid intermediate, redundant buffer copies that degrade performance. However, current solutions are either closed-source or tied to proprietary frameworks, limiting P2P communication across most PCIe-based devices and requiring significant technical effort to enable P2P capabilities on supported hardware. For this reason, we propose RoPeerTo, a fully open-source, datacenter-scale architecture for P2P DMA communication over PCIe, validated on both GPUs and FPGAs. The goal is to provide a general, open alternative that ensures flexibility, efficiency, and usability. To this end, we design a complete HW/SW stack operating across different layers, supporting standard protocols for DMA-based memory sharing, advanced tools for device virtualization, memory address translation, and access protection. The result is a unified framework exposing a high-level API to end users, that enables direct communication between accelerators such as FPGAs and GPUs, and abstracts away the underlying hardware setup and management. We validate the system across different scenarios. First, we isolate the communication layer, observing a 5.61× speedup and a 37.99% reduction in GPU power consumption during data transfer. Next, we leverage the system for a compute-intensive workload where communication is only a partial bottleneck, achieving a 6.77% speedup without any compute-side modifications. Finally, we evaluate communication-heavy distributed computing workloads, demonstrating up to a 21.79× speedup in network-bound data scattering.
Marco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer, Lucian Petrica, Dario Korolija, Marco D. Santambrogio, Davide Conficconi, Gustavo Alonso, Kenneth O'Brien
EuroSys8
2026 ReFHE-NTT: Resource-Driven NTT FPGA Architecture for Fully Homomorphic Encryption
abstract
Fully Homomorphic Encryption (FHE) enables privacy-preserving computation on encrypted data at a high computational cost. Among the existing schemes, CKKS is gaining traction thanks to its support for approximate arithmetic over real numbers. In such a scheme, the Number Theoretic Transform (NTT) is dominant, involving intensive modular arithmetic, twiddle-factor handling, and nontrivial memory access patterns. As a result, NTT acceleration has gained significant interest, with many FPGA designs achieving high throughput by aggressively exploiting device resources. While effective for NTT-centric workloads, this approach is ill-suited to full FHE pipelines, where the NTT must coexist with other compute-intensive kernels. Thus, we propose ReFHE-NTT, a resource-efficient NTT accelerator tailored to such settings. Our design combines on-the-fly twiddle-factor fusion with specialized modular arithmetic for pseudo-Mersenne primes, reducing memory footprint and arithmetic cost without storing precomputed tables. We co-design and validate the accelerator on the KV260 MPSoC, supporting polynomial degrees log N ∈ 12…16 and CKKS parameter sets with moduli up to 64 bits per prime. To the best of our knowledge, ReFHE-NTT is the first solution targeting embedded platform to support full-scale CKKS parameter sets. Compared to prior FPGA designs, it achieves up to a 20.2× improvement in slice-equivalent efficiency over the fastest open-source accelerator and a 1.98× improvement over the most resource-efficient one. Integrated into HEAAN CKKS library, ReFHE-NTT delivers top end-to-end speedup of 15× for encoding and 7.9× for encryption, demonstrating how resource-driven NTT design can substantially improve FHE performance on embedded platforms.
Valentino Guerrini, Giuseppe Sorrentino, Alessandro Barenghi, Davide Conficconi
FCCM4
2026 To Infinity and Beyond: A Reconfigurable Architecture for Space Vision Computing
abstract
On-board satellite vision requires executing workloads such as Earth Observation (EO) and visual Guidance, Navigation, and Control (GNC) within a power budget that dynamically varies across mission phases. While these two workloads share a core of computational kernels, it is important to note that no single hardware configuration is optimal across all mission phases. To address this gap, we present three main contributions. First, we introduce STAR-Bench, an open-source framework that characterizes vision workloads across platforms and scenarios. Next, we define the Quality-Adjusted Cost (QAC), a composite metric to model the multidimensional latency-accuracy-energy Pareto frontier. Finally, UCHIHA is introduced as a domain-specific Coarse-Grained Reconfigurable Array (CGRA) whose Processing Element (PE) composition is derived from STAR-Bench profiles and whose spatial configuration adapts at runtime, driven by QAC-based Pareto analysis. In preliminary work, we applied our proposed contribution to on-board Image Registration (IR), revealing that no single algorithm dominates the Pareto frontier across scenarios. Furthermore, an adaptive orchestrator built on these insights achieved up to a 12.03× improvement over the best static configuration.
Claudio Di Salvo, Davide Conficconi
FCCM2
2026 Unleashing Heterogeneous Systems Capabilities to Enhance Compute-Intensive Workloads
abstract
Heterogeneous systems are a key enabler for meeting the tight performance and energy-efficiency demands of modern artificial intelligence workloads. By integrating diverse compute units, such as scalar cores, vector and matrix processors, and Programmable Logic (PL), these systems can exploit multiple, complementary forms of parallelism. However, current design methodologies rarely exploit this heterogeneity effectively. In practice, developers face fragmented toolchains, limited abstractions, and a lack of systematic design guidelines, leading to suboptimal resource utilization, longer development time, and inefficient application mapping across heterogeneous components. To address these challenges, this research proposes a unified approach that combines analytical modelling and practical design methodologies for heterogeneous systems. The objective is to provide designers with structured guidelines to effectively leverage heterogeneity, with validation on modern AMD Versal platforms that combine PL with hardened VLIW processors, namely AI Engines. The proposed approach introduces a methodology paired with an open-source development framework, proving a reduction in development effort while enabling the systematic design of efficient heterogeneous accelerators.
Giuseppe Sorrentino, Davide Conficconi
FCCM2
2026 Adaptive AIE-PL Systems for Efficient End-to-End Pyramidal 3D Image Registration
abstract
Modern accelerators maximize throughput through aggressive specialization. However, in many real-world applications, workloads often vary at runtime, requiring multiple bitstreams to handle such changes. As a result, frequent reconfigurations introduce substantial overhead that can dominate end-to-end execution time. This issue is particularly evident in AIE–PL systems, where statically scheduled AI Engines (AIEs) achieve high performance through compile-time optimization and are therefore typically tailored to fixed workloads. Although AIEs support Runtime Parameters (RTPs) under Processing System (PS) orchestration, RTPs are impractical for discrete hosts. For this reason, we present a structured approach to designing single-bitstream, runtime-adaptable AIE–PL accelerators that does not rely on RTPs, suitable for discrete hosts. We exploit the Programmable Logic (PL) to generate and stream a compact metadata packet that distributes workload configuration across a directed AIE graph before computation. By doing so, we deliberately trade a fraction of fixed-instance efficiency for flexibility. We validate our approach by devising PeterPan, a software-programmable AIE–PL accelerator for 3D image registration. PeterPan supports runtime-varying problem sizes and integrates seamlessly into multi-stage pipelines, such as pyramidal (coarse-to-fine) registration. To maximize PeterPan utilization, we couple it with an ad-hoc software module that employs a novel heuristic to rapidly select informative sub-volumes, keeping the accelerator continuously fed and preventing input-side stalls. On a VCK5000, PeterPan matches state-of-the-art accelerator performance while retaining software programmability. In the end-to-end task, instead, PeterPan delivers a 3.06× speedup and a 2.74× higher energy efficiency than the state-of-the-art AIE-PL accelerator.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Claudio Di Salvo, Eleonora D'Arnese, Davide Conficconi
FCCM5
2025 Combining MLIR Dialects with Domain-Specific Architecture for Efficient Regular Expression Matching
abstract
Pattern matching based on Regular Expressions (REs) is a pervasive and challenging computational kernel used in several applications to identify critical information in a data stream. Due to the sequential data dependency of REs and the increasing data volume growth, hardware acceleration is gaining attention to address the limitation of general-purpose architectures. RE-oriented Domain-Specific Architectures (DSAs) combine the flexibility of translating REs into binary code with the efficiency of a specialized architecture, filling the gap between frozen hardware accelerators and the versatility of CPUs/GPUs. However, existing DSAs focus mainly on the efficiency execution challenge while missing the optimization opportunities that a structured compilation infrastructure can provide. This paper proposes a RE-tailored multi-level intermediate representation strategy embodied by the MLIR framework at the compiler level to exploit different abstraction optimizations via two domain-specific dialects, one targeting the abstract representation of REs and the other targeting the underlying domain-specific ISA. Moreover, this paper proposes a novel architectural organization of an open-source state-of-the-art DSA to maximize the parallelization capabilities. Overall, the proposed approach significantly improves execution time by up to 2.26×, energy efficiency by up to 2.30×, and resource usage.
Andrea Somaini, Filippo Carloni, Giovanni Agosta, Marco D. Santambrogio, Davide Conficconi
CGO5
2025 Soaring with TRILLI: An HW/SW Heterogeneous Accelerator for Multi-Modal Image Registration
abstract
3D rigid image registration is a pivotal procedure in computer vision that aligns a floating volume with a reference one to correct positional and rotational distortions. It serves either as a stand-alone process or as a pre-processing step for non-rigid registration, where the rigid part dominates the computational cost. Various hardware accelerators have been proposed to optimize its compute-intensive components: geometric transformation with interpolation and similarity metric computation. However, existing solutions fail to address both components effectively, as GPUs excel at image transformation, while FPGAs in similarity metric computation. To close this gap, we propose TRILLI, a novel Versal-based accelerator for image transformation and interpolation. TRILLI optimally maps each computational step on the proper heterogeneous hardware component. TRILLI achieves speedup of 5.32× against the top performing GPU-based solution, and an energy efficiency improvement of 36.75 × against the most efficient one. Moreover, we integrate it with an FPGA-based similarity metric from literature to complete a rigid image registration step (i.e., transformation, interpolation, and similarity metric) attaining a speedup of 18.60 × against the top performing GPU-based solution, while being 36.11 ×more efficient than the most energy efficient one.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
FCCM4
2025 Accelerating K-Means: A Vectorized Approach for AI Engines & Neural Processing Units
abstract
K-Means is a clustering technique widely employed in AI workloads, from image processing to data mining. Given its importance, researchers propose different algorithms and hardware-accelerated implementations. While algorithm suitability can depend on the target use case, there is much less doubt about the architecture: FPGAs are the de facto standard, as the design can be perfectly tailored to the target use case. Despite this, AI accelerators such as GPUs and Neural Processing Units (NPUs) are gaining traction. The former attains remarkable performance at the cost of low energy efficiency. The latter, instead, promises to maximize both, but they are strongly underutilized due to the lack of a clear approach for K-Means acceleration. Considering AMD NPU, for example, the main computing cores are AI Engines that require algorithm reshaping and code optimization to harness data parallelism effectively. Thus, this research analyzes different K-Means versions to propose a vectorized algorithm that fully uses AI Engine (AIE) features. We validate our vectorized K-Means on Versal VCK5000, using FPGAs for data movement only, as the Memory Transfer Engines and Shim Tiles of NPUs, and the AI Engine for computation. This design reflects features of modern NPUs, making the validation fair. We attain up to$59.5 \times$speedup against Torch library on GPUs while being comparable but more energy efficient than further optimized GPU solutions.
Eleonora Cabai, Giuseppe Sorrentino, Marco D. Santambrogio, Davide Conficconi
FPL4
2025 Towards Accelerated Healthcare Federated System Through Heterogeneous Accelerators
abstract
Federated Learning (FL) enables collaborative model training across multiple hosts without exposing sensitive data. Yet, privacy-preserving is critical, introducing a nonnegligible overhead. Furthermore, while GPUs perfectly suit AI workloads, FPGAs outperform them in networking and cryptographic tasks as Fully Homomorphic Encryption (FHE), widely employed in FL to safeguard privacy. In this context, applications like Image Registration (IR), requiring different computeintensive steps, each suitable for a different architecture, even worsen the hardware dichotomy. Thus, this research explores innovative federated network structures to reduce privacy risks while assigning roles to each host, guaranteeing complete hardware acceleration per each compute-intensive task by partitioning each step on different hardware layers of modern heterogeneous platforms as Neural Processing Unitss (NPUs), combining AI Engines (AIEs) and GPUs, Versal VCK5000, combining FPGA and AIEs, and multi-platform peer-to-peer systems.
Giuseppe Sorrentino, Davide Conficconi
FPL2
2025 VOTED - Versal Optimization Toolkit for Education and Heterogeneous Systems Development
abstract
Despite classic educational approaches proving their effectiveness for classic hardware acceleration, they suffer novel heterogeneous systems such as Versal, requiring a deeper system-level awareness. Versal devices merge FPGA on-field programma-bility with the performance of hardened VLIW processors, namely AI Engine, at the cost of facing novel challenges when integrating these two different layers. Given the system complexity of such devices and the low-level knowledge required to leverage them, Versal potentialities have yet to be fully exploited. Therefore, we present VOTED, a Versal Optimization Toolkit for Education and Heterogeneous Systems Development, that guides users of any expertise in discovering, using, and optimizing Versal-based applications. VOTED proved to be effective, leading five different groups of students, at their first experience with heterogeneous system design, at devising quite complex Versal-based applications capable of reaching the final stages of a design competition and even winning it with four months of work.
Giuseppe Sorrentino, Paolo Salvatore Galfano, Eleonora D'Arnese, Davide Conficconi
ISCAS4
2025 DFlows: A Flow-Based Programming Approach for a Polyglot Design-Space Exploration Framework
abstract
Current architectural Design-Space Exploration (DSE) tools specify the exploration problem through annotations or pragmas. However, this approach is inherently language-dependent and limits the applicability to one specific target language and synthesis toolchain. Additionally, the rapid development of new hardware Domain-Specific Languages, programming models, and different exploration heuristics calls for a language-agnostic and modular approach. To address this need, we present a DSE formalization to facilitate the integration of new components and customized flows and leverage it to implement DFlows , a flow-based-programming DSE tool that decouples problem definition, code generation, exploration, and evaluation strategies. DFlows ’s compiler-based frontend provides language-agnostic generation of design points through Abstract Syntax Tree manipulation. We show how DFlows can integrate custom performance models from complex state-of-the-art accelerators for Verilog, VHDL, Chisel, and HLS designs. We compare the runtimes of our DSE process against a state-of-the-art Chisel-based DSE tool, achieving up to 3.74× speedup while identifying the same set of optimal solutions. Additionally, we integrate in DFlows a custom exploration heuristic leveraging genetic algorithms and a novel online learning fitness function approximation methodology. This approximation yields a negligible hypervolume difference with the exhaustive search Pareto-front while improving DSE runtime by up to 2.67×.
Francesco Peverelli, Daniele Paletti, Davide Conficconi
ACM Trans. Reconfigurable Technol. Syst.3
2025 QUEKUF: An FPGA Union Find Decoder for Quantum Error Correction on the Toric Code
abstract
Quantum computing represents an exciting computing paradigm that promises to solve problems untractable for a classical computer. The main limiting factor for quantum devices is the noise impacting qubits, which hinders the superpolynomial speedup promise. Thus, although Quantum Error Correction (QEC) mechanisms are paramount, QEC demands high speed and low latency to scale quantum computations to real-life-sized problems. Within this context, hardware accelerators, such as Field Programmable Gate Arrays (FPGAs), represent a valuable approach to fulfilling QEC requirements. Nevertheless, the literature falls short in proposing solutions targeting the toric code, a type of quantum Low-Density Parity Check code capable of encoding two logical qubits, thus requiring fewer physical qubits. This manuscript presents QUEKUF , an FPGA-based QEC dataflow architecture dealing with the toric code. QUEKUF disposes of parallel processing units to spatially parallelize QEC, which a centralized controller orchestrates for data movement and operation decisions. We also provide a latency-oriented resource optimization model to identify the best theoretical configuration of QUEKUF that minimizes latency and optimizes resource requirements based upon high-level quantum parameters. Experimental results show that QUEKUF attains up to \(7.30\times\) speedup and \(81.51\times\) improvement in energy efficiency over a C++ implementation with error-free syndromes while keeping high accuracy.
Federico Valentino, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.3
2025 Rock the QASBA: Quantum Error Correction Acceleration via the Sparse Blossom Algorithm on FPGAs
abstract
Quantum computing is a new paradigm of computation that exploits principles from quantum mechanics to achieve an exponential speedup compared to classical logic. However, noise strongly limits current quantum hardware, reducing achievable performance and limiting the scaling of the applications. For this reason, current noisy intermediate-scale quantum devices require Quantum Error Correction (QEC) mechanisms to identify errors occurring in the computation and correct them in real time. Nevertheless, the high computational complexity of QEC algorithms is incompatible with the tight time constraints of quantum devices. Thus, hardware acceleration is paramount to achieving real-time QEC. This work presents QASBA, an FPGA-based hardware accelerator for the Sparse Blossom Algorithm (SBA), a state-of-the-art decoding algorithm. After profiling the state-of-the-art software counterpart, we developed a design methodology for hardware development based on the SBA. We also devised an automation process to help users without expertise in hardware design in deploying architectures based on QASBA. We implement QASBA on different FPGA architectures and experimentally evaluate resource usage, execution time, and energy efficiency of our solution. Our solution attains up to \(25.05\times\) speedup and \(304.16\times\) improvement in energy efficiency compared to the software baseline.
Marco Venere, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.3
2024 One Automaton to Rule Them All: Beyond Multiple Regular Expressions Execution
abstract
Regular Expressions (REs) matching is crucial to identify strings exhibiting certain morphological properties in a data stream, resulting paramount in contexts such as deep packet inspection in computer security and genome analysis in bioinformatics. Yet, due to their intrinsic data-dependence characteristics, REs represent a complex computational kernel, and numerous solutions investigate pattern-matching efficiency in different directions. However, most of them lack a comprehensive ruleset optimization approach to truly push the pattern matching performance when considering multiple REs together. Thus, exploiting REs morphological similarities within the same dataset allows memory reduction when storing the patterns and drastically improves the dataset-matching throughput. Based on this observation, we propose the Multi-RE Finite State Automata (MFSA) that extends the Finite State Automata (FSA) model to improve REs parallelization by leveraging similarities within a specific application ruleset. We design a multi-level compilation framework to manage REs merging and optimization to produce MFSA(s). Furthermore, we extend iNFAnt algorithm for MFSAs execution with the novel iMFAnt engine. Our evaluation investigates the MFSA size-reduction impact and the execution throughput compared with the one of multiple FSA in both single-and multi-threaded configurations. This approach shows an average 71.95% compression in terms of states, introducing limited compilation time overhead. Besides, best iMFAnt achieves a geomean$5.99\times$throughput improvement and$4.05\times$speedup against single and multiple parallel FSAs.
Luisa Cicolini, Filippo Carloni, Marco D. Santambrogio, Davide Conficconi
CGO4
2024 ALVEARE: a Domain-Specific Framework for Regular Expressions
abstract
Regular Expression (RE) matching enables the identification of patterns in datastreams of heterogeneous fields ranging from proteomics to computer security. These scenarios require massive data analysis that, combined with the high data dependency of the REs, leads to long computational times and high energy consumption. Currently, RE engines rely on either (1) flexibility in run-time RE changes and broad operators support impairing performance or (2) fixed high-performing accelerators implementing few simple RE operators. To overcome these limitations, we propose ALVEARE: a hardware-software approach combining a Domain-Specific Language (DSL) with an embedded Domain-Specific Architecture. We exploit REs as a DSL by translating them into flexible executables through our RISC-based Instruction Set Architecture that expresses from simple to advanced primitives. Then, we design a speculation-based microarchitecture to execute real benchmarks efficiently. ALVEARE provides RE-domain flexibility and broad operators' support and achieves up to 34× speedup and 57× energy efficiency improvements against the state-of-the-art RE2 and Bluefield DPU 2 with its RE accelerator.
Filippo Carloni, Davide Conficconi, Marco D. Santambrogio
DAC2
2024 Co-Designing a 3D Transformation Accelerator for Versal-Based Image Registration
abstract
Rigid image registration is pivotal in modern imaging for correcting distortions of images acquired with different modalities or at different time instants. Literature accelerates its main compute-intensive steps, image transformation and similarity metric computation, through GPUs or FPGAs to meet performance-efficiency constraints. However, GPUs lack energy efficiency while FPGAs lack performance for image transformation. Therefore, we adopt a single heterogeneous system through PEGASO, a methodology and its implementation to co-design the image transformation algorithm on Versal system. We maximize performance by combining custom data layout and hardware optimizations, attaining a 19x speedup over the best GPU-based transformation accelerator. When integrated with FPGA-based similarity metric, PEGASO achieves 82.73x and 20.73x speedup against FPGA- and GPU-based solutions while improving the corresponding energy efficiency of 318.18x and 52.62x.
Paolo Salvatore Galfano, Giuseppe Sorrentino, Eleonora D'Arnese, Davide Conficconi
ICCD4
2024 SATL: A Spatial Architecture Rapid Prototyping Framework for Irregular Applications Acceleration
abstract
Modern FPGA HLS tools are proficient at accelerating datapath applications, but they generate considerable overhead when dealing with irregular, control-driven workloads. Conversely, RTL-based approaches significantly increase development time and system integration effort. To address this tooling gap, we present SATL, a Chisel-based rapid prototyping framework for building FPGA-based spatial architectures targeting irregular workloads. We use it to re-implement YoseUe, a state-of-the-art accelerator for inferring Decision Tree Ensemble Machine Learning models. Compared to the original HLS-based work, SATL yields an average 3.4 x throughput improvement and reduces the architecture's resource consumption, allowing inference on Ensemble Models with up to 3.6 x more trees, enabling the deployment of larger models on resource-constrained devices.
Francesco Peverelli, Alessandro Verosimile, Davide Conficconi, Andrea Damiani, Marco D. Santambrogio
ICCD3
2024 Starlight: A kernel optimizer for GPU processing
abstract
Over the past few years, GPUs have found widespread adoption in many scientific domains, offering notable performance and energy efficiency advantages compared to CPUs. However, optimizing GPU high-performance kernels poses challenges given the complexities of GPU architectures and programming models. Moreover, current GPU development tools provide few high-level suggestions and overlook the underlying hardware. Here we present Starlight, an open-source, highly flexible tool for enhancing GPU kernel analysis and optimization. Starlight autonomously describes Roofline Models, examines performance metrics, and correlates these insights with GPU architectural bottlenecks. Additionally, Starlight predicts potential performance enhancements before altering the source code. We demonstrate its efficacy by applying it to literature genomics and physics applications, attaining speedups from 1.1× to 2.5× over state-of-the-art baselines. Furthermore, Starlight supports the development of new GPU kernels, which we exemplify through an image processing application, showing speedups of 12.7× and 140× when compared against state-of-the-art FPGA- and GPU-based solutions.
Alberto Zeni, Emanuele Del Sozzo, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio
J. Parallel Distributed Comput.4
2024 NERONE: The Fast Way to Efficiently Execute Your Deep Learning Algorithm at the Edge
abstract
Semantic segmentation and classification are pivotal in many clinical applications, such as radiation dose quantification and surgery planning. While manually labeling images is highly time-consuming, the advent of Deep Learning (DL) has introduced a valuable alternative. Nowadays, DL models inference is run on Graphics Processing Units (GPUs), which are power-hungry devices, and, therefore, are not the most suited solution in constrained environments where Field Programmable Gate Arrays (FPGAs) become an appealing alternative given their remarkable performance per watt ratio. Unfortunately, FPGAs are hard to use for non-experts, and the creation of tools to open their employment to the computer vision community is still limited. For these reasons, we propose NERONE, which allows end users to seamlessly benefit from FPGA acceleration and energy efficiency without modifying their DL development flows. To prove the capability of NERONE to cover different network architectures, we have developed four models, one for each of the chosen datasets (three for segmentation and one for classification), and we deployed them, thanks to NERONE, on three different embedded FPGA-powered boards achieving top average energy efficiency improvements of 3.4× and 1.9× against a mobile and a datacenter GPU devices, respectively.
Raffaele Berzoini, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio
IEEE J. Biomed. Health Informatics3
2024 Across Time and Space: Senju's Approach for Scaling Iterative Stencil Loop Accelerators on Single and Multiple FPGAs
abstract
Stencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional point grid multiple times or until convergence. However, due to their iterative and intensive nature, ISLs are highly performance-hungry, demanding specialized solutions. Here, Field Programmable Gate Arrays (FPGAs) represent a valid architectural choice as they enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this article introduces Senju , an automation framework for the design of highly parallel ISL accelerators targeting single-/multi-FPGA systems. Given an input description, Senju automates the entire design process and provides accurate performance estimations. The experimental evaluation shows remarkable and scalable results, outperforming single- and multi-FPGA literature approaches under different metrics. Finally, we present a new analysis of temporal and spatial parallelism trade-offs in a real-case scenario and discuss our performance through a single- and novel specialized multi-FPGA formulation of the Roofline Model.
Emanuele Del Sozzo, Davide Conficconi, Kentaro Sano
ACM Trans. Reconfigurable Technol. Syst.2
2023 Senju: A Framework for the Design of Highly Parallel FPGA-based Iterative Stencil Loop Accelerators
abstract
Stencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving, seismic simulations, and financial option pricing, to name a few. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional system of points until it reaches convergence. However, due to their iterative and computationally intensive nature, these workloads are highly performance-hungry, demanding specialized solutions to boost performance and reduce power consumption. Here, FPGAs represent a valid architectural choice as their peculiar features enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this paper introduces Senju, an automation framework for FPGA-based ISL accelerators. Starting from an input description, Senju builds highly parallel hardware modules and automatizes all their design phases. The experimental evaluation shows remarkable and scalable results, reaching significant performance and energy efficiency improvements compared to the other single-FPGA literature approaches.
Emanuele Del Sozzo, Davide Conficconi, Marco D. Santambrogio, Kentaro Sano
FPGA2
2023 YARB: a Methodology to Characterize Regular Expression Matching on Heterogeneous Systems
abstract
The continuous growth of data pushes novel and efficient approaches for information retrieval. In this context, Regular Expression (RE) matching is widely employed and represents a relevant computational kernel that carries control-and memory-related issues. Among the several solutions to relieve these burdens, accelerators seem a promising alternative to general-purpose systems. However, state-of-the-art benchmarking presents a highly fragmented scenario without consensus on the approach and lacks an open-source strategy. Therefore, to fairly characterize existing execution engines, this work presents YARB, an open benchmarking methodology. It builds upon literature solutions, a comprehensive approach, and an in-depth characterization of heterogeneous systems. Moreover, YARB's openness will enable future integrations and engines comparison.
Filippo Carloni, Davide Conficconi, Ilaria Moschetto, Marco D. Santambrogio
ISCAS2
2023 Hephaestus: Codesigning and Automating 3D Image Registration on Reconfigurable Architectures
abstract
Healthcare is a pivotal research field, and medical imaging is crucial in many applications. Therefore finding new architectural and algorithmic solutions would benefit highly repetitive image processing procedures. One of the most complex tasks in this sense is image registration, which finds the optimal geometric alignment among 3D image stacks and is widely employed in healthcare and robotics. Given the high computational demand of such a procedure, hardware accelerators are promising real-time and energy-efficient solutions, but they are complex to design and integrate within software pipelines. Therefore, this work presents an automation framework called Hephaestus that generates efficient 3D image registration pipelines combined with reconfigurable accelerators. Moreover, to alleviate the burden from the software, we codesign software-programmable accelerators that can adapt at run-time to the image volume dimensions. Hephaestus features a cross-platform abstraction layer that enables transparently high-performance and embedded systems deployment. However, given the computational complexity of 3D image registration, the embedded devices become a relevant and complex setting being constrained in memory; thus, they require further attention and tailoring of the accelerators and registration application to reach satisfactory results. Therefore, with Hephaestus , we also propose an approximation mechanism that enables such devices to perform the 3D image registration and even achieve, in some cases, the accuracy of the high-performance ones. Overall, Hephaestus demonstrates 1.85× of maximum speedup, 2.35× of efficiency improvement with respect to the State of the Art, a maximum speedup of 2.51× and 2.76× efficiency improvements against our software, while attaining state-of-the-art accuracy on 3D registrations.
Giuseppe Sorrentino, Marco Venere, Davide Conficconi, Eleonora D'Arnese, Marco D. Santambrogio
ACM Trans. Embed. Comput. Syst.3
2023 Faber: A Hardware/SoftWare Toolchain for Image Registration
abstract
Image registration is a well-defined computation paradigm widely applied to align one or more images to a target image. This paradigm, which builds upon three main components, is particularly compute-intensive and represents many image processing pipelines’ bottlenecks. State-of-the-art solutions leverage hardware acceleration to speed up image registration, but they are usually limited to implementing a single component. We present Faber, an open-source HW/SW CAD toolchain tailored to image registration. The Faber toolchain comprises HW/SW highly-tunable registration components, supports users with different expertise in building custom pipelines, and automates the design process. In this direction, Faber provides both default settings for entry-level users and latency and resource models to guide HW experts in customizing the different components. Finally, Faber achieves from 1.5× to 54× in speedup and from 2× to 177× in energy efficiency against state-of-the-art tools on a Xeon Gold.
Eleonora D'Arnese, Davide Conficconi, Emanuele Del Sozzo, Luigi Fusco, Donatella Sciuto, Marco D. Santambrogio
IEEE Trans. Parallel Distributed Syst.2
2021 A Framework for Customizable FPGA-based Image Registration Accelerators
abstract
Image Registration is a highly compute-intensive optimization procedure that determines the geometric transformation to align a floating image to a reference one. Generally, the registration targets are images taken from different time instances, acquisition angles, and/or sensor types. Several methodologies are employed in the literature to address the limiting factors of this class of algorithms, among which hardware accelerators seem the most promising solution to boost performance. However, most hardware implementations are either closed-source or tailored to a specific context, limiting their application to different fields. For these reasons, we propose an open-source hardware-software framework to generate a configurable architecture for the most compute-intensive part of registration algorithms, namely the similarity metric computation. This metric is the Mutual Information, a well-known calculus from the Information Theory, used in several optimization procedures. Through different design parameters configurations, we explore several design choices of our highly-customizable architecture and validate it on multiple FPGAs. We evaluated various architectures against an optimized Matlab implementation on an Intel Xeon Gold, reaching a speedup up to 2.86x, and remarkable performance and power efficiency against other state-of-the-art approaches.
Davide Conficconi, Eleonora D'Arnese, Emanuele Del Sozzo, Donatella Sciuto, Marco D. Santambrogio
FPGA1
2021 CICERO: A Domain-Specific Architecture for Efficient Regular Expression Matching
abstract
Regular Expression (RE) matching is a computational kernel used in several applications. Since RE complexity and data volumes are steadily increasing, hardware acceleration is gaining attention also for this problem. Existing approaches have limited flexibility as they require a different implementation for each RE. On the other hand, it is complex to map efficient RE representations like non-deterministic finite-state automata onto software-programmable engines or parallel architectures. In this work, we present CICERO , an end-to-end framework composed of a domain-specific architecture and a companion compilation framework for RE matching. Our solution is suitable for many applications, such as genomics/proteomics and natural language processing. CICERO aims at exploiting the intrinsic parallelism of non-deterministic representations of the REs. CICERO can trade-off accelerators’ efficiency and processors’ flexibility thanks to its programmable architecture and the compilation framework. We implemented CICERO prototypes on embedded FPGA achieving up to 28.6× and 20.8× more energy efficiency than embedded and mainstream processors, respectively. Since it is a programmable architecture, it can be implemented as a custom ASIC that is orders of magnitude more energy-efficient than mainstream processors.
Daniele Parravicini, Davide Conficconi, Emanuele Del Sozzo, Christian Pilato, Marco D. Santambrogio
ACM Trans. Embed. Comput. Syst.2
2021 Enhancing the Scalability of Multi-FPGA Stencil Computations via Highly Optimized HDL Components
abstract
Stencil-based algorithms are a relevant class of computational kernels in high-performance systems, as they appear in a plethora of fields, from image processing to seismic simulations, from numerical methods to physical modeling. Among the various incarnations of stencil-based computations,Iterative Stencil Loops (ISLs)andConvolutional Neural Networks (CNNs)represent two well-known examples of kernels belonging to the stencil class. Indeed, ISLs apply the same stencil several times until convergence, while CNN layers leverage stencils to extract features from an image. The computationally intensive essence of ISLs, CNNs, and in general stencil-based workloads, requires solutions able to produce efficient implementations in terms of throughput and power efficiency. In this context, FPGAs are ideal candidates for such workloads, as they allow design architectures tailored to the stencil regular computational pattern. Moreover, the ever-growing need for performance enhancement leads FPGA-based architectures to scale to multiple devices to benefit from a distributed acceleration. For this reason, we propose a library of HDL components to effectively compute ISLs and CNNs inference on FPGA, along with a scalable multi-FPGA architecture, based on custom PCB interconnects. Our solution eases the design flow and guarantees both scalability and performance competitive with state-of-the-art works.
Enrico Reggiani, Emanuele Del Sozzo, Davide Conficconi, Giuseppe Natale, Carlo Moroni, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.3
2019 Optimizing Bit-Serial Matrix Multiplication for Reconfigurable Computing
abstract
Matrix-matrix multiplication is a key computational kernel for numerous applications in science and engineering, with ample parallelism and data locality that lends itself well to high-performance implementations. Many matrix multiplication-dependent applications can use reduced-precision integer or fixed-point representations to increase their performance and energy efficiency while still offering adequate quality of results. However, precision requirements may vary between different application phases or depend on input data, rendering constant-precision solutions ineffective. BISMO, a vectorized bit-serial matrix multiplication overlay for reconfigurable computing, previously utilized the excellent binary-operation performance of FPGAs to offer a matrix multiplication performance that scales with required precision and parallelism. We show how BISMO can be scaled up on Xilinx FPGAs using an arithmetic architecture that better utilizes six-input LUTs. The improved BISMO achieves a peak performance of 15.4 binary TOPS on the Ultra96 board with a Xilinx UltraScale+ MPSoC.
Yaman Umuroglu, Davide Conficconi, Lahiru Rasnayake, Thomas B. Preußer, Magnus Själander
ACM Trans. Reconfigurable Technol. Syst.2