EDBT 2026 Demo / reviewers in the wild / expert
Enrico Calore
dblp:147/6877
· DBLP profile ↗
8ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0002-2301-3838ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking a DNN for aortic valve calcium lesions segmentation on FPGA-based DPU using the vitis AI toolchainabstractSemantic segmentation assigns a class to every pixel of an image to automatically locate objects in the context of computer vision applications for autonomous vehicles, robotics, agriculture, gaming, and medical imaging. Deep Neural Network models, such as Convolutional Neural Networks (CNNs), are widely used for this purpose. Among the plethora of models, the U-Net is a standard in biomedical imaging. Nowadays, GPUs efficiently perform segmentation and are the reference architectures for running CNNs, and FPGAs compete for inferences among alternative platforms, promising higher energy efficiency and lower latency solutions. In this contribution, we evaluate the performance of FPGA-based Deep Processing Units (DPUs) implemented on the AMD Alveo U55C for the inference task, using calcium segmentation in cardiac aortic valve computer tomography scans as a benchmark. We design and implement a U-Net-based application, optimize the hyperparameters to maximize the prediction accuracy, perform pruning to simplify the model, and use different numerical quantizations to exploit low-precision operations supported by the DPUs and GPUs to boost the computation time. We describe how to port and deploy the U-Net model on DPUs, and we compare accuracy, throughput, and energy efficiency achieved with four generations of GPUs and a recent dual 32-core high-end CPU platform. Our results show that a complex DNN like the U-Net can run effectively on DPUs using 8-bit integer computation, achieving a prediction accuracy of approximately 95 % in Dice and 91 % in IoU scores. These results are comparable to those measured when running the floating-point models on GPUs and CPUs. On the one hand, in terms of computing performance, the DPUs achieves a inference latency of approximately 3.5 ms and a throughput of approximately 4.2 kPFS, boosting the performance of a 64-core CPU system by approximately 10 % in terms of latency and a factor 2 X in terms of throughput, but still do not overcoming the performance of GPUs when using the same numerical precision. On the other hand, considering the energy efficiency, the improvements are approximately a factor 6.7 X compared to the CPU, and 1.6 X compared to the P100 GPU manufactured with the same technological process (16 nm). Valentina Sisini, Andrea Miola, Giada Minghini, Enrico Calore, Armando Ugo Cavallo, Sebastiano Fabio Schifano, Cristian Zambelli |
Future Gener. Comput. Syst. | 4 |
| 2025 | High throughput edit distance computation on FPGA-based accelerators using HLSabstractEdit distance is a computational grand challenge problem to quantify the minimum number of editing operations required to modify one string of characters to the other, finding many applications of natural language processing. In recent years, relevant and increasing interest has also emerged from deoxyribonucleic acid (DNA) applications, like Next Generation Sequencing and DNA storage technologies. Both applications share two crucial features: i) the information is coded into the four bases of DNA and ii) the level of operational noise is still high causing errors in the data, requiring inclusion in the workflow of the computation of algorithms such as the edit distance for finding similarities between sequences. To boost this computation many solutions are available in the literature. Among them, the FPGAs are largely used since the data domain of those applications is strings of 4 characters represented as two-bit values, inconveniently fitting the basic data types of ordinary CPUs and GPUs, with additional benefits of providing a high level of parallelism and low processing latency. This contribution presents a computing- and energy-efficient design implementing the edit distance algorithm combining metaprogramming and High-Level Synthesis. We also assess the performance of our design targeting recent FPGA-based accelerators. Our solution uses nearly 90% of FPGA basic-block hardware resources achieving about 90% of computing efficiency delivering a maximum throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule, enabling the use of FPGAs as a new class of accelerators for High Performance Computing in DNA applications. • Computing- and energy-efficient edit distance algorithms running on FPGA accelerators. • Targeting short reads for DNA data storage and long reads for genomics applications. • Combining metaprogramming and HLS to design and optimize C codes for FPGA devices. • Achieving about 90% of computing efficiency and near 90% FPGA resources occupancy. • Delivering a peak throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule. Sebastiano Fabio Schifano, Marco Reggiani, Enrico Calore, Rino Micheloni, Alessia Marelli, Cristian Zambelli |
Future Gener. Comput. Syst. | 3 |
| 2021 | Performance assessment of FPGAs as HPC accelerators using the FPGA Empirical RooflineabstractHardware accelerators are nowadays very common in HPC systems, and GPUs are playing a major role in this. More recently, also FPGAs started to be adopted in few data centers to accelerate specific workloads, and it is expected that in the near future they will be increasingly used also in general purpose HPC systems. FPGAs are already known to provide interesting speedups in several application fields, but to estimate their expected performance in the context of typical HPC workloads is not straightforward. To ease this task, in this paper we present the FPGA Empirical Roofline (FER), a benchmarking tool able to empirically estimate the computing throughput and memory bandwidth of FPGAs, when used as hardware accelerators for HPC applications. Using the widely known Roofline Model as a theoretical foundation, FER allows to measure FPGAs computing throughput and bandwidth upper-bounds, allowing to estimate the performance of function kernels developed using high level synthesis tools, according to their arithmetic intensity. We implemented FER using two different high level paradigms: OmpSs@FPGA and Xilinx Vitis workflow, two promising approaches to develop HPC applications enabling exploitation of FPGAs as hardware accelerators. In this paper we describe the theoretical model on which the FER benchmark relies, as well as its implementation details, and we provide performance results measured on the Xilinx Alveo U250 FPGA. Enrico Calore, Sebastiano Fabio Schifano |
FPL | 1 |
| 2017 | Evaluation of DVFS techniques on modern HPC processors and accelerators for energy-aware applicationsabstractSummary Energy efficiency is becoming increasingly important for computing systems, in particular for large scale High Performance Computing (HPC) facilities. In this work, we evaluate, from a user perspective, the use of Dynamic Voltage and Frequency Scaling techniques, assisted by the power and energy monitoring capabilities of modern processors to tune applications for energy efficiency. We run selected kernels and a full HPC application on 2 high‐end processors widely used in the HPC context, namely, an NVIDIA K80 GPU and an Intel Haswell CPU. We evaluate the available trade‐offs between energy‐to‐solution and time‐to‐solution, attempting a function‐by‐function frequency tuning. We finally estimate the benefits obtainable running the full code on an HPC multi‐GPU node, with respect to default clock frequency governors. We instrument our code to accurately monitor power consumption and execution time without the need of any additional hardware, and we enable it to change CPUs and GPUs clock frequencies while running. We analyze our results on the different architectures using a simple energy‐performance model and derive a number of energy saving strategies, which can be easily adopted on recent high‐end HPC systems for generic applications. Enrico Calore, Alessandro Gabbana, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Performance and portability of accelerated lattice Boltzmann applications with OpenACCabstractSummary An increasingly large number of HPC systems rely on heterogeneous architectures combining traditional multi‐core CPUs with power efficient accelerators. Designing efficient applications for these systems have been troublesome in the past as accelerators could usually be programmed using specific programming languages threatening maintainability, portability, and correctness. Several new programming environments try to tackle this problem. Among them, OpenACC offers a high‐level approach based on compiler directives to mark regions of existing C, C++, or Fortran codes to run on accelerators. This approach directly addresses code portability, leaving to compilers the support of each different accelerator, but one has to carefully assess the relative costs of portable approaches versus computing efficiency. In this paper, we address precisely this issue, using as a test‐bench a massively parallel lattice Boltzmann algorithm. We first describe our multi‐node implementation and optimization of the algorithm, using OpenACC and MPI. We then benchmark the code on a variety of processors, including traditional CPUs and GPUs, and make accurate performance comparisons with other GPU implementations of the same algorithm using CUDA and OpenCL. We also asses the performance impact associated with portable programming, and the actual portability and performance‐portability of OpenACC‐based applications across several state‐of‐the‐art architectures. Copyright © 2016 John Wiley & Sons, Ltd. Enrico Calore, Alessandro Gabbana, Jiri Kraus, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Massively parallel lattice-Boltzmann codes on large GPU clusters
Enrico Calore, Alessandro Gabbana, Jiri Kraus, E. Pellegrini, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Parallel Comput. | 1 |
| 2015 | Accelerating Lattice Boltzmann Applications with OpenACC
Enrico Calore, Jiri Kraus, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Euro-Par | 1 |
| 2014 | Accelerometer-based correction of skewed horizon and keystone distortion in digital photography
Enrico Calore, Iuri Frosio |
Image Vis. Comput. | 1 |