EDBT 2026 Demo / reviewers in the wild / expert
Sebastiano Fabio Schifano
dblp:01/5335
· DBLP profile ↗
13ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-0132-9196ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Outcomes of the ICSC Flagship 2 Project on Architectures and Design Methodologies to Accelerate AI WorkloadsabstractEnergy-efficient hardware accelerators specialized for AI tasks are now being deployed from low-power edge devices to large-scale high-performance computing systems and data centers. This paper presents the main outcomes of the Flagship 2 project of the ICSC Italian National Research Center for High Performance Computing, which focuses on the design techniques for heterogeneous hardware optimized for AI acceleration from the edge to the HPC. In particular, we describe the main challenges addressed and highlight some advances in architectures, technologies, and design methodologies tailored to accelerate deep learning, transformer-based, and generative AI models. We also summarize the most significant outcomes achieved through the close collaboration among the project partners, including the development of design techniques, tools, prototypes, IP cores, and models that collectively advance AI acceleration from the edge to the HPC contexts. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Cristian Zambelli, Sebastiano Fabio Schifano, Francesco Conti 0001, Angelo Garofalo, Luca Benini, Maurizio Palesi, Giuseppe Ascia, Enrico Russo 0002, Fanny Spagnolo, Pasquale Corsonello, Stefania Perri, Fabio Frustaci |
DATE | 6 |
| 2026 | Benchmarking a DNN for aortic valve calcium lesions segmentation on FPGA-based DPU using the vitis AI toolchainabstractSemantic segmentation assigns a class to every pixel of an image to automatically locate objects in the context of computer vision applications for autonomous vehicles, robotics, agriculture, gaming, and medical imaging. Deep Neural Network models, such as Convolutional Neural Networks (CNNs), are widely used for this purpose. Among the plethora of models, the U-Net is a standard in biomedical imaging. Nowadays, GPUs efficiently perform segmentation and are the reference architectures for running CNNs, and FPGAs compete for inferences among alternative platforms, promising higher energy efficiency and lower latency solutions. In this contribution, we evaluate the performance of FPGA-based Deep Processing Units (DPUs) implemented on the AMD Alveo U55C for the inference task, using calcium segmentation in cardiac aortic valve computer tomography scans as a benchmark. We design and implement a U-Net-based application, optimize the hyperparameters to maximize the prediction accuracy, perform pruning to simplify the model, and use different numerical quantizations to exploit low-precision operations supported by the DPUs and GPUs to boost the computation time. We describe how to port and deploy the U-Net model on DPUs, and we compare accuracy, throughput, and energy efficiency achieved with four generations of GPUs and a recent dual 32-core high-end CPU platform. Our results show that a complex DNN like the U-Net can run effectively on DPUs using 8-bit integer computation, achieving a prediction accuracy of approximately 95 % in Dice and 91 % in IoU scores. These results are comparable to those measured when running the floating-point models on GPUs and CPUs. On the one hand, in terms of computing performance, the DPUs achieves a inference latency of approximately 3.5 ms and a throughput of approximately 4.2 kPFS, boosting the performance of a 64-core CPU system by approximately 10 % in terms of latency and a factor 2 X in terms of throughput, but still do not overcoming the performance of GPUs when using the same numerical precision. On the other hand, considering the energy efficiency, the improvements are approximately a factor 6.7 X compared to the CPU, and 1.6 X compared to the P100 GPU manufactured with the same technological process (16 nm). Valentina Sisini, Andrea Miola, Giada Minghini, Enrico Calore, Armando Ugo Cavallo, Sebastiano Fabio Schifano, Cristian Zambelli |
Future Gener. Comput. Syst. | 6 |
| 2025 | Multi-Partner Project: Architectures and Design Methodologies to Accelerate AI Workloads. The ICSC Flagship 2 ProjectabstractRecent pre-exascale and exascale supercomputers have driven the development of increasingly sophisticated AI models for diverse applications, including image recognition and classification, natural language processing, and generative AI. These applications require specialized hardware accelerators, to handle the heavy computational demands of AI algorithms in an energy-efficient manner. Today, AI accelerators are deployed across various systems, from low-power edge devices to large-scale servers, high-performance computing (HPC) infrastructures, and data centers. The primary objective of the ICSC Flagship 2 project, discussed in this paper, is to develop heterogeneous hardware platforms optimized to accelerate HPC and big data applications. Specifically, this paper provides an overview of the key challenges addressed and the achievements realized at the current intermediate stage of the ICSC Flagship 2 project focused on architectures, technologies, and design methodologies to design efficient hardware accelerators for AI workloads, such as deep learning (DL) and transformer models. Cristina Silvano, Fabrizio Ferrandi, Serena Curzel, Daniele Ielmini, Stefania Perri, Fanny Spagnolo, Pasquale Corsonello, Sebastiano Fabio Schifano, Cristian Zambelli, Angelo Garofalo, Francesco Conti 0001, Luca Benini |
DATE | 8 |
| 2025 | Transforming Digital Electronics Education: Integrating Sustainability Through Eco-Design, Modularity, and Circular PracticesabstractIntegrating eco-design principles and open hardware into digital electronics education is crucial to equip future engineers for addressing the environmental issues of the digital era. This research is part of the DEEPCEL (Digital Electronics with Eco-Designed Paradigm in Collaborative Enhanced Learning) ERASMUS+ project. Initially, it explores the knowledge and perceptions of participants through a targeted survey, emphasizing methodologies, the environmental impact of open hardware, and their role in fostering sustainability in digital electronics. The research also explores effective teaching strategies, tools, and platforms that can improve the adoption of eco-transition principles, with the aim of creating guidelines for curriculum modernization focused on sustainability. Participants affirm that open hardware offers transformative potential in improving innovation, accessibility, and eco-conscious practices and community-driven development. This study underscores the need for a collaborative effort between academia, professors, and industry leaders, along with their demands and needs to prioritize and integrate sustainability principles into digital electronics education. However, significant barriers remain, including insufficient training for professors, limited resources, and resistance to changes in traditional pedagogical practices. The findings should be interpreted with caution due to potential limitations in sample representation. Further research will validate the effectiveness of proposed strategies and assess long-term impacts on educational practices. Ignacio Bravo Muñoz, Ernesto Martín, Etienne Lemaire, Jean-Paul Chemla, Cristian Zambelli, Sebastiano Fabio Schifano, Hélio Sousa Mendonça, José Carlos Alves |
EDUCON | 7 |
| 2025 | High throughput edit distance computation on FPGA-based accelerators using HLSabstractEdit distance is a computational grand challenge problem to quantify the minimum number of editing operations required to modify one string of characters to the other, finding many applications of natural language processing. In recent years, relevant and increasing interest has also emerged from deoxyribonucleic acid (DNA) applications, like Next Generation Sequencing and DNA storage technologies. Both applications share two crucial features: i) the information is coded into the four bases of DNA and ii) the level of operational noise is still high causing errors in the data, requiring inclusion in the workflow of the computation of algorithms such as the edit distance for finding similarities between sequences. To boost this computation many solutions are available in the literature. Among them, the FPGAs are largely used since the data domain of those applications is strings of 4 characters represented as two-bit values, inconveniently fitting the basic data types of ordinary CPUs and GPUs, with additional benefits of providing a high level of parallelism and low processing latency. This contribution presents a computing- and energy-efficient design implementing the edit distance algorithm combining metaprogramming and High-Level Synthesis. We also assess the performance of our design targeting recent FPGA-based accelerators. Our solution uses nearly 90% of FPGA basic-block hardware resources achieving about 90% of computing efficiency delivering a maximum throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule, enabling the use of FPGAs as a new class of accelerators for High Performance Computing in DNA applications. • Computing- and energy-efficient edit distance algorithms running on FPGA accelerators. • Targeting short reads for DNA data storage and long reads for genomics applications. • Combining metaprogramming and HLS to design and optimize C codes for FPGA devices. • Achieving about 90% of computing efficiency and near 90% FPGA resources occupancy. • Delivering a peak throughput of 16.8 TCUPS and an energy efficiency of 46 Mpair/Joule. Sebastiano Fabio Schifano, Marco Reggiani, Enrico Calore, Rino Micheloni, Alessia Marelli, Cristian Zambelli |
Future Gener. Comput. Syst. | 1 |
| 2021 | Performance assessment of FPGAs as HPC accelerators using the FPGA Empirical RooflineabstractHardware accelerators are nowadays very common in HPC systems, and GPUs are playing a major role in this. More recently, also FPGAs started to be adopted in few data centers to accelerate specific workloads, and it is expected that in the near future they will be increasingly used also in general purpose HPC systems. FPGAs are already known to provide interesting speedups in several application fields, but to estimate their expected performance in the context of typical HPC workloads is not straightforward. To ease this task, in this paper we present the FPGA Empirical Roofline (FER), a benchmarking tool able to empirically estimate the computing throughput and memory bandwidth of FPGAs, when used as hardware accelerators for HPC applications. Using the widely known Roofline Model as a theoretical foundation, FER allows to measure FPGAs computing throughput and bandwidth upper-bounds, allowing to estimate the performance of function kernels developed using high level synthesis tools, according to their arithmetic intensity. We implemented FER using two different high level paradigms: OmpSs@FPGA and Xilinx Vitis workflow, two promising approaches to develop HPC applications enabling exploitation of FPGAs as hardware accelerators. In this paper we describe the theoretical model on which the FER benchmark relies, as well as its implementation details, and we provide performance results measured on the Xilinx Alveo U250 FPGA. Enrico Calore, Sebastiano Fabio Schifano |
FPL | 2 |
| 2017 | Evaluation of DVFS techniques on modern HPC processors and accelerators for energy-aware applicationsabstractSummary Energy efficiency is becoming increasingly important for computing systems, in particular for large scale High Performance Computing (HPC) facilities. In this work, we evaluate, from a user perspective, the use of Dynamic Voltage and Frequency Scaling techniques, assisted by the power and energy monitoring capabilities of modern processors to tune applications for energy efficiency. We run selected kernels and a full HPC application on 2 high‐end processors widely used in the HPC context, namely, an NVIDIA K80 GPU and an Intel Haswell CPU. We evaluate the available trade‐offs between energy‐to‐solution and time‐to‐solution, attempting a function‐by‐function frequency tuning. We finally estimate the benefits obtainable running the full code on an HPC multi‐GPU node, with respect to default clock frequency governors. We instrument our code to accurately monitor power consumption and execution time without the need of any additional hardware, and we enable it to change CPUs and GPUs clock frequencies while running. We analyze our results on the different architectures using a simple energy‐performance model and derive a number of energy saving strategies, which can be easily adopted on recent high‐end HPC systems for generic applications. Enrico Calore, Alessandro Gabbana, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Performance and portability of accelerated lattice Boltzmann applications with OpenACCabstractSummary An increasingly large number of HPC systems rely on heterogeneous architectures combining traditional multi‐core CPUs with power efficient accelerators. Designing efficient applications for these systems have been troublesome in the past as accelerators could usually be programmed using specific programming languages threatening maintainability, portability, and correctness. Several new programming environments try to tackle this problem. Among them, OpenACC offers a high‐level approach based on compiler directives to mark regions of existing C, C++, or Fortran codes to run on accelerators. This approach directly addresses code portability, leaving to compilers the support of each different accelerator, but one has to carefully assess the relative costs of portable approaches versus computing efficiency. In this paper, we address precisely this issue, using as a test‐bench a massively parallel lattice Boltzmann algorithm. We first describe our multi‐node implementation and optimization of the algorithm, using OpenACC and MPI. We then benchmark the code on a variety of processors, including traditional CPUs and GPUs, and make accurate performance comparisons with other GPU implementations of the same algorithm using CUDA and OpenCL. We also asses the performance impact associated with portable programming, and the actual portability and performance‐portability of OpenACC‐based applications across several state‐of‐the‐art architectures. Copyright © 2016 John Wiley & Sons, Ltd. Enrico Calore, Alessandro Gabbana, Jiri Kraus, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Massively parallel lattice-Boltzmann codes on large GPU clusters
Enrico Calore, Alessandro Gabbana, Jiri Kraus, E. Pellegrini, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Parallel Comput. | 5 |
| 2015 | Accelerating Lattice Boltzmann Applications with OpenACC
Enrico Calore, Jiri Kraus, Sebastiano Fabio Schifano, Raffaele Tripiccione |
Euro-Par | 3 |
| 2013 | Benchmarking MIC architectures with Monte Carlo simulations of spin glass systemsabstractSpin glasses - theoretical models used to capture several physical properties of real glasses - are mostly studied by Monte Carlo simulations. The associated algorithms have a very large and easily identifiable degree of available parallelism, that can also be easily cast in SIMD form. State-of-the-art multi- and many-core processors and accelerators are therefore a promising computational platform to support these Grand Challenge applications. In this paper we port and optimize for many-core processors a Monte Carlo code for the simulation of the 3D Edwards Anderson spin glass, focusing on a dual eight-core Sandy Bridge processor, and on a Xeon-Phi co-processor based on the new Many Integrated Core architecture. We present performance results, discuss bottlenecks preventing further performance gains and compare with the corresponding figures for GPU-based implementations and for application-specific dedicated machines. Alessandro Gabbana, Marcello Pivanti, Sebastiano Fabio Schifano, Raffaele Tripiccione |
HiPC | 3 |
| 2013 | Benchmarking GPUs with a Parallel Lattice-Boltzmann CodeabstractAccelerators are an increasingly common option to boost performance of codes that require extensive number crunching. In this paper we report on our experience with NVIDIA accelerators to study fluid systems using the Lattice Boltzmann (LB) method. The regular structure of LB algorithms makes them suitable for processor architectures with a large degree of parallelism, such as recent multi- and many-core processors and GPUs; however, the challenge of exploiting a large fraction of the theoretically available performance of this new class of processors is not easily met. We consider a state-of-theart two-dimensional LB model based on 37 populations (a D2Q37 model), that accurately reproduces the thermo-hydrodynamics of a 2D-fluid obeying the equation-of-state of a perfect gas. The computational features of this model make it a significant benchmark to analyze the performance of new computational platforms, since critical kernels in this code require both high memory-bandwidth on sparse memory addressing patterns and floating-point throughput. In this paper we consider two recent classes of GPU boards based on the Fermi and Kepler architectures; we describe in details all steps done to implement and optimize our LB code and analyze its performance first on single- GPU systems, and then on parallel multi-GPU systems based on one node as well as on a cluster of many nodes; in the latter case we use CUDA-aware MPI as an abstraction layer to assess the advantages of advanced GPU-to-GPU communication technologies like GPUDirect. On our implementation, aggregate sustained performance of the most compute intensive part of the code breaks the 1 double-precision Tflops barrier on a single-host system with two GPUs. Jiri Kraus, Marcello Pivanti, Sebastiano Fabio Schifano, Raffaele Tripiccione, Marco Zanella |
SBAC-PAD | 3 |
| 2005 | The Potential of On-Chip Multiprocessing for QCD Machines
Gianfranco Bilardi, Andrea Pietracaprina, Geppino Pucci, Sebastiano Fabio Schifano, Raffaele Tripiccione |
HiPC | 4 |