Lorenzo Carpentieri

dblp:344/3586 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0001-2041-7618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 SYprox: Combining Host and Device Perforation with Mixed Precision Approximation on Heterogeneous Architectures
abstract
Approximate computing is an emerging paradigm that aims to exploit the inherent error tolerance of many applications, particularly in domains such as image processing and machine learning.Taking advantage of this property, applications can trade off accuracy for significant gains in performance and power consumption.Existing approximation techniques for GPUs are limited to very specific approaches, do not fully exploit the host-device execution model, and are often restricted in terms of programming models and supported target hardware.This paper introduces SYprox, a new approximate computing framework based on SYCL that allows programmers to easily implement heterogeneous approximated applications.SYprox supports multiple techniques, including data perforation, signal reconstruction, and mixed precision, and allows them to be combined to support a wide range of approximations.In particular, SYprox extends existing perforation approaches to allow both host and device data perforation.Experimental results show that SYprox's approximations are Pareto dominant with respect to state-of-the-art approaches and are portable to AMD, Intel and NVIDIA GPUs.
Lorenzo Carpentieri, Biagio Cosenza
ICS1
2025 Phase-Based Frequency Scaling for Energy-Efficient Heterogeneous Computing
abstract
Energy efficiency has been a major challenge for exascale computing. Frequency scaling is a powerful technique to achieve energy savings in modern heterogeneous systems, and can be applied either at a coarse granularity, by application, or at a fine granularity, by setting the frequency for each computational kernel. The chosen granularity significantly impacts the performance and energy consumption of applications due to frequency-change overhead. We propose a novel phase-based method that minimizes the frequency-change overhead and improves performance and energy efficiency on heterogeneous multi-GPU systems. Our approach detects different phases through application profiling and DAG analysis, and sets an optimal frequency for each phase. Our methodology also considers MPI programs, where the overhead can be hidden by overlapping frequency-change with communication. Experimental results show up to 37 % energy saving and$1.87 \times$speedup for various benchmarks on a single GPU, and 68 % energy saving and$3.63 \times$speedup on two multiGPU applications.
Lorenzo Carpentieri, Antonio De Caro, Majid Salimi Beni, Kaijie Fan, Biagio Cosenza
IPDPS1
2025 A Performance Analysis of Autovectorization on RVV RISC-V Boards
Lorenzo Carpentieri, Mohammad VazirPanah, Biagio Cosenza
PDP1
2024 Enabling performance portability on the LiGen drug discovery pipeline
abstract
In recent years, there has been a growing interest in developing high-performance implementations of drug discovery processing software. To target modern GPU architectures, such applications are mostly written in proprietary languages such as CUDA or HIP. However, with the increasing heterogeneity of modern HPC systems and the availability of accelerators from multiple hardware vendors, it has become critical to be able to efficiently execute drug discovery pipelines on multiple large-scale computing systems, with the ultimate goal of working on urgent computing scenarios. This article presents the challenges of migrating LiGen, an industrial drug discovery software pipeline, from CUDA to the SYCL programming model, an industry standard based on C++ that enables heterogeneous computing. We perform a structured analysis of the performance portability of the SYCL LiGen platform, focusing on different aspects of the approach from different perspectives. First, we analyze the performance portability provided by the high-level semantics of SYCL, including the most recent group algorithms and subgroups of SYCL 2020. Second, we analyze how low-level aspects such as kernel occupancy and register pressure affect the performance portability of the overall application. The experimental evaluation is performed on two different versions of LiGen, implementing two different parallelization patterns, by comparing them with a manually optimized CUDA version, and by evaluating performance portability using both known and ad hoc metrics. The results show that, thanks to the combination of high-level SYCL semantics and some manual tuning, LiGen achieves native-comparable performance on NVIDIA, while also running on AMD GPUs.
Luigi Crisci, Lorenzo Carpentieri, Biagio Cosenza, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Gianluca Palermo, Andrea Beccari
Future Gener. Comput. Syst.2
2023 SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy Saving
abstract
Energy-efficient computing uses power management techniques such as frequency scaling to save energy. Implementing energy-efficient techniques on large-scale computing systems is challenging for several reasons. While most modern architectures, including GPUs, are capable of frequency scaling, these features are often not available on large systems. In addition, achieving higher energy savings requires precise energy tuning because not only applications but also different kernels can have different energy characteristics. We propose SYnergy, a novel energy-efficient approach that spans languages, compilers, runtimes, and job schedulers to achieve unprecedented fine-grained energy savings on large-scale heterogeneous clusters. SYnergy defines an extension to the SYCL programming model that allows programmers to define a specific energy goal for each kernel. For example, a kernel can aim to minimize well-known energy metrics such as EDP and ED2P or to achieve predefined energy-performance tradeoffs, such as the best performance with 25% energy savings. Through compiler integration and a machine learning model, each kernel is statically optimized for the specific target. On large computing systems, a SLURM plug-in allows SYnergy to run on all available devices in the cluster, providing scalable energy savings. The methodology is inherently portable and has been evaluated on both NVIDIA and AMD GPUs. Experimental results show unprecedented improvements in energy and energy-related metrics on real-world applications, as well as scalable energy savings on a 64-GPU cluster.
Kaijie Fan, Marco D'Antonio, Lorenzo Carpentieri, Biagio Cosenza, Federico Ficarelli, Daniele Cesarini
SC3