Biagio Cosenza

dblp:07/527 · DBLP profile ↗
← Back
42ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-8869-6705ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 4 first-author · 19 since 2021Software engineering, systems software and programming languages · 2
YearPublicationVenuePosition
2026 miniLB: Benchmarking Lattice Boltzmann simulations on AMD, Intel, and NVIDIA GPUs
abstract
In computational fluid dynamics, the Lattice Boltzmann method is a computational technique that has gained popularity due to its flexibility in handling complex geometries and turbulence models, and its unique suitability for massive parallel processing. The method, which discretizes both space and velocity into a lattice structure, has been the subject of highly sophisticated implementation by academia and industry, resulting in very large and engineered code bases. This article introduces miniLB , to the best of our knowledge the first SYCL-based mini-application for the Lattice Boltzmann method. Thanks to its minimalist structure, miniLB is a perfect benchmark to address four key aspects of lattice Boltzmann implementations: (a) GPU acceleration, thanks to an efficient implementation in SYCL capable of abstracting complex fluid dynamics simulations across heterogeneous computing systems; (b) performance portability, with an efficient mapping to SYCL semantics focused on performance portability, evaluated on GPUs from different vendors; (c) mixed precision, with four different variations exploiting combinations of double, single, and half floating-point representations; (d) flexibility, demonstrated through four different use cases, including Lid-driven cavity, Von Karmann street, Poiseuille flow, and Taylor–Green vortex. The results of miniLB , compared to a manually tuned FORTRAN version, demonstrate the effectiveness of miniLB in assessing performance portability across different hardware, while providing valuable insights for optimizing large-scale lattice Boltzmann simulations in modern massively parallel computing systems.
Biagio Cosenza, Luigi Crisci, Giorgio Amati, Matteo Turisini
Future Gener. Comput. Syst.1
2026 A Portable Compiler-Runtime Approach for Scalability Prediction
abstract
Highly scalable parallel applications can efficiently solve expensive computational problems when run on a large number of compute nodes. However, selecting the optimal number of nodes for a compute job of a given size is non-trivial, and allocating too few or too many nodes may not yield the expected performance. Knowing the scaling behavior of an application in advance enables us, for example, to make optimal use of the available hardware resources. We introduce a novel, portable approach to predict the scalability of parallel applications written in modern high-level programming models. We propose a predictive compiler-runtime framework based on Celerity, a task-based distributed runtime system that enables executing SYCL codes on clusters. The framework targets a broad range of computing systems, from CPU to GPU clusters, and proposes a model that combines machine learning, communication modeling and DAG heuristics. Experimental results on two large-scale clusters, JUWELS and Marconi-100, show accurate scalability prediction of unseen single and multi-task applications.
Nicolai Stawinoga, Sohan Lal, Biagio Cosenza, Philip Salzmann, Peter Thoman, Thomas Fahringer
Future Gener. Comput. Syst.3
2025 SYgraph: A Portable Heterogeneous Graph Analytics Framework for GPUs
abstract
Graph analytics play a crucial role in a wide range of fields, including social network analysis, bioinformatics, and scientific computing, due to their ability to model and explore complex relationships. However, optimizing graph algorithms is inherently difficult due to their memory-bound constraints, often resulting in poor performance on modern massively parallel hardware. In addition, most state-of-the-art implementations are designed in CUDA for NVIDIA GPUs, and thus they can not run on supercomputers equipped with AMD and Intel GPUs. To address these challenges, we propose SYgraph, a portable heterogeneous graph analytics framework written in SYCL. SYgraph provides an efficient two-layer bitmap data layout optimized for GPU memory, eliminates the need for pre- or post-processing steps, and abstracts the complexity of working with diverse target platforms. Experimental results demonstrate that SYgraph delivers competitive performance against state-of-the-art frameworks on datasets with up to 21 million nodes and 530 million edges on NVIDIA GPUs while being able to target any SYCL-supported device, such as AMD and Intel GPUs.
Antonio De Caro, Gennaro Cordasco, Biagio Cosenza
ICPP3
2025 SYprox: Combining Host and Device Perforation with Mixed Precision Approximation on Heterogeneous Architectures
abstract
Approximate computing is an emerging paradigm that aims to exploit the inherent error tolerance of many applications, particularly in domains such as image processing and machine learning.Taking advantage of this property, applications can trade off accuracy for significant gains in performance and power consumption.Existing approximation techniques for GPUs are limited to very specific approaches, do not fully exploit the host-device execution model, and are often restricted in terms of programming models and supported target hardware.This paper introduces SYprox, a new approximate computing framework based on SYCL that allows programmers to easily implement heterogeneous approximated applications.SYprox supports multiple techniques, including data perforation, signal reconstruction, and mixed precision, and allows them to be combined to support a wide range of approximations.In particular, SYprox extends existing perforation approaches to allow both host and device data perforation.Experimental results show that SYprox's approximations are Pareto dominant with respect to state-of-the-art approaches and are portable to AMD, Intel and NVIDIA GPUs.
Lorenzo Carpentieri, Biagio Cosenza
ICS2
2025 Phase-Based Frequency Scaling for Energy-Efficient Heterogeneous Computing
abstract
Energy efficiency has been a major challenge for exascale computing. Frequency scaling is a powerful technique to achieve energy savings in modern heterogeneous systems, and can be applied either at a coarse granularity, by application, or at a fine granularity, by setting the frequency for each computational kernel. The chosen granularity significantly impacts the performance and energy consumption of applications due to frequency-change overhead. We propose a novel phase-based method that minimizes the frequency-change overhead and improves performance and energy efficiency on heterogeneous multi-GPU systems. Our approach detects different phases through application profiling and DAG analysis, and sets an optimal frequency for each phase. Our methodology also considers MPI programs, where the overhead can be hidden by overlapping frequency-change with communication. Experimental results show up to 37 % energy saving and$1.87 \times$speedup for various benchmarks on a single GPU, and 68 % energy saving and$3.63 \times$speedup on two multiGPU applications.
Lorenzo Carpentieri, Antonio De Caro, Majid Salimi Beni, Kaijie Fan, Biagio Cosenza
IPDPS5
2025 A Performance Analysis of Autovectorization on RVV RISC-V Boards
Lorenzo Carpentieri, Mohammad VazirPanah, Biagio Cosenza
PDP3
2025 SIGMo: High-Throughput Batched Subgraph Isomorphism on GPUs for Molecular Matching
abstract
Subgraph isomorphism is a fundamental graph problem with applications in diverse domains from biology to social network analysis. Of particular interest is molecular matching, which uses a subgraph isomorphism formulation for the drug discovery process. While subgraph isomorphism is known to be NP-complete and computationally expensive, in the molecular matching formulation a number of domain constraints allow for efficient implementations. This paper presents SIGMo, a high-throughput, portable subgraph isomorphism framework for GPUs, specifically designed for batch molecular matching. SIGMo takes advantage of the specific domain formulation to provide a more efficient filter-and-join strategy: the framework introduces a novel multi-level iterative filtering technique based on neighborhood signature encoding to efficiently prune candidates prior to a GPU-optimized join phase using a stack-based DFS traversal. The GPU implementation is written in SYCL, allowing portable execution on AMD, Intel, and NVIDIA GPUs. Our experimental evaluation on a large dataset from ZINC demonstrates up to 1470 × speedup over state-of-the-art subgraph isomorphism frameworks, and achieves a throughput of 7.7 billion matches per second on a cluster with 256 GPUs.
Antonio De Caro, Gennaro Cordasco, Federico Ficarelli, Biagio Cosenza
SC4
2024 MPI Collective Algorithm Selection in the Presence of Process Arrival Patterns
abstract
The Message Passing Interface (MPI) is a programming model for developing high-performance applications on large-scale machines. A key component of MPI is its collective communication operations. While the MPI standard defines the semantics of these operations, it leaves the algorithmic implementation to the MPI libraries. Each MPI library contains various algorithms for each collective, and selecting the best algorithm typically relies on performance metrics obtained from micro-benchmarks. In such micro-benchmarks, processes are typically synchronized using an MPI_Barrier before invoking a collective operation. However, in real-world scenarios, processes often arrive at a collective in diverse patterns, often due to resource contention. The performance of collective algorithms can vary significantly depending on the arrival pattern type. In this work, we address the challenge of selecting the most efficient algorithm for a given collective, taking into account process arrival patterns. First, we demonstrate through a simulation study that arrival patterns significantly influence the choice of the optimal collective algorithm for specific communication instances. Second, we conduct a comprehensive micro-benchmark analysis to illustrate the sensitivity of MPI collectives to these arrival patterns. Third, we show that our innovative micro-benchmarking methodology is effective in selecting the best-performing collective algorithm for real-world applications.
Majid Salimi Beni, Biagio Cosenza, Sascha Hunold
CLUSTER2
2024 Enabling performance portability on the LiGen drug discovery pipeline
abstract
In recent years, there has been a growing interest in developing high-performance implementations of drug discovery processing software. To target modern GPU architectures, such applications are mostly written in proprietary languages such as CUDA or HIP. However, with the increasing heterogeneity of modern HPC systems and the availability of accelerators from multiple hardware vendors, it has become critical to be able to efficiently execute drug discovery pipelines on multiple large-scale computing systems, with the ultimate goal of working on urgent computing scenarios. This article presents the challenges of migrating LiGen, an industrial drug discovery software pipeline, from CUDA to the SYCL programming model, an industry standard based on C++ that enables heterogeneous computing. We perform a structured analysis of the performance portability of the SYCL LiGen platform, focusing on different aspects of the approach from different perspectives. First, we analyze the performance portability provided by the high-level semantics of SYCL, including the most recent group algorithms and subgroups of SYCL 2020. Second, we analyze how low-level aspects such as kernel occupancy and register pressure affect the performance portability of the overall application. The experimental evaluation is performed on two different versions of LiGen, implementing two different parallelization patterns, by comparing them with a manually optimized CUDA version, and by evaluating performance portability using both known and ad hoc metrics. The results show that, thanks to the combination of high-level SYCL semantics and some manual tuning, LiGen achieves native-comparable performance on NVIDIA, while also running on AMD GPUs.
Luigi Crisci, Lorenzo Carpentieri, Biagio Cosenza, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Gianluca Palermo, Andrea Beccari
Future Gener. Comput. Syst.3
2024 Out of kernel tuning and optimizations for portable large-scale docking experiments on GPUs
abstract
Abstract Virtual screening is an early stage in the drug discovery process that selects the most promising candidates. In the urgent computing scenario, finding a solution in the shortest time frame is critical. Any improvement in the performance of a virtual screening application translates into an increase in the number of candidates evaluated, thereby raising the probability of finding a drug. In this paper, we show how we can improve application throughput using Out-of-kernel optimizations. They use input features, kernel requirements, and architectural features to rearrange the kernel inputs, executing them out of order, to improve the computation efficiency. These optimizations’ implementations are designed on an extreme-scale virtual screening application, named LiGen, that can hinge on CUDA and SYCL kernels to carry out the computation on modern supercomputer nodes. Even if they are tailored to a single application, they might also be of interest for applications that share a similar design pattern. The experimental results show how these optimizations can increase kernel performance by 2 $$\times$$ × , respectively, up to 2.2 $$\times$$ × in CUDA and up to 1.9 $$\times$$ × , in SYCL. Moreover, the reported speedup can be achieved with the best-proposed parameterization, as shown by the data we collected and reported in this manuscript.
Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Luigi Crisci, Biagio Cosenza, Andrea Beccari, Gianluca Palermo
J. Supercomput.5
2024 Analysis and prediction of performance variability in large-scale computing systems
abstract
Abstract The development of new exascale supercomputers has dramatically increased the need for fast, high-performance networking technology. Efficient network topologies, such as Dragonfly+, have been introduced to meet the demands of data-intensive applications and to match the massive computing power of GPUs and accelerators. However, these supercomputers still face performance variability mainly caused by the network that affects system and application performance. This study comprehensively analyzes performance variability on a large-scale HPC system with Dragonfly+ network topology, focusing on factors such as communication patterns, message size, job placement locality, MPI collective algorithms, and overall system workload. The study also proposes an easy-to-measure metric for estimating network background traffic generated by other users, which can be used to estimate the performance of our job accurately. The insights gained from this study contribute to improving performance predictability, enhancing job placement policies and MPI algorithm selection, and optimizing resource management strategies in supercomputers.
Majid Salimi Beni, Sascha Hunold, Biagio Cosenza
J. Supercomput.3
2023 EMPI: Enhanced Message Passing Interface in Modern C++
abstract
Message Passing Interface (MPI) is a well-known standard for programming distributed and HPC systems. While the community has been continuously improving MPI to address the requirements of next-generation architectures and applications, its interface has not substantially evolved. In fact, MPI only provides an interface to C and Fortran and does not support recent features of modern C++. Moreover, MPI programs are error-prone and subject to different syntactic and semantic errors. This paper introduces EMPI, an Enhanced Message Passing Interface based on modern C++, which is directly mapped to the OpenMPI implementation and exploits modern C++ for safe and efficient distributed programming. EMPI proposes novel C++RAII-based semantics and constant specialization to prevent error-prone code patterns such as parameter mismatch, and reduce the overhead of handling multiple objects and perinvocation time. Consequently, EMPI programs are safer: six out of nine well-known MPI error patterns do not occur while correctly using EMPI semantics. Experimental results on five microbenchmarks and two applications on a large-scale cluster using up to 1024 processes show that EMPI's performance is very similar to native MPI and considerably faster than the MPL C++ interface.
Majid Salimi Beni, Luigi Crisci, Biagio Cosenza
CCGrid3
2023 An Asynchronous Dataflow-Driven Execution Model For Distributed Accelerator Computing
abstract
While domain-specific HPC software packages continue to thrive and are vital to many scientific communities, a general purpose high-productivity GPU cluster programming model that facilitates experimentation for non-experts remains elusive. We demonstrate how Celerity, a high-level C++ programming model for distributed accelerator computing based on the open SYCL standard, allows for the quick development of - and experimentation with - distributed applications. To achieve scalability on large machines, we replace Celerity's existing master/worker scheduling model with a fully distributed scheme that reduces the worst-case scheduling complexity from quadratic to linear while maintaining the existing programming interface. We then show how this declarative, data-flow based API paired with a point-to-point communication model with eager data pushing can effectively expose and leverage opportunities for latency hiding and computation/communication overlapping with minimal or no manual guidance. We demonstrate how Celerity exhibits very good scalability on multiple benchmarks from several scientific domains and up to 128 GPUs.
Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Biagio Cosenza, Thomas Fahringer
CCGrid5
2023 Tunable and Portable Extreme-Scale Drug Discovery Platform at Exascale: the LIGATE Approach
abstract
Today digital revolution is having a dramatic impact on the pharmaceutical industry and the entire healthcare system. The implementation of machine learning, extreme-scale computer simulations, and big data analytics in the drug design and development process offers an excellent opportunity to lower the risk of investment and reduce the time to the patient.
Gianluca Palermo, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Cristina Silvano, Bruno Guindani, Danilo Ardagna, Andrea Beccari, Domenico Bonanni, Carmine Talarico, Filippo Lunghini, Jan Martinovic, Paulo Silva 0002, Ada Böhm, Jakub Beránek, Jan Krenek, Branislav Jansik, Biagio Cosenza, Luigi Crisci, Peter Thoman, Philip Salzmann, Thomas Fahringer, Leila Tamara Alexander, Gerardo Tauriello, Torsten Schwede, Janani Durairaj, Andrew Emerson, Federico Ficarelli, Sebastian Wingbermühle, Erik Lindahl, Daniele Gregori, Emanuele Sana, Silvano Coletti, Philipp Gschwandtner
CF18
2023 SYnergy: Fine-grained Energy-Efficient Heterogeneous Computing for Scalable Energy Saving
abstract
Energy-efficient computing uses power management techniques such as frequency scaling to save energy. Implementing energy-efficient techniques on large-scale computing systems is challenging for several reasons. While most modern architectures, including GPUs, are capable of frequency scaling, these features are often not available on large systems. In addition, achieving higher energy savings requires precise energy tuning because not only applications but also different kernels can have different energy characteristics. We propose SYnergy, a novel energy-efficient approach that spans languages, compilers, runtimes, and job schedulers to achieve unprecedented fine-grained energy savings on large-scale heterogeneous clusters. SYnergy defines an extension to the SYCL programming model that allows programmers to define a specific energy goal for each kernel. For example, a kernel can aim to minimize well-known energy metrics such as EDP and ED2P or to achieve predefined energy-performance tradeoffs, such as the best performance with 25% energy savings. Through compiler integration and a machine learning model, each kernel is statically optimized for the specific target. On large computing systems, a SLURM plug-in allows SYnergy to run on all available devices in the cluster, providing scalable energy savings. The methodology is inherently portable and has been evaluated on both NVIDIA and AMD GPUs. Experimental results show unprecedented improvements in energy and energy-related metrics on real-world applications, as well as scalable energy savings on a 64-GPU cluster.
Kaijie Fan, Marco D'Antonio, Lorenzo Carpentieri, Biagio Cosenza, Federico Ficarelli, Daniele Cesarini
SC4
2022 FLEXDP: flexible frequency scaling for energy-delay product optimization of GPU applications
abstract
Dynamic frequency scaling is broadly available among different modern computer architectures, making it possible to improve the performance and energy efficiency of an application by carefully setting the core frequency. However, while an exhaustive tuning is feasible on simple single-kernel applications, in real-world applications comprised of multiple tasks, the set of possible frequency setting combinations is too large to be exhaustively evaluated.
Kaijie Fan, Biagio Cosenza, Ben H. H. Juurlink
CF2
2022 An Analysis of Performance Variability on Dragonfly+topology
abstract
Large-scale compute clusters are highly affected by performance variability that originates from different sources. Among these sources, the network plays an essential role as a shared resource between users and their jobs in a supercomputer. In this paper, we analyze the effect of some network-related sources on the performance variability of a modern compute cluster equipped with a Dragonfly+ interconnect. Specifically, we focus on the impacts of job placement, communication patterns, routing strategy, and network background traffic on the performance variability of communication-intensive workloads. To quantify the effect of network congestion (background traffic) on the performance variability, we propose a heuristic that can successfully estimate the amount of communication on the network produced by other jobs running on the cluster simultaneously. Then, we show how this network congestion contributes to the performance variability of different communication patterns and real-world communication-intensive applications.
Majid Salimi Beni, Biagio Cosenza
CLUSTER2
2021 The Italian research on HPC key technologies across EuroHPC
abstract
High-Performance Computing (HPC) is one of the strategic priorities for research and innovation worldwide due to its relevance for industrial and scientific applications. We envision HPC as composed of three pillars: infrastructures, applications, and key technologies and tools. While infrastructures are by construction centralized in large-scale HPC centers, and applications are generally within the purview of domain-specific organizations, key technologies fall in an intermediate case where coordination is needed, but design and development are often decentralized. A large group of Italian researchers has started a dedicated laboratory within the National Interuniversity Consortium for Informatics (CINI) to address this challenge. The laboratory, albeit young, has managed to succeed in its first attempts to propose a coordinated approach to HPC research within the EuroHPC Joint Undertaking, participating in the calls 2019--20 to five successful proposals for an aggregate total cost of 95M€. In this paper, we outline the working group's scope and goals and provide an overview of the five funded projects, which become fully operational in March 2021, and cover a selection of key technologies provided by the working group partners, highlighting their usage development within the projects.
Marco Aldinucci, Giovanni Agosta, Antonio Andreini, Claudio A. Ardagna, Andrea Bartolini, Alessandro Cilardo, Biagio Cosenza, Marco Danelutto, Roberto Esposito, William Fornaciari, Roberto Giorgi, Davide Lengani, Raffaele Montella, Mauro Olivieri, Sergio Saponara, Daniele Simoni, Massimo Torquati
CF7
2021 ALONA: Automatic Loop Nest Approximation with Reconstruction and Space Pruning
Daniel Maier 0002, Biagio Cosenza, Ben H. H. Juurlink
Euro-Par2
2021 Easy and efficient agent-based simulations with the OpenABL language and compiler
Biagio Cosenza, Nikita Popov, Ben H. H. Juurlink, Paul Richmond, Mozhgan Chimeh, Carmine Spagnuolo, Gennaro Cordasco, Vittorio Scarano
Future Gener. Comput. Syst.1
2020 SYCL-Bench: A Versatile Cross-Platform Benchmark Suite for Heterogeneous Computing
Sohan Lal, Aksel Alpay, Philip Salzmann, Biagio Cosenza, Alexander Hirsch, Nicolai Stawinoga, Peter Thoman, Thomas Fahringer, Vincent Heuveline
Euro-Par4
2020 Vectorization cost modeling for NEON, AVX and SVE
Angela Pohl, Biagio Cosenza, Ben H. H. Juurlink
Perform. Evaluation2
2019 Celerity: High-Level C++ for Accelerator Clusters
Peter Thoman, Philip Salzmann, Biagio Cosenza, Thomas Fahringer
Euro-Par3
2019 Predictable GPUs Frequency Scaling for Energy and Performance
abstract
Dynamic voltage and frequency scaling (DVFS) is an important solution to balance performance and energy consumption, and hardware vendors provide management libraries that allow the programmer to change both memory and core frequencies. The possibility to manually set these frequencies is a great opportunity for application tuning, which can focus on the best application-dependent setting. However, this task is not straightforward because of the large set of possible configurations and because of the multi-objective nature of the problem, which minimizes energy consumption and maximizes performance.
Kaijie Fan, Biagio Cosenza, Ben H. H. Juurlink
ICPP2
2019 Portable Cost Modeling for Auto-Vectorizers
abstract
Compiler optimization passes employ cost models to determine if a code transformation will yield performance improvements. When this assessment is inaccurate, compilers apply transformations that are not beneficial, or refrain from applying ones that would have improved the code. We analyze the accuracy of the cost models used in LLVM's and GCC's vectorization passes for two different instruction set architectures. In general, speedup is over-estimated, resulting in mispredictions and a weak to medium correlation between predicted and actual performance gain. We therefore propose a novel cost model that is based on a code's intermediate representation with refined memory access pattern features. Using linear regression techniques, this platform independent model is fitted to an AVX2 and a NEON hardware. Results show that the fitted model significantly improves the correlation between predicted and measured speedup (AVX2: +52% for training data, +13% for validation data), as well as the number of mispredictions (NEON: -15 for training data, -12 for validation data) for more than 80 code patterns.
Angela Pohl, Biagio Cosenza, Ben H. H. Juurlink
MASCOTS2
2018 Local memory-aware kernel perforation
abstract
Many applications provide inherent resilience to some amount of error and can potentially trade accuracy for performance by using approximate computing. Applications running on GPUs often use local memory to minimize the number of global memory accesses and to speed up execution. Local memory can also be very useful to improve the way approximate computation is performed, e.g., by improving the quality of approximation with data reconstruction techniques. This paper introduces local memory-aware perforation techniques specifically designed for the acceleration and approximation of GPU kernels. We propose a local memory-aware kernel perforation technique that first skips the loading of parts of the input data from global memory, and later uses reconstruction techniques on local memory to reach higher accuracy while having performance similar to state-of-the-art techniques. Experiments show that our approach is able to accelerate the execution of a variety of applications from 1.6× to 3× while introducing an average error of 6%, which is much smaller than that of other approaches. Results further show how much the error depends on the input data and application scenario, the impact of local memory tuning and different parameter configurations.
Daniel Maier 0002, Biagio Cosenza, Ben H. H. Juurlink
CGO2
2018 Cost Modelling for Vectorization on ARM
abstract
When applying a code transformation to optimize for performance, compilers need to assess its profitability beforehand. For this purpose, they utilize cost models, which compare the cost, an abstract measure of the code, before and after the transformation. If the cost is lower after the transformation, it will be applied. Exact cost modelling is therefore critical to avoid slowdowns or missed opportunities for speedups. In this work, we analyze the accuracy of LLVM's loop-level vectorization (LLV) cost model, and show the benefit of modelling speedup instead of instruction costs for higher vectorization rates and smaller execution times. The presented approach is portable to other compilers and hardwares as well.
Angela Pohl, Biagio Cosenza, Ben H. H. Juurlink
CLUSTER2
2018 OpenABL: A Domain-Specific Language for Parallel and Distributed Agent-Based Simulations
Biagio Cosenza, Nikita Popov, Ben H. H. Juurlink, Paul Richmond, Mozhgan Chimeh, Carmine Spagnuolo, Gennaro Cordasco, Vittorio Scarano
Euro-Par1
2018 Accelerating the RICH Particle Detector Algorithm on Intel Xeon Phi
abstract
At the LHC, particles are collided in order to understand how the universe was created. Those collisions are called events and generate large quantities of data, which have to be pre-filtered before they are stored to hard disks. This paper presents a parallel implementation of these algorithms that is specifically designed for the Intel Xeon Phi Knights Landing platform, exploiting its 64 cores and AVX-512 instruction set. It shows that a linear speedup up until approximately 64 threads is attainable when vectorization is used, data is aligned to cache line boundaries, program execution is pinned to MCDRAM, mathematical expressions are transformed to a more efficient equivalent formulation, and OpenMP is used for parallelization. The code was transformed from being compute bound to memory bound. Overall, a speedup of 36.47x was reached while obtaining an error which is smaller than the detector resolution.
Christina Quast, Angela Pohl, Biagio Cosenza, Ben H. H. Juurlink, Rainer Schwemmer
PDP3
2018 Control Flow Vectorization for ARM NEON
abstract
Single Instruction Multiple Data (SIMD) extensions in processors enable in-core parallelism for operations on vectors of data. From the compiler perspective, SIMD instructions require automatic techniques to determine how and when it is possible to express computations in terms of vector operations. When this is not possible automatically, a user may still write code in a manner that allows the compiler to deduce that vectorization is possible, or by explicitly define how to vectorize by using intrinsics.
Angela Pohl, Biagio Cosenza, Ben H. H. Juurlink
SCOPES2
2017 Static optimization in PHP 7
Nikita Popov, Biagio Cosenza, Ben H. H. Juurlink, Dmitry Stogov
CC2
2017 Autotuning Stencil Computations with Structural Ordinal Regression Learning
abstract
Stencil computations expose a large and complex space of equivalent implementations. These computations often rely on autotuning techniques, based on iterative compilation or machine learning (ML), to achieve high performance. Iterative compilation autotuning is a challenging and time-consuming task that may be unaffordable in many scenarios. Meanwhile, traditional ML autotuning approaches exploiting classification algorithms (such as neural networks and support vector machines) face difficulties in capturing all features of large search spaces. This paper proposes a new way of automatically tuning stencil computations based on structural learning. By organizing the training data in a set of partially-sorted samples (i.e., rankings), the problem is formulated as a ranking prediction model, which translates to an ordinal regression problem. Our approach can be coupled with an iterative compilation method or used as a standalone autotuner. We demonstrate its potential by comparing it with state-of-the-art iterative compilation methods on a set of nine stencil codes and by analyzing the quality of the obtained ranking in terms of Kendall rank correlation coefficients.
Biagio Cosenza, Juan José Durillo, Stefano Ermon, Ben H. H. Juurlink
IPDPS1
2017 Stencil Autotuning with Ordinal Regression: Extended Abstract
abstract
The increasing performance of today's computer architecture comes with an unprecedented augment of hardware complexity. Unfortunately this results in difficult-to-tune software and consequentially in a gap between the potential peak performance and the actual performance. Automatic tuning is an emerging approach that assists the programmer in managing this complexity. State-of-the-art autotuners are limited, though: they either require long tuning times, e.g., due to iterative searches, or cannot tackle the complexity of the problem due to the limitation of the supervised machine learning (ML) methodologies used. In particular, traditional ML autotuning approaches exploiting classification algorithms (such as neural networks and support vector machines) face difficulties in capturing all features of large search spaces. We propose a new way of performing automatic tuning based on structural learning: the tuning problem is formulated as a version ranking prediction modeling and solved using ordinal regression. We demonstrate its potential on a well-known autotuning problem: stencil computations. We compare state-of-the-art iterative compilation methods with our ordinal regression approach and analyze the quality of the obtained ranking in terms of Kendall rank correlation coefficients.
Biagio Cosenza, Juan José Durillo, Stefano Ermon, Ben H. H. Juurlink
SCOPES1
2015 Automatic Data Layout Optimizations for GPUs
Klaus Kofler, Biagio Cosenza, Thomas Fahringer
Euro-Par2
2015 Spectral turning bands for efficient Gaussian random fields generation on GPUs and accelerators
abstract
Summary A random field (RF) is a set of correlated random variables associated with different spatial locations. RF generation algorithms are of crucial importance for many scientific areas, such as astrophysics, geostatistics, computer graphics, and many others. Current approaches commonly make use of 3D fast Fourier transform (FFT), which does not scale well for RF bigger than the available memory; they are also limited to regular rectilinear meshes. We introduce random field generation with the turning band method (RAFT), an RF generation algorithm based on the turning band method that is optimized for massively parallel hardware such as GPUs and accelerators. Our algorithm replaces the 3D FFT with a lower‐order, one‐dimensional FFT followed by a projection step and is further optimized with loop unrolling and blocking. RAFT can easily generate RF on non‐regular (non‐uniform) meshes and efficiently produce fields with mesh sizes bigger than the available device memory by using a streaming, out‐of‐core approach. Our algorithm generates RF with the correct statistical behavior and is tested on a variety of modern hardware, such as NVIDIA Tesla, AMD FirePro and Intel Phi. RAFT is faster than the traditional methods on regular meshes and has been successfully applied to two real case scenarios: planetary nebulae and cosmological simulations. Copyright © 2015 John Wiley & Sons, Ltd.
Lars Hunger, Biagio Cosenza, Stefan Kimeswenger, Thomas Fahringer
Concurr. Comput. Pract. Exp.2
2014 Random Fields Generation on the GPU with the Spectral Turning Bands Method
Lars Hunger, Biagio Cosenza, Stefan Kimeswenger, Thomas Fahringer
Euro-Par2
2014 A uniform approach for programming distributed heterogeneous computing systems
abstract
Large-scale compute clusters of heterogeneous nodes equipped with multi-core CPUs and GPUs are getting increasingly popular in the scientific community. However, such systems require a combination of different programming paradigms making application development very challenging. In this article we introduce libWater, a library-based extension of the OpenCL programming model that simplifies the development of heterogeneous distributed applications. libWater consists of a simple interface, which is a transparent abstraction of the underlying distributed architecture, offering advanced features such as inter-context and inter-node device synchronization. It provides a runtime system which tracks dependency information enforced by event synchronization to dynamically build a DAG of commands, on which we automatically apply two optimizations: collective communication pattern detection and device-host-device copy removal. We assess libWater's performance in three compute clusters available from the Vienna Scientific Cluster, the Barcelona Supercomputing Center and the University of Innsbruck, demonstrating improved performance and scaling with different test applications and configurations.
Ivan Grasso, Simone Pellegrini, Biagio Cosenza, Thomas Fahringer
J. Parallel Distributed Comput.3
2013 LibWater: heterogeneous distributed computing made easy
abstract
Clusters of heterogeneous nodes composed of multi-core CPUs and GPUs are increasingly being used for High Performance Computing (HPC) due to the benefits in peak performance and energy efficiency. In order to fully harvest the computational capabilities of such architectures, application developers often employ a combination of different parallel programming paradigms (e.g. OpenCL, CUDA, MPI and OpenMP), also known in literature as hybrid programming, which makes application development very challenging. Furthermore, these languages offer limited support to orchestrate data and computations for heterogeneous systems.
Ivan Grasso, Simone Pellegrini, Biagio Cosenza, Thomas Fahringer
ICS3
2013 An automatic input-sensitive approach for heterogeneous task partitioning
abstract
Unleashing the full potential of heterogeneous systems, consisting of multi-core CPUs and GPUs, is a challenging task due to the difference in processing capabilities, memory availability, and communication latencies of different computational resources.
Klaus Kofler, Ivan Grasso, Biagio Cosenza, Thomas Fahringer
ICS3
2013 Automatic problem size sensitive task partitioning on heterogeneous parallel systems
abstract
In this paper we propose a novel approach which automatizes task partitioning in heterogeneous systems. Our framework is based on the Insieme Compiler and Runtime infrastructure. The compiler translates a single-device OpenCL program into a multi-device OpenCL program. The runtime system then performs dynamic task partitioning based on an offline-generated prediction model. In order to derive the prediction model, we use a machine learning approach that incorporates static program features as well as dynamic, input sensitive features. Our approach has been evaluated over a suite of 23 programs and achieves performance improvements compared to an execution of the benchmarks on a single CPU and a single GPU only.
Ivan Grasso, Klaus Kofler, Biagio Cosenza, Thomas Fahringer
PPoPP3
2011 Distributed Load Balancing for Parallel Agent-Based Simulations
abstract
We focus on agent-based simulations where a large number of agents move in the space, obeying to some simple rules. Since such kind of simulations are computational intensive, it is challenging, for such a contest, to let the number of agents to grow and to increase the quality of the simulation. A fascinating way to answer to this need is by exploiting parallel architectures. In this paper, we present a novel distributed load balancing schema for a parallel implementation of such simulations. The purpose of such schema is to achieve an high scalability. Our approach to load balancing is designed to be lightweight and totally distributed: the calculations for the balancing take place at each computational step, and influences the successive step. To the best of our knowledge, our approach is the first distributed load balancing schema in this context. We present both the design and the implementation that allowed us to perform a number of experiments, with up-to 1,000,000 agents. Tests show that, in spite of the fact that the load balancing algorithm is local, the workload distribution is balanced while the communication overhead is negligible.
Biagio Cosenza, Gennaro Cordasco, Rosario De Chiara, Vittorio Scarano
PDP1
2008 Load Balancing in Mesh-like Computations using Prediction Binary Trees
abstract
We present a load-balancing technique that exploits the temporal coherence, among successive computation phases, in mesh-like computations to be mapped on a cluster of processors. Our method partitions the computation in balanced tasks and distributes them to independent processors through the prediction binary tree (PBT). At each new phase, current PBT is updated by using previous phase computing time (for each task) as (next phase) cost estimate. The PBT is designed so that it balances the load across the tasks as well as reduce {\em dependency} among processors for higher performances. Reducing dependency is obtained by using rectangular tiles of the mesh, of almost-square shape (i.e. one dimension is at most twice the other). By reducing dependency, one can reduce inter-processors communication or exploit local dependencies among tasks (such as data locality).Our strategy has been assessed on a significant problem, parallel ray tracing. Our implementation shows a good scalability, and improves over coherence-oblivious implementations. We report different measurements showing that granularity of tasks is a key point for the performances of our decomposition/mapping strategy.
Biagio Cosenza, Gennaro Cordasco, Rosario De Chiara, Ugo Erra, Vittorio Scarano
ISPDC1