Paolo Viviani 0001

dblp:55/10024-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-8947-9481ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications
Marco Cipollini, Simone Rizzo, Sergio Iserte, Paolo Viviani 0001, Giacomo Vitali, Matteo Barbieri, Gabriella Bettonte, Elisabetta Boella, Fulvio Ganz, Roberto Rocco, Orazio Spina, Antonio J. Peña, Petter Sandås, Iacopo Colonnelli, Alberto Scionti, Chiara Vercellino, Emanuele Dri, Jonathan Frassineti, Sara Marzella, Andrea Muratori, Daniele Ottaviani, Olivier Terzo, Bartolomeo Montrucchio, Daniele Gregori
Future Gener. Comput. Syst.4
2023 Neural optimization for quantum architectures: graph embedding problems with Distance Encoder Networks
abstract
Quantum machines are among the most promising technologies expected to provide significant improvements in the following years. However, bridging the gap between real-world applications and their implementation on quantum hardware is still a complicated task. One of the main challenges is to represent through qubits (i.e., the basic units of quantum information) the problems of interest. According to the specific technology under-lying the quantum machine, it is necessary to implement a proper representation strategy, generally referred to as embedding. This paper introduces a neural-enhanced optimization framework to solve the constrained unit disk problem, which arises in the context of qubits positioning for neutral atoms-based quantum hardware. The proposed approach involves a modified autoencoder model, i.e., the Distances Encoder Network, and a custom loss, i.e., the Embedding Loss Function, respectively, to compute Euclidean distances and model the optimization constraints. The core idea behind this design relies on the capability of neural networks to approximate non-linear transformations to make the Distances Encoder Network learn the spatial transformation that maps initial non-feasible solutions of the constrained unit disk problem into feasible ones. The proposed approach outperforms classical solvers, given fixed comparable computation times, and paves the way to address other optimization problems through a similar strategy.
Chiara Vercellino, Giacomo Vitali, Paolo Viviani 0001, Alberto Scionti, Andrea Scarabosio, Olivier Terzo, Edoardo Giusto, Bartolomeo Montrucchio
COMPSAC3
2023 A Machine Learning Approach for an HPC Use Case: the Jobs Queuing Time Prediction
abstract
High-Performance Computing (HPC) domain provided the necessary tools to support the scientific and industrial advancements we all have seen during the last decades. HPC is a broad domain targeting to provide both software and hardware solutions as well as envisioning methodologies that allow achieving goals of interest, such as system performance and energy efficiency. In this context, supercomputers have been the vehicle for developing and testing the most advanced technologies since their first appearance. Unlike cloud computing resources that are provided to the end-users in an on-demand fashion in the form of virtualized resources (i.e., virtual machines and containers), supercomputers’ resources are generally served through State-of-the-Art batch schedulers (e.g., SLURM, PBS, LSF, HTCondor). As such, the users submit their computational jobs to the system, which manages their execution with the support of queues. In this regard, predicting the behaviour of the jobs in the batch scheduler queues becomes worth it. Indeed, there are many cases where a deeper knowledge of the time experienced by a job in a queue (e.g., the submission of check-pointed jobs or the submission of jobs with execution dependencies) allows exploring more effective workflow orchestration policies. In this work, we focused on applying machine learning (ML) techniques to learn from the historical data collected from the queuing system of real supercomputers, aiming at predicting the time spent on a queue by a given job. Specifically, we applied both unsupervised learning (UL) and supervised learning (SL) techniques to define the most effective features for the prediction task and the actual prediction of the queue waiting time. For this purpose, two approaches have been explored: on one side, the prediction of ranges on jobs’ queuing times (classification approach) and, on the other side, the prediction of the waiting time at the minutes level (regression approach). Experimental results highlight the strong relationship between the SL models’ performances and the way the dataset is split. At the end of the prediction step, we present the uncertainty quantification approach, i.e., a tool to associate the predictions with reliability metrics, based on variance estimation.
Chiara Vercellino, Alberto Scionti, Giuseppe Varavallo, Paolo Viviani 0001, Giacomo Vitali, Olivier Terzo
Future Gener. Comput. Syst.4
2023 Accelerating legacy applications with spatial computing devices
abstract
Abstract Heterogeneous computing is the major driving factor in designing new energy-efficient high-performance computing systems. Despite the broad adoption of GPUs and other specialized architectures, the interest in spatial architectures like field-programmable gate arrays (FPGAs) has grown. While combining high performance, low power consumption and high adaptability constitute an advantage, these devices still suffer from a weak software ecosystem, which forces application developers to use tools requiring deep knowledge of the underlying system, often leaving legacy code (e.g., Fortran applications) unsupported. By realizing this, we describe a methodology for porting Fortran (legacy) code on modern FPGA architectures, with the target of preserving performance/power ratios. Aimed as an experience report, we considered an industrial computational fluid dynamics application to demonstrate that our methodology produces synthesizable OpenCL codes targeting Intel Arria10 and Stratix10 devices. Although performance gain is not far beyond that of the original CPU code (we obtained a relative speedup of $$\times$$ × 0.59 and $$\times$$ × 0.63, respectively, for a single optimized main kernel, while only on the Stratix10 we achieved $$\times$$ × 2.56 by replicating the main optimized kernel 4 times), our results are quite encouraging to drawn the path for further investigations. This paper also reports some major criticalities in porting Fortran code on FPGA architectures.
Paolo Savio, Alberto Scionti, Giacomo Vitali, Paolo Viviani 0001, Chiara Vercellino, Olivier Terzo, Huy-Nam Nguyen, Donato Magarielli, Ennio Spano, Michele Marconcini, Francesco Poli
J. Supercomput.4
2022 Dynamic Job Allocation on Federated Cloud-HPC Environments
Giacomo Vitali, Alberto Scionti, Paolo Viviani 0001, Chiara Vercellino, Olivier Terzo
CISIS3
2022 Taming Multi-node Accelerated Analytics: An Experience in Porting MATLAB to Scale with Python
Paolo Viviani 0001, Giacomo Vitali, Davide Lengani, Alberto Scionti, Chiara Vercellino, Olivier Terzo
CISIS1
2019 Accelerating Spectral Graph Analysis Through Wavefronts of Linear Algebra Operations
abstract
The wavefront pattern captures the unfolding of a parallel computation in which data elements are laid out as a logical multidimensional grid and the dependency graph favours a diagonal sweep across the grid. In the emerging area of spectral graph analysis, the computing often consists in a wavefront running over a tiled matrix, involving expensive linear algebra kernels. While these applications might benefit from parallel heterogeneous platforms (multi-core with GPUs), programming wavefront applications directly with high-performance linear algebra libraries yields code that is complex to write and optimize for the specific application. We advocate a methodology based on two abstractions (linear algebra and parallel pattern-based run-time), that allows to develop portable, self-configuring, and easy-to-profile code on hybrid platforms.
Maurizio Drocco, Paolo Viviani 0001, Iacopo Colonnelli, Marco Aldinucci, Marco Grangetto
PDP2
2019 Deep Learning at Scale
abstract
This work presents a novel approach to distributed training of deep neural networks (DNNs) that aims to overcome the issues related to mainstream approaches to data parallel training. Established techniques for data parallel training are discussed from both a parallel computing and deep learning perspective, then a different approach is presented that is meant to allow DNN training to scale while retaining good convergence properties. Moreover, an experimental implementation is presented as well as some preliminary results.
Paolo Viviani 0001, Maurizio Drocco, Daniele Baccega, Iacopo Colonnelli, Marco Aldinucci
PDP1
2018 HPC4AI: an AI-on-demand federated platform endeavour
abstract
In April 2018, under the auspices of the POR-FESR 2014-2020 program of Italian Piedmont Region, the Turin's Centre on High-Performance Computing for Artificial Intelligence (HPC4AI) was funded with a capital investment of 4.5M€ and it began its deployment. HPC4AI aims to facilitate scientific research and engineering in the areas of Artificial Intelligence and Big Data Analytics. HPC4AI will specifically focus on methods for the on-demand provisioning of AI and BDA Cloud services to the regional and national industrial community, which includes the large regional ecosystem of Small-Medium Enterprises (SMEs) active in many different sectors such as automotive, aerospace, mechatronics, manufacturing, health and agrifood.
Marco Aldinucci, Sergio Rabellino, Marco Pironti, Filippo Spiga, Paolo Viviani 0001, Maurizio Drocco, Marco Guerzoni, Guido Boella, Marco Mellia, Paolo Margara, Idilio Drago, Roberto Marturano, Guido Marchetto, Elio Piccolo, Stefano Bagnasco, Stefano Lusso, Sara Vallero, Giuseppe Attardi, Alex Barchiesi, Alberto Colla, Fulvio Galeazzi
CF5
2018 Scaling Dense Linear Algebra on Multicore and Beyond: A Survey
abstract
The present trend in big-data analytics is to exploit algorithms with (sub-)linear time complexity, in this sense it is usually worth to investigate if the available techniques can be approximated to reach an affordable complexity. However, there are still problems in data science and engineering that involve algorithms with higher time complexity, like matrix inversion or Singular Value Decomposition (SVD). This work presents the results of a survey that reviews a number of tools meant to perform dense linear algebra at “Big Data” scale: namely, the proposed approach aims first to define a feasibility boundary for the problem size of shared-memory matrix factorizations, then to understand whether it is convenient to employ specific tools meant to scale out such dense linear algebra tasks on distributed platforms. The survey will eventually discuss the presented tools from the point of view of domain experts (data scientist, engineers), hence focusing on the trade-off between usability and performance.
Paolo Viviani 0001, Maurizio Drocco, Marco Aldinucci
PDP1