Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Salvatore Filippone

dblp:f/SalvatoreFilippone · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-5859-7538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Theory of computation · 5 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 39% Electronic design automation · 30% Emerging computing paradigms · 30%
Theoretical computer science
1 paper
Algorithms and data structures · 50% Coding theory · 50%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms › approximate and stochastic computing › stochastic computing
random number generation
0.011990
A vectorized long-period shift-register random number generator · SC 1990
High-performance computing
scientific computing systems
0.011990
A vectorized long-period shift-register random number generator · SC 1990
Electronic design automation › circuit simulation › probabilistic simulation
statistical simulation
0.011990
A vectorized long-period shift-register random number generator · SC 1990
Coding theory
linear feedback shift register
0.011990
A vectorized long-period shift-register random number generator · SC 1990
Algorithms and data structures
pseudorandom number generation
0.011990
A vectorized long-period shift-register random number generator · SC 1990
High-performance computing › code optimization
vectorization
0.011990
A vectorized long-period shift-register random number generator · SC 1990

Methods — techniques the papers use, named apart from their topics

vectorization · 0.0statistical testing · 0.0
YearPublicationVenuePosition
2024 Alya toward exascale: algorithmic scalability using PSCToolkit
abstract
Abstract In this paper, we describe an upgrade of the Alya code with up-to-date parallel linear solvers capable of achieving reliability, efficiency and scalability in the computation of the pressure field at each time step of the numerical procedure for solving a Large Eddy Simulation formulation of the incompressible Navier–Stokes equations. We developed a software module in the Alya’s kernel to interface the libraries included in the current version of , a framework for the iterative solution of sparse linear systems, on parallel distributed-memory computers, by Krylov methods coupled to Algebraic MultiGrid preconditioners. The Toolkit has undergone various extensions within the EoCoE-II project with the primary goal of facing the exascale challenge. Results on a realistic benchmark for airflow simulations in wind farm applications show that the solvers significantly outperform the original versions of the Conjugate Gradient method available in the Alya’s kernel in terms of scalability and parallel efficiency and represent a very promising software layer to move the Alya code toward exascale.
Herbert Owen, Oriol Lehmkuhl, Pasqua D'Ambra, Fabio Durastante, Salvatore Filippone
J. Supercomput.5
2023 AMG Preconditioners based on Parallel Hybrid Coarsening and Multi-objective Graph Matching
abstract
We describe preliminary results from a multi-objective graph matching algorithm, in the coarsening step of an aggregation-based Algebraic MultiGrid (AMG) preconditioner, for solving large and sparse linear systems of equations on high-end parallel computers. We have two objectives. First, we wish to improve the convergence behavior of the AMG method when applied to highly anisotropic problems. Second, we wish to extend the parallel package PSCToolkit to exploit multi-threaded parallelism at the node level on multi-core processors. Our matching proposal balances the need to simultaneously compute high weights and large cardinalities by a new formulation of the weighted matching problem combining both these objectives using a parameter$\lambda$. We compute the matching by a parallel$2/3-\varepsilon$-approximation algorithm for maximum weight matchings. Results with the new matching algorithm show that for a suitable choice of the parameter$\lambda$we compute effective preconditioners in the presence of anisotropy, i.e., smaller solve times, setup times, iterations counts, and operator complexity.
Pasqua D'Ambra, Fabio Durastante, S. M. Ferdous, Salvatore Filippone, Mahantesh Halappanavar, Alex Pothen
PDP4
2022 Social ski driver conditional autoregressive-based deep learning classifier for flight delay prediction
abstract
Abstract The importance of robust flight delay prediction has recently increased in the air transportation industry. This industry seeks alternative methods and technologies for more robust flight delay prediction because of its significance for all stakeholders. The most affected are airlines that suffer from monetary and passenger loyalty losses. Several studies have attempted to analysed and solve flight delay prediction problems using machine learning methods. This research proposes a novel alternative method, namely social ski driver conditional autoregressive-based (SSDCA-based) deep learning. Our proposed method combines the Social Ski Driver algorithm with Conditional Autoregressive Value at Risk by Regression Quantiles. We consider the most relevant instances from the training dataset, which are the delayed flights. We applied data transformation to stabilise the data variance using Yeo-Johnson. We then perform the training and testing of our data using deep recurrent neural network (DRNN) and SSDCA-based algorithms. The SSDCA-based optimisation algorithm helped us choose the right network architecture with better accuracy and less error than the existing literature. The results of our proposed SSDCA-based method and existing benchmark methods were compared. The efficiency and computational time of our proposed method are compared against the existing benchmark methods. The SSDCA-based DRNN provides a more accurate flight delay prediction with 0.9361 and 0.9252 accuracy rates on both dataset-1 and dataset-2, respectively. To show the reliability of our method, we compared it with other meta-heuristic approaches. The result is that the SSDCA-based DRNN outperformed all existing benchmark methods tested in our experiment.
Desmond Bala Bisandu, Irene Moulitsas, Salvatore Filippone
Neural Comput. Appl.3
2018 BootCMatch: A Software Package for Bootstrap AMG Based on Graph Weighted Matching
abstract
This article has two main objectives: one is to describe some extensions of an adaptive Algebraic Multigrid (AMG) method of the form previously proposed by the first and third authors, and a second one is to present a new software framework, named BootCMatch , which implements all the components needed to build and apply the described adaptive AMG both as a stand-alone solver and as a preconditioner in a Krylov method. The adaptive AMG presented is meant to handle general symmetric and positive definite (SPD) sparse linear systems, without assuming any a priori information of the problem and its origin; the goal of adaptivity is to achieve a method with a prescribed convergence rate. The presented method exploits a general coarsening process based on aggregation of unknowns, obtained by a maximum weight matching in the adjacency graph of the system matrix. More specifically, a maximum product matching is employed to define an effective smoother subspace (complementary to the coarse space), a process referred to as compatible relaxation, at every level of the recursive two-level hierarchical AMG process. Results on a large variety of test cases and comparisons with related work demonstrate the reliability and efficiency of the method and of the software.
Pasqua D'Ambra, Salvatore Filippone, Panayot S. Vassilevski
ACM Trans. Math. Softw.2
2017 Coarray-based load balancing on heterogeneous and many-core architectures
Valeria Cardellini, Alessandro Fanfarillo, Salvatore Filippone
Parallel Comput.3
2017 Sparse Matrix-Vector Multiplication on GPGPUs
abstract
The multiplication of a sparse matrix by a dense vector (SpMV) is a centerpiece of scientific computing applications: it is the essential kernel for the solution of sparse linear systems and sparse eigenvalue problems by iterative methods. The efficient implementation of the sparse matrix-vector multiplication is therefore crucial and has been the subject of an immense amount of research, with interest renewed with every major new trend in high-performance computing architectures. The introduction of General-Purpose Graphics Processing Units (GPGPUs) is no exception, and many articles have been devoted to this problem. With this article, we provide a review of the techniques for implementing the SpMV kernel on GPGPUs that have appeared in the literature of the last few years. We discuss the issues and tradeoffs that have been encountered by the various researchers, and a list of solutions, organized in categories according to common features. We also provide a performance comparison across different GPGPU models and on a set of test matrices coming from various application domains.
Salvatore Filippone, Valeria Cardellini, Davide Barbieri, Alessandro Fanfarillo
ACM Trans. Math. Softw.1
2014 Coarrays in GNU Fortran
abstract
Coarray Fortran is a set of features of the Fortran 2008 standard which makes Fortran a PGAS language. Currently, the coarray support is provided mainly by commercial compilers like Cray and Intel. In this work we present two coarray implementations on the GNU Fortran compiler. We present a performance comparison between our coarray implementations and those provided by Cray and Intel. Such comparison includes synthetic benchmarks and real, commonly used, scientific applications.
Alessandro Fanfarillo, Tobias Burnus, Valeria Cardellini, Salvatore Filippone, Dan Nagle, Damian W. I. Rouson
PACT4
2012 Object-Oriented Techniques for Sparse Matrix Computations in Fortran 2003
abstract
The efficiency of a sparse linear algebra operation heavily relies on the ability of the sparse matrix storage format to exploit the computing power of the underlying hardware. Since no format is universally better than the others across all possible kinds of operations and computers, sparse linear algebra software packages should provide facilities to easily implement and integrate new storage formats within a sparse linear algebra application without the need to modify it; it should also allow to dynamically change a storage format at run-time depending on the specific operations to be performed. Aiming at these important features, we present an Object Oriented design model for a sparse linear algebra package which relies on Design Patterns. We show that an implementation of our model can be efficiently achieved through some of the unique features of the Fortran 2003 language. Experimental results show that the proposed software infrastructure improves the modularity and ease of use of the code at no performance loss.
Salvatore Filippone, Alfredo Buttari
ACM Trans. Math. Softw.1
2010 MLD2P4: A Package of Parallel Algebraic Multilevel Domain Decomposition Preconditioners in Fortran 95
abstract
Domain decomposition ideas have long been an essential tool for the solution of PDEs on parallel computers. In recent years many research efforts have been focused on recursively employing domain decomposition methods to obtain multilevel preconditioners to be used with Krylov solvers. In this context, we developed MLD2P4 (MultiLevel Domain Decomposition Parallel Preconditioners Package based on PSBLAS), a package of parallel multilevel preconditioners that combines additive Schwarz domain decomposition methods with a smoothed aggregation technique to build a hierarchy of coarse-level corrections in an algebraic way. The design of MLD2P4 was guided by objectives such as extensibility, flexibility, performance, portability, and ease of use. They were achieved by following an object-based approach while using the Fortran 95 language, as well as by employing the PSBLAS library as a basic framework. In this article, we present MLD2P4 focusing on its design principles, software architecture, and use.
Pasqua D'Ambra, Daniela di Serafino, Salvatore Filippone
ACM Trans. Math. Softw.3
2006 An Enhanced Parallel Version of Kiva-3V, Coupled with a 1D CFD Code, and Its Use in General Purpose Engine Applications
Gino Bella, Fabio Bozza, Alessandro De Maio, Francesco Del Citto, Salvatore Filippone
HPCC5
2005 FAST-EVP: An Engine Simulation Tool
Gino Bella, Alfredo Buttari, Alessandro De Maio, Francesco Del Citto, Salvatore Filippone, Fabiano Gasperini
HPCC5
2000 PSBLAS: a library for parallel linear algebra computation on sparse matrices
abstract
Many computationally intensive problems in engineering and science give rise to the solution of large, sparse, linear systems of equations. Fast and efficient methods for their soltion are very important because these systems usually occur in the innermost loop of the computational scheme. Parallelization is often necessary to achieve an acceptable level of performance. This paper presents the design, implementation, and interface of a library of Basic Linear Algebra Subroutines for sparse matrices (PSBLAS) which is specifically tailored to distributed-memory computers. PSBLAS enables easy, efficient, and portable implementations of parallel iterative solvers for linear systems. The interface keeps in view a Single Program Multiple Data programming model on distributed-memory machines. However, the architecture of the library does not exclude an implementation in different paradigms, such as those based on the shared-memory model.
Salvatore Filippone, Michele Colajanni
ACM Trans. Math. Softw.1
1990 A vectorized long-period shift-register random number generator
abstract
A pseudorandom number generator, based on a linear-feedback shift-register sequence, is presented. The very long period of the generator, 2/sup 1279/-1, makes it useful in modern statistical simulations. The proposed generator overcomes the limitations of multiplicative-congruential generators with modulus 2/sup 31/-1. The properties of linear-feedback shift-register sequences are reviewed, and a sequence of order p=1279 is proposed as a source of pseudorandom numbers. Results of the vectorization of shift register algorithm and of statistical tests are presented.>
Salvatore Filippone, Paolo Santangelo, Marcello Vitaletti
SC1