Filippo Mantovani

dblp:88/5952 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-3559-4825ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 8 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
Fabio Banchelli, Marta Garcia-Gasulla, Filippo Mantovani, Joan Vinyals-Ylla-Catala, Josep Pocurull, David Vicente, Beatriz Eguzkitza, Flavio Cesar Cunha Galeazzo, Mario C. Acosta, Sergi Girona
Future Gener. Comput. Syst.3
2026 Exploring RISC-V long vector capabilities: A case study in Earth Sciences
Fabio Banchelli, David Jurado, Marta Garcia-Gasulla, Filippo Mantovani
Future Gener. Comput. Syst.4
2026 Designing a QEMU plugin to profile multicore long vector RISC-V architectures: RAVE
Pablo Vizcaino, Roger Ferrer, Jesús Labarta, Filippo Mantovani
Future Gener. Comput. Syst.4
2024 Exploiting long vectors with a CFD code: a co-design show case
abstract
A current trend in HPC systems is the utilization of architectures with SIMD or vector extensions to exploit data parallelism. There are several ways to take advantage of such modern vector architectures, each with a different impact on the code and its portability. For example, the use of intrinsics, guided vectorization via pragmas, or compiler autovectorization. Our objectives are to maximize vectorization efficiency and minimize code specialization. To achieve these objectives, we rely on compiler autovectorization. We leverage a set of hardware and software tools that allow us to analyze in detail where autovectorization is suboptimal. Thus, we apply an iterative methodology that allows us to incrementally improve the efficient use of the underlying hardware. In this paper, we apply this methodology to a CFD production code. We evaluate the performance on an innovative configurable platform powered by a RISC-V core coupled with a wide vector unit capable of operating with up to 256 double precision elements. Following the vectorization process, we demonstrate a single-core speedup of 7.6× compared to its scalar implementation. Furthermore, we show that code portability is not compromised, as our solution continues to exhibit performance benefits, or at the very least, no drawbacks, on other HPC architectures such as Intel x86 and NEC SX-Aurora.
Marc Blancafort, Roger Ferrer, Guillaume Houzeaux, Marta Garcia-Gasulla, Filippo Mantovani
IPDPS5
2023 Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extension
abstract
Summary Novel architectures leveraging long and variable vector lengths like the NEC SX‐Aurora or the vector extension of RISCV are appearing as promising solutions on the supercomputing market. These architectures often require re‐coding of scientific kernels. For example, traditional implementations of algorithms for computing the fast Fourier transform (FFT) cannot take full advantage of vector architectures. In this article, we present the implementation of FFT algorithms able to leverage these novel architectures. We evaluate these codes on NEC SX‐Aurora , comparing them with the optimized NEC libraries; and in a prototype of a RISC‐V core with a vector processing unit. We present the benefits and limitations of two approaches of RADIX‐2 FFT vector implementations. We show that our approach makes better use of the vector unit of the NEC SX‐Aurora , reaching higher or equal performance than the optimized NEC library. More generally, we prove the importance of maximizing the vector length usage of the algorithm, taking advantage of the FFT properties to reduce long‐latency vector operations, and reordering the instructions according to the specific hardware features to boost the performance of FFT‐like computational kernels.
Pablo Vizcaino, Filippo Mantovani, Roger Ferrer, Jesús Labarta
Concurr. Comput. Pract. Exp.2
2023 HPCG on long-vector architectures: Evaluation and optimization on NEC SX-Aurora and RISC-V
Constantino Gómez, Filippo Mantovani, Erich Focht, Marc Casas
Future Gener. Comput. Syst.2
2022 Asymmetric HMMs for Online Ball-Bearing Health Assessments
abstract
The degradation of critical components inside large industrial assets, such as ball-bearings, has a negative impact on production facilities, reducing the availability of assets due to an unexpectedly high failure rate. Machine learning-based monitoring systems can estimate the remaining useful life (RUL) of ball bearings, reducing the downtime by early failure detection. However, traditional approaches for predictive systems require run-to-failure (RTF) data as training data, which in real scenarios can be scarce and expensive to obtain as the expected useful life could be measured in years. Therefore, to overcome the need of RTF, we propose a new methodology based on online novelty detection and asymmetrical hidden Markov models (As-HMMs) to work out the health assessment. This new methodology does not require previous RTF data and can adapt to natural degradation of mechanical components over time in data-stream and online environments. As the system is designed to work online within the electrical cabinet of machines, it has to be deployed using embedded electronics. Therefore, a performance analysis of As-HMM is presented to detect the strengths and critical points of the algorithm. To validate our approach, we use real life ball-bearing data sets and compare our methodology with other methodologies where no RTF data are needed and check the advantages in RUL prediction and health monitoring. As a result, we showcase a complete end-to-end solution from the sensor to actionable insights regarding RUL estimation toward maintenance application in real industrial environments.
Carlos Puerto-Santana, Concha Bielza, Javier Diaz-Rozo, Guillem Ramirez-Gargallo, Filippo Mantovani, Gaizka Virumbrales, Jesús Labarta, Pedro Larrañaga
IEEE Internet Things J.5
2021 Cluster of emerging technology: evaluation of a production HPC system based on A64FX
abstract
Clusters of emerging technologies are appearing with more and more frequency in HPC. After years of skepticism, data-centers are adopting them as production systems thanks to several geopolitical and technological factors. The most honorable example is the Fugaku supercomputer, powered by the latest Fujitsu A64FX CPU. Which is the behavior of mature HPC codes on such emerging technology clusters? Which performance will obtain scientists when running their HPC applications “as is” on these clusters? This paper presents the evaluation of CTE-Arm, a Fugaku-like system, including both fine-tuned micro-benchmarks and five scientific applications run without prior fine-tuning: Alya, NEMO, Gromacs, OpenIFS, and WRF. Results show that while micro-architectural benchmarks show performance as expected, the performance obtained running HPC applications not tuned for a specific architecture are between $2\times $ and $4\times $ slower compared with a standard Intel-based HPC system. Therefore further effort is needed to improve tools (e.g., compilers) and system software (e.g., MPI libraries) to ease applications deployment and improve their performance.
Fabio Banchelli, Kilian Peiro, Guillem Ramirez-Gargallo, Joan Vinyals-Ylla-Catala, David Vicente, Marta Garcia-Gasulla, Filippo Mantovani
CLUSTER7
2021 Efficiently running SpMV on long vector architectures
abstract
Sparse Matrix-Vector multiplication (SpMV) is an essential kernel for parallel numerical applications. SpMV displays sparse and irregular data accesses, which complicate its vectorization. Such difficulties make SpMV to frequently experiment non-optimal results when run on long vector ISAs exploiting SIMD parallelism. In this context, the development of new optimizations becomes fundamental to enable high performance SpMV executions on emerging long vector architectures. In this paper, we improve the state-of-the-art SELL-C-σ sparse matrix format by proposing several new optimizations for SpMV. We target aggressive long vector architectures like the NEC Vector Engine. By combining several optimizations, we obtain an average 12% improvement over SELL-C-σ considering a heterogeneous set of 24 matrices. Our optimizations boost performance in long vector architectures since they expose a high degree of SIMD parallelism.
Constantino Gómez, Filippo Mantovani, Erich Focht, Marc Casas
PPoPP2
2020 CoreNEURON: Performance and Energy Efficiency Evaluation on Intel and Arm CPUs
abstract
The simulation of detailed neuronal circuits is based on computationally expensive software simulations and requires access to a large computing cluster. The appearance of new Instruction Set Architectures (ISAs) in most recent High-Performance Computing (HPC) systems, together with the layers of system software and complex scientific applications running on top of them, makes the performance and power figures challenging to evaluate. In this paper, we focus on evaluating CoreNEURON on two HPC systems powered by Intel and Arm architectures. CoreNEURON is a computational engine of the widely used NEURON simulator adapted to run on emerging architectures while maintaining compatibility with existing NEURON models developed by the neuroscience community. The evaluation is based on the analysis of the dynamic instruction mix on two versions of CoreNEURON. It focuses on the performance gain obtained by exploiting the Single Instruction Multiple Data (SIMD) unit and includes energy measurements. Our results show that using a tool for increasing data-level parallelism (ISPC) boosts the performance up to 2× independently on the ISA. Its combination with vendor-specific compilers can further speed up the neural simulation time. Also, the performance/price ratio is higher for Arm-based systems than for Intel ones making them more cost-efficient keeping the same usability level of other HPC systems.
Joel Criado, Marta Garcia-Gasulla, Pramod S. Kumbhar, Omar Awile, Ioannis Magkanaris, Filippo Mantovani
CLUSTER6
2020 Performance study of HPC applications on an Arm-based cluster using a generic efficiency model
abstract
HPC systems and parallel applications are increasing their complexity. Therefore the possibility of easily study and project at large scale the performance of scientific applications is of paramount importance. In this paper we describe a performance analysis method and we apply it to four complex HPC applications. We perform our study on a pre-production HPC system powered by the latest Arm-based CPUs for HPC, the Marvell ThunderX2. For each application we spot inefficiencies and factors that limit their scalability. The results show that in several cases the bottlenecks do not come from the hardware but from the way applications are programmed or the way the system software is configured.
Fabio Banchelli, Kilian Peiro, Andrea Querol, Guillem Ramirez-Gargallo, Guillem Ramirez-Miranda, Joan Vinyals-Ylla-Catala, Pablo Vizcaino, Marta Garcia-Gasulla, Filippo Mantovani
PDP9
2020 Performance and energy consumption of HPC workloads on a cluster based on Arm ThunderX2 CPU
Filippo Mantovani, Marta Garcia-Gasulla, José Gracia, Esteban Stafford, Fabio Banchelli, Marc Josep-Fabrego, Joel Criado, Mathias Nachtmann
Future Gener. Comput. Syst.1
2019 TensorFlow on State-of-the-Art HPC Clusters: A Machine Learning use Case
abstract
The recent rapid growth of the data-flow programming paradigm enabled the development of specific architectures, e.g., for machine learning. The most known example is the Tensor Processing Unit (TPU) by Google. Standard data-centers, however, still can not foresee large partitions dedicated to machine learning specific architectures. Within data-centers, the High-Performance Computing (HPC) clusters are highly parallel machines targeting a broad class of compute-intensive workflows, as such they can be used for tackling machine learning challenges. On top of this, HPC architectures are rapidly changing, including accelerators and instruction sets other than the classical x86 CPUs. In this blurry scenario, identifying which are the best hardware/software configurations to efficiently support machine learning workloads on HPC clusters is not trivial. In this paper, we considered the workflow of TensorFlow for image recognition. We highlight the strong dependency of the performance in the training phase on the availability of arithmetic libraries optimized for the underlying architecture. Following the example of Intel leveraging the MKL libraries for improving the TensorFlow performance, we plugged the Arm Performance Libraries into TensorFlow and tested on an HPC cluster based on Marvell ThunderX2 CPUs. Also, we performed a scalability study on three state-of-the-art HPC clusters based on different CPU architectures, x86 Intel Skylake, Arm-v8 Marvell ThunderX2, and PowerPC IBM Power9.
Guillem Ramirez-Gargallo, Marta Garcia-Gasulla, Filippo Mantovani
CCGRID3
2019 Design Space Exploration of Next-Generation HPC Machines
abstract
The landscape of High Performance Computing (HPC) system architectures keeps expanding with new technologies and increased complexity. With the goal of improving the efficiency of next-generation large HPC systems, designers require tools for analyzing and predicting the impact of new architectural features on the performance of complex scientific applications at scale. We simulate five hybrid (MPI+OpenMP) applications over 864 architectural proposals based on stateof-the-art and emerging HPC technologies, relevant both in industry and research. This paper significantly extends our previous work with MUltiscale Simulation Approach (MUSA) enabling accurate performance and power estimations of large-scale HPC systems. We reveal that several applications present critical scalability issues mostly due to the software parallelization approach. Looking at speedup and energy consumption exploring the design space (i.e., changing memory bandwidth, number of cores, and type of cores), we provide evidence-based architectural recommendations that will serve as hardware and software codesign guidelines.
Constantino Gómez, Francesc Martínez, Adrià Armejach, Miquel Moretó, Filippo Mantovani, Marc Casas
IPDPS5
2019 Containers in HPC: A Scalability and Portability Study in Production Biological Simulations
abstract
Since the appearance of Docker in 2013, container technologies for computers have evolved and gained importance in cloud data centers. However, adoption of containers in High-Performance Computing (HPC) centers is still under discussion: on one hand, the ease in portability is very well accepted; on the other hand, the performance penalties and security issues introduced by the added software layers are often under scrutiny. Since very little evaluation of large production HPC codes running in containers is available, we provide in this paper a comparative study using a production simulation of a biological system. The simulation is performed using Alya, which is a computational fluid dynamics (CFD) code optimized for HPC environments and enabled to run multiphysics problems. In the paper, we analyze the productivity advantages of adopting containers for large HPC codes, and we quantify performance overhead induced by the use of three different container technologies (Docker, Singularity and Shifter) comparing it to native execution. Given the results of these tests, we selected Singularity as best technology, based on performance and portability. We show scalability results of Alya using singularity up to 256 computational nodes (up to 12k cores) of MareNostrum4 and present a study of performance and portability on three different HPC architectures (Intel Skylake, IBM Power9, and Arm-v8).
Oleksandr Rudyy, Marta Garcia-Gasulla, Filippo Mantovani, Alfonso Santiago, Raül Sirvent, Mariano Vázquez
IPDPS3
2018 Efficient CFD code implementation for the ARM-based Mont-Blanc architecture
abstract
Since 2011, the European project Mont-Blanc has been focused on enabling ARM-based technology for HPC, developing both hardware platforms and system software. The latest Mont-Blanc prototypes use system-on-chip (SoC) devices that combine a CPU and a GPU sharing a common main memory. Specific developments of parallel computing software and well-suited implementation approaches are of crucial importance for such a heterogeneous architecture in order to efficiently exploit its potential. This paper is devoted to the optimizations carried out in the TermoFluids CFD code to efficiently run it on the Mont-Blanc system. The underlying numerical method is based on an unstructured finite-volume discretization of the Navier–Stokes equations for the numerical simulation of incompressible turbulent flows. It is implemented using a portable and modular operational approach based on a minimal set of linear algebra operations. An architecture-specific heterogeneous multilevel MPI+OpenMP+OpenCL implementation of such kernels is proposed. It includes optimizations of the storage formats, dynamic load balancing between the CPU and GPU devices and hiding of communication overheads by overlapping computations and data transfers. A detailed performance study shows time reductions of up to 2.1× on the kernels’ execution with the new heterogeneous implementation, its scalability on up to 128 Mont-Blanc nodes and the energy savings (around 40%) achieved with the Mont-Blanc system versus the high-end hybrid supercomputer MinoTauro.
Guillermo Oyarzun, Ricard Borrell, Andrey V. Gorobets, Filippo Mantovani, Assensi Oliva
Future Gener. Comput. Syst.4
2016 The mont-blanc prototype: an alternative approach for HPC systems
abstract
High-performance computing (HPC) is recognized as one of the pillars for further progress in science, industry, medicine, and education. Current HPC systems are being developed to overcome emerging architectural challenges in order to reach Exascale level of performance, projected for the year 2020. The much larger embedded and mobile market allows for rapid development of intellectual property (IP) blocks and provides more flexibility in designing an application-specific system-on-chip (SoC), in turn providing the possibility in balancing performance, energy-efficiency, and cost. In the Mont-Blanc project, we advocate for HPC systems being built from such commodity IP blocks, currently used in embedded and mobile SoCs. As a first demonstrator of such an approach, we present the Mont-Blanc prototype; the first HPC system built with commodity SoCs, memories, and network interface cards (NICs) from the embedded and mobile domain, and off-the-shelf HPC networking, storage, cooling, and integration solutions. We present the system's architecture and evaluate both performance and energy efficiency. Further, we compare the system's abilities against a production level supercomputer. At the end, we discuss parallel scalability and estimate the maximum scalability point of this approach across a set of applications.
Nikola Rajovic, Alejandro Rico, Filippo Mantovani, Daniel Ruiz 0003, Josep Oriol Vilarrubi, Constantino Gómez, Luna Backes, Diego Nieto, Harald Servat, Xavier Martorell, Jesús Labarta, Eduard Ayguadé, Chris Adeniyi-Jones, Said Derradji, Hervé Gloaguen, Piero Lanucara, Nico Sanna, Jean-François Méhaut, Kevin Pouget, Brice Videau, Eric Boyer, Momme Allalen, Axel Auweter, David Brayford, Daniele Tafani, Volker Weinberg, Dirk Brömmel, René Halver, Jan H. Meinke, Ramón Beivide, Mariano Benito, Enrique Vallejo 0001, Mateo Valero, Alex Ramírez
SC3
2006 Poster reception - IANUS: scientific computing on an FPGA-based architecture
abstract
IANUS is a massively parallel system based on a 2D array of FPGA-based processors with nearest-neighbor connections. Processors are also directly connected to a central hub attached to a host computer.The prototype, available in October 2006 uses an array of 4x4 Xilinx Virtex4LX160 FPGA's.We map onto the array the computational kernels of scientific applications characterized by regular control flow, unconventional mix of data-manipulation operations and limited memory usage.Careful VHDL coding of the kernel algorithms relevant for Monte Carlo simulation of spin-glass systems (our first application) yields impressive performances: single processor tests concurrently update ~1000 spins, so average spin-update time is 15 psec. This is ~60 times faster than accurately programmed 3,2 GHz PC's. We plan to build a 256 nodes system, roughly equivalent to 15000 PC's.This poster describes the architecture, the implementation and the methodology with which a specific application is mapped onto the system.
Filippo Mantovani
SC1