EDBT 2026 Demo / reviewers in the wild / expert
Valentina Cesare
dblp:265/2835
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-1119-4237ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Covariances computation in the Gaia AVU-GSR Parallel Solver with I/O techniques: a performance study as a function of writing cycle lengthabstractThe solver module of the Astrometric Verification Unit-Global Sphere Reconstruction (AVU-GSR) pipeline aims to find the astrometric parameters of $\sim 10^{8}$ stars in the Milky Way, the attitude and instrumental settings of the Gaia satellite, and the parametrized post Newtonian parameter $\gamma$ with a resolution of 1 0 - 100 micro – arc seconds. To perform this task, the code, which runs in production on Leonardo CINECA infrastructure, solves a system of linear equations with the iterative LSQR algorithm, where the coefficient matrix is large (10-50 TB) and sparse and the iterations stop when least square convergence is reached. The solver was ported to GPU with CUDA, obtaining a $\sim 14 x$ acceleration factor over an original version CPU-parallelized with OpenMP. This work concentrates on a code section dedicated to covariances calculation, representing an important scientific task for Gaia mission, since the problems unknowns present strong correlations. Given the number of unknowns at mission end, the variances-covariances matrix is expected to occupy $\sim 1$ EB, which represents a substantial “Big Data” issue. To compute a subset of the total covariances, we defined an I/Obased pipeline made of two jobs. The first job, the LSQR, writes the files every $i \operatorname{tnCov} C P$ iterations, and the second job reads them and calculates the corresponding covariances. The two jobs can be launched either in sequence or concurrently. Previous studies demonstrated that the covariances calculation does not significantly slowdown the AVU-GSR production up to $\sim 3 \times 10^{7}$ covariances. Here we investigate the performance of the covariances pipeline as a function of $i t n \operatorname{Cov} C P$. The results show that writing smaller files more frequently or writing larger files less frequently does not affect the global performance of the solver, whose speed only depends on the number of covariances to calculate and of system unknowns. Valentina Cesare, Ugo Becciani, Alberto Vecchiato, Mario Gilberto Lattanzi, Marco Aldinucci, Beatrice Bucciarelli |
PDP | 1 |
| 2025 | Performance Portability Assessment in GaiaabstractModern scientific experiments produce ever-increasing amounts of data, soon requiring ExaFLOPs computing capacities for analysis. Reaching such performance requires purpose-built supercomputers with$O(10^{3})$nodes, each hosting multicore CPUs and multiple GPUs, and applications designed to exploit this hardware optimally. Given that each supercomputer is generally a one-off project, the need for computing frameworks portable across diverse CPU and GPU architectures without performance losses is increasingly compelling. We investigate the performance portability (ȹ) of a real-world application: the solver module of the AVU–GSR pipeline for the ESA Gaia mission. This code finds the astrometric parameters of$\sim$$10^{8}$stars in the Milky Way using the LSQR iterative algorithm. LSQR is widely used to solve linear systems of equations across a wide range of high-performance computing applications, elevating the study beyond its astrophysical relevance. The code is memory-bound, with six main compute kernels implementing sparse matrix-by-vector products. We optimize the previous CUDA implementation and port the code to further six GPU-acceleration frameworks: C++ PSTL, SYCL, OpenMP, HIP, KOKKOS, and OpenACC. We evaluate each framework's performance portability across multiple GPUs (NVIDIA and AMD) and problem sizes in terms of application and architectural efficiency. Architectural efficiency is estimated through the roofline model of the six most computationally expensive GPU kernels. Our results show that C++ library-based (C++ PSTL and KOKKOS), pragma-based (OpenMP and OpenACC), and language-specific (CUDA, HIP, and SYCL) frameworks achieve increasingly better performance portability across the supported platforms with larger problem sizes providing better ȹ scores due to higher GPU occupancies. Giulio Malenza, Valentina Cesare, Marco Edoardo Santimaria, Robert Birke, Alberto Vecchiato, Ugo Becciani, Marco Aldinucci |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Toward HPC application portability via C++ PSTL: the Gaia AVU-GSR code assessmentabstractAbstract The computing capacity needed to process the data generated in modern scientific experiments is approaching ExaFLOPs. Currently, achieving such performances is only feasible through GPU-accelerated supercomputers. Different languages were developed to program GPUs at different levels of abstraction. Typically, the more abstract the languages, the more portable they are across different GPUs. However, the less abstract and co-designed with the hardware, the more room for code optimization and, eventually, the more performance. In the HPC context, portability and performance are a fairly traditional dichotomy. The current C++ Parallel Standard Template Library (PSTL) has the potential to go beyond this dichotomy. In this work, we analyze the main performance benefits and limitations of PSTL using as a use-case the Gaia Astrometric Verification Unit-Global Sphere Reconstruction parallel solver developed by the European Space Agency Gaia mission. The code aims to find the astrometric parameters of $$\sim10^8$$ ∼ 10 8 stars in the Milky Way by iteratively solving a linear system of equations with the LSQR algorithm, originally GPU-ported with the CUDA language. We show that the performance obtained with the PSTL version, which is intrinsically more portable than CUDA, is comparable to the CUDA one on NVIDIA GPU architecture. Giulio Malenza, Valentina Cesare, Marco Aldinucci, Ugo Becciani, Alberto Vecchiato |
J. Supercomput. | 2 |
| 2021 | Practical parallelization of scientific applications with OpenMP, OpenACC and MPI
Marco Aldinucci, Valentina Cesare, Iacopo Colonnelli, Alberto Riccardo Martinelli, Gianluca Mittone, Barbara Cantalupo, Carlo Cavazzoni, Maurizio Drocco |
J. Parallel Distributed Comput. | 2 |
| 2020 | Practical Parallelization of Scientific ApplicationsabstractThis work aims at distilling a systematic methodology to modernize existing sequential scientific codes with a limited re-designing effort, turning an old codebase into modern code, i.e., parallel and robust code. We propose an automatable methodology to parallelize scientific applications designed with a purely sequential programming mindset, thus possibly using global variables, aliasing, random number generators, and stateful functions. We demonstrate the methodology by way of an astrophysical application, where we model at the same time the kinematic profiles of 30 disk galaxies with a Monte Carlo Markov Chain (MCMC), which is sequential by definition. The parallel code exhibits a 12 times speedup on a 48-core platform. Valentina Cesare, Iacopo Colonnelli, Marco Aldinucci |
PDP | 1 |