Marta Garcia-Gasulla

dblp:40/7785 · also Marta Garcia 0001 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2026 A Compiler-Assisted Workflow for Efficiency-Guided Selective Tracing
Sebastian Kreutzer, Valentin Seitz, Joan Vinyals-Ylla-Catala, Tim Heldmann, Christian Iwainsky, Marta Garcia-Gasulla, Jesús Labarta, Christian H. Bischof
Euro-Par (1)6
2026 Introducing MareNostrum5: A European pre-exascale energy-efficient system designed to serve a broad spectrum of scientific workloads
Fabio Banchelli, Marta Garcia-Gasulla, Filippo Mantovani, Joan Vinyals-Ylla-Catala, Josep Pocurull, David Vicente, Beatriz Eguzkitza, Flavio Cesar Cunha Galeazzo, Mario C. Acosta, Sergi Girona
Future Gener. Comput. Syst.2
2026 Exploring RISC-V long vector capabilities: A case study in Earth Sciences
Fabio Banchelli, David Jurado, Marta Garcia-Gasulla, Filippo Mantovani
Future Gener. Comput. Syst.3
2025 NSYS2PRV: Detailed and Quantitative Analysis of Large-Scale GPU Execution Traces with Paraver
abstract
This work presents a tool, a methodology, a set of metrics, and practical examples for evaluating the performance of large-scale AI and traditional HPC applications using GPUs. NSYS2PRV is a tool that converts NVIDIA Nsight Systems reports into traces compatible with Paraver, enabling significantly enhanced insight compared to current performance analysis practices. By leveraging the capabilities of a well-established HPC performance analysis tool, we enable the comparison of execution traces and the quantification of microscopic-level differences to explain behaviors across hundreds or more computing devices. We argue that large-scale GPU applications and AI workloads can greatly benefit from the type of large-scale performance analysis introduced here, an approach that is not yet widely adopted in this domain. Translating nsys-generated traces to Paraver allows analysts to combine the fine-grained, highly accurate execution data obtainable from proprietary tools with the flexibility and scalability of an open-source, parallel performance analysis environment. Paraver also enables easy, customizable computation of efficiency metrics. This work demonstrates a more effective and insightful analysis experience than that offered by the native visualization tools in Nsight Systems. Additionally, we introduce a set of Paravercompatible metrics that guide the analysis process, and we showcase examples where these metrics were successfully applied to real-world AI and HPC workloads.
Marc Clascà, Jesús Labarta, Marta Garcia-Gasulla
CLUSTER3
2024 Exploiting long vectors with a CFD code: a co-design show case
abstract
A current trend in HPC systems is the utilization of architectures with SIMD or vector extensions to exploit data parallelism. There are several ways to take advantage of such modern vector architectures, each with a different impact on the code and its portability. For example, the use of intrinsics, guided vectorization via pragmas, or compiler autovectorization. Our objectives are to maximize vectorization efficiency and minimize code specialization. To achieve these objectives, we rely on compiler autovectorization. We leverage a set of hardware and software tools that allow us to analyze in detail where autovectorization is suboptimal. Thus, we apply an iterative methodology that allows us to incrementally improve the efficient use of the underlying hardware. In this paper, we apply this methodology to a CFD production code. We evaluate the performance on an innovative configurable platform powered by a RISC-V core coupled with a wide vector unit capable of operating with up to 256 double precision elements. Following the vectorization process, we demonstrate a single-core speedup of 7.6× compared to its scalar implementation. Furthermore, we show that code portability is not compromised, as our solution continues to exhibit performance benefits, or at the very least, no drawbacks, on other HPC architectures such as Intel x86 and NEC SX-Aurora.
Marc Blancafort, Roger Ferrer, Guillaume Houzeaux, Marta Garcia-Gasulla, Filippo Mantovani
IPDPS4
2024 Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities
abstract
With the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research.
Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001
IEEE Trans. Parallel Distributed Syst.14
2022 Transparent load balancing of MPI programs using [email protected] and DLB
abstract
Abstract Load imbalance is a long-standing source of inefficiency in high performance computing. The situation has only got worse as applications and systems increase in complexity, e.g., adaptive mesh refinement, DVFS, memory hierarchies, power and thermal management, and manufacturing processes. Load balancing is often implemented in the application, but it obscures application logic and may need extensive code refactoring. This paper presents an automated and transparent dynamic load balancing approach for MPI applications with OmpSs-2 tasks, which relieves applications from this burden. Only local and trivial changes are required to the application. Our approach exploits the ability of [email protected] to offload tasks for execution on other nodes, and it reallocates compute resources among ranks using the Dynamic Load Balancing (DLB) library. It employs LeWI to react to fine-grained load imbalances and DROM to address coarse-grained load imbalances by reserving cores on other nodes that can be reclaimed on demand. We use an expander graph to limit the amount of point-to-point communication and state. The results show 46% reduction in time-to-solution for micro-scale solid mechanics on 32 nodes and a 20% reduction beyond DLB for n-body on 16 nodes, when one node is running slow. A synthetic benchmark shows that performance is within 10% of optimal for an imbalance of up to 2.0 on 8 nodes. All software is released open source.
Jimmy Aguilar Mena, Omar Shaaban, Víctor López 0003, Marta Garcia-Gasulla, Paul M. Carpenter, Eduard Ayguadé, Jesús Labarta
ICPP4
2022 Hybrid parallelization of molecular dynamics simulations to reduce load imbalance
Julian Morillo, Maxime Vassaux, Peter V. Coveney, Marta Garcia-Gasulla
J. Supercomput.4
2021 Cluster of emerging technology: evaluation of a production HPC system based on A64FX
abstract
Clusters of emerging technologies are appearing with more and more frequency in HPC. After years of skepticism, data-centers are adopting them as production systems thanks to several geopolitical and technological factors. The most honorable example is the Fugaku supercomputer, powered by the latest Fujitsu A64FX CPU. Which is the behavior of mature HPC codes on such emerging technology clusters? Which performance will obtain scientists when running their HPC applications “as is” on these clusters? This paper presents the evaluation of CTE-Arm, a Fugaku-like system, including both fine-tuned micro-benchmarks and five scientific applications run without prior fine-tuning: Alya, NEMO, Gromacs, OpenIFS, and WRF. Results show that while micro-architectural benchmarks show performance as expected, the performance obtained running HPC applications not tuned for a specific architecture are between $2\times $ and $4\times $ slower compared with a standard Intel-based HPC system. Therefore further effort is needed to improve tools (e.g., compilers) and system software (e.g., MPI libraries) to ease applications deployment and improve their performance.
Fabio Banchelli, Kilian Peiro, Guillem Ramirez-Gargallo, Joan Vinyals-Ylla-Catala, David Vicente, Marta Garcia-Gasulla, Filippo Mantovani
CLUSTER6
2021 knowlEdge Project -Concept, Methodology and Innovations for Artificial Intelligence in Industry 4.0
abstract
AI is one of the biggest megatrends towards the 4th industrial revolution. Although these technologies promise business sustainability as well as product and process quality, it seems that the ever-changing market demands, the complexity of technologies and fair concerns about privacy, impede broad application and reuse of Artificial Intelligence (AI) models across the industry. To break the entry barriers for these technologies and unleash its full potential, the knowlEdge project will develop a new generation of AI methods, systems, and data management infrastructure. Subsequently, as part of the knowlEdge project we propose several major innovations in the areas of data management, data analytics and knowledge management including (i) a set of AI services that allows the usage of edge deployments as computational and live data infrastructure as well as a continuous learning execution pipeline on the edge, (ii) a digital twin of the shop-floor able to test AI models, (iii) a data management framework deployed along the edge-to-cloud continuum ensuring data quality, privacy and confidentiality, (iv) Human-AI Collaboration and Domain Knowledge Fusion tools for domain experts to inject their experience into the system, (v) a set of standardisation mechanisms for the exchange of trained AI models from one context to another, and (vi) a knowledge marketplace platform to distribute and interchange trained AI models. In this paper, we present a short overview of the EU Project knowlEdge –Towards Artificial Intelligence powered manufacturing services, processes, and products in an edge-to-cloud-knowledge continuum for humans [in-the-loop], which is funded by the Horizon 2020 (H2020) Framework Programme of the European Commission under Grant Agreement 957331. Our overview includes a description of the project’s main concept and methodology as well as the envisioned innovations.
Sergio Álvarez-Napagao, Boki Ashmore, Marta Barroso, Cristian Barrué, Christian Beecks, Fabian Berns, Ilaria Bosi, Sisay Adugna Chala, Nicola Ciulli, Marta Garcia-Gasulla, Alexander Graß, Dimosthenis Ioannidis, Natalia Jakubiak, Karl Köpke, Ville Lämsä, Pedro Megias, Alexandros Nizamis, Claudio Pastrone, Rosaria Rossini, Miquel Sànchez-Marrè, Luca Ziliotti
INDIN10
2020 CoreNEURON: Performance and Energy Efficiency Evaluation on Intel and Arm CPUs
abstract
The simulation of detailed neuronal circuits is based on computationally expensive software simulations and requires access to a large computing cluster. The appearance of new Instruction Set Architectures (ISAs) in most recent High-Performance Computing (HPC) systems, together with the layers of system software and complex scientific applications running on top of them, makes the performance and power figures challenging to evaluate. In this paper, we focus on evaluating CoreNEURON on two HPC systems powered by Intel and Arm architectures. CoreNEURON is a computational engine of the widely used NEURON simulator adapted to run on emerging architectures while maintaining compatibility with existing NEURON models developed by the neuroscience community. The evaluation is based on the analysis of the dynamic instruction mix on two versions of CoreNEURON. It focuses on the performance gain obtained by exploiting the Single Instruction Multiple Data (SIMD) unit and includes energy measurements. Our results show that using a tool for increasing data-level parallelism (ISPC) boosts the performance up to 2× independently on the ISA. Its combination with vendor-specific compilers can further speed up the neural simulation time. Also, the performance/price ratio is higher for Arm-based systems than for Intel ones making them more cost-efficient keeping the same usability level of other HPC systems.
Joel Criado, Marta Garcia-Gasulla, Pramod S. Kumbhar, Omar Awile, Ioannis Magkanaris, Filippo Mantovani
CLUSTER2
2020 Performance study of HPC applications on an Arm-based cluster using a generic efficiency model
abstract
HPC systems and parallel applications are increasing their complexity. Therefore the possibility of easily study and project at large scale the performance of scientific applications is of paramount importance. In this paper we describe a performance analysis method and we apply it to four complex HPC applications. We perform our study on a pre-production HPC system powered by the latest Arm-based CPUs for HPC, the Marvell ThunderX2. For each application we spot inefficiencies and factors that limit their scalability. The results show that in several cases the bottlenecks do not come from the hardware but from the way applications are programmed or the way the system software is configured.
Fabio Banchelli, Kilian Peiro, Andrea Querol, Guillem Ramirez-Gargallo, Guillem Ramirez-Miranda, Joan Vinyals-Ylla-Catala, Pablo Vizcaino, Marta Garcia-Gasulla, Filippo Mantovani
PDP8
2020 Heterogeneous CPU/GPU co-execution of CFD simulations on the POWER9 architecture: Application to airplane aerodynamics
Ricard Borrell, Damien Dosimont, Marta Garcia-Gasulla, Guillaume Houzeaux, Oriol Lehmkuhl, V. Mehta, Herbert Owen, Mariano Vázquez, Guillermo Oyarzun
Future Gener. Comput. Syst.3
2020 Performance and energy consumption of HPC workloads on a cluster based on Arm ThunderX2 CPU
Filippo Mantovani, Marta Garcia-Gasulla, José Gracia, Esteban Stafford, Fabio Banchelli, Marc Josep-Fabrego, Joel Criado, Mathias Nachtmann
Future Gener. Comput. Syst.2
2019 TensorFlow on State-of-the-Art HPC Clusters: A Machine Learning use Case
abstract
The recent rapid growth of the data-flow programming paradigm enabled the development of specific architectures, e.g., for machine learning. The most known example is the Tensor Processing Unit (TPU) by Google. Standard data-centers, however, still can not foresee large partitions dedicated to machine learning specific architectures. Within data-centers, the High-Performance Computing (HPC) clusters are highly parallel machines targeting a broad class of compute-intensive workflows, as such they can be used for tackling machine learning challenges. On top of this, HPC architectures are rapidly changing, including accelerators and instruction sets other than the classical x86 CPUs. In this blurry scenario, identifying which are the best hardware/software configurations to efficiently support machine learning workloads on HPC clusters is not trivial. In this paper, we considered the workflow of TensorFlow for image recognition. We highlight the strong dependency of the performance in the training phase on the availability of arithmetic libraries optimized for the underlying architecture. Following the example of Intel leveraging the MKL libraries for improving the TensorFlow performance, we plugged the Arm Performance Libraries into TensorFlow and tested on an HPC cluster based on Marvell ThunderX2 CPUs. Also, we performed a scalability study on three state-of-the-art HPC clusters based on different CPU architectures, x86 Intel Skylake, Arm-v8 Marvell ThunderX2, and PowerPC IBM Power9.
Guillem Ramirez-Gargallo, Marta Garcia-Gasulla, Filippo Mantovani
CCGRID2
2019 Containers in HPC: A Scalability and Portability Study in Production Biological Simulations
abstract
Since the appearance of Docker in 2013, container technologies for computers have evolved and gained importance in cloud data centers. However, adoption of containers in High-Performance Computing (HPC) centers is still under discussion: on one hand, the ease in portability is very well accepted; on the other hand, the performance penalties and security issues introduced by the added software layers are often under scrutiny. Since very little evaluation of large production HPC codes running in containers is available, we provide in this paper a comparative study using a production simulation of a biological system. The simulation is performed using Alya, which is a computational fluid dynamics (CFD) code optimized for HPC environments and enabled to run multiphysics problems. In the paper, we analyze the productivity advantages of adopting containers for large HPC codes, and we quantify performance overhead induced by the use of three different container technologies (Docker, Singularity and Shifter) comparing it to native execution. Given the results of these tests, we selected Singularity as best technology, based on performance and portability. We show scalability results of Alya using singularity up to 256 computational nodes (up to 12k cores) of MareNostrum4 and present a study of performance and portability on three different HPC architectures (Intel Skylake, IBM Power9, and Arm-v8).
Oleksandr Rudyy, Marta Garcia-Gasulla, Filippo Mantovani, Alfonso Santiago, Raül Sirvent, Mariano Vázquez
IPDPS2
2017 DJSB: Dynamic Job Scheduling Benchmark
Víctor López 0003, Ana Jokanovic, Marco D'Amico, Marta Garcia-Gasulla, Raül Sirvent, Julita Corbalán
JSSPP4
2014 Hints to improve automatic load balancing with LeWI for hybrid applications
Marta Garcia-Gasulla, Jesús Labarta, Julita Corbalán
J. Parallel Distributed Comput.1
2009 LeWI: A Runtime Balancing Algorithm for Nested Parallelism
abstract
We present LeWI: a novel load balancing algorithm, that can balance applications with very different patterns of imbalance. Our algorithm can balance fine grain imbalances, non iterative applications and applications with irregular imbalance. To achieve this LeWI reassigns the computational resources of blocked processes to other processes more loaded. We have implemented LeWI within DLB a Dynamic Load Balancing Library developed by us. DLB helps parallel programming models to make the most of the computational power available with the minimum effort. It solves the imbalance among processes in applications with two levels of parallelism using the malleability of the inner level. The performance evaluation shows that LeWI, the novel balancing algorithm we are presenting in this paper, together with DLB is able to improve the performance of a different range of unbalanced applications and when applied to well balanced applications it does not introduce significant overhead. Therefore we present a mechanism that can be used with any hybrid application without needing a programmer to analyze the application nor modify it.
Marta Garcia-Gasulla, Julita Corbalán, Jesús Labarta
ICPP1