Alberto Cascajo

dblp:290/5763 · also Alberto Cascajo García · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-5506-1431ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future Opportunities
abstract
With the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research.
Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001
IEEE Trans. Parallel Distributed Syst.3
2023 Dynamic management of processes and communicators in malleable MPI applications
abstract
A malleable application is defined as one that can increase or decrease its resources dynamically based on workload variations. These applications often leverage a job manager to handle the resources. The MPI standard incorporates strategies to increase the number of processes connected to a running application. However, it does not clarify how to use these same features to reduce said processes. This paper presents a strategy compatible with the latest version of the MPI standard to develop malleable MPI applications. It allows the addition and removal of resources at the computing node level in a discretionary manner. The results obtained from the evaluation show that the proposed strategy has acceptable performance and scalability for the functionality it provides.
Javier Fernández 0001, Alberto Cascajo, Jesús Carretero 0001
ICPADS2
2023 Evaluating the spread of Omicron COVID-19 variant in Spain
abstract
This work analyzes the propagation the highly transmissible COVID-19 variant Omicron across Spain via simulation by using EpiGraph. EpiGraph is an agent-based parallel simulator that reproduces the COVID-19 propagation over wide areas. In this work we consider a population of 19,574,086 individuals of the 63 most populated cities of Spain, for the time interval between May 15th 2021 and March 6th 2022. The main variants existing at the start of the simulation were the Alpha and Delta, with prevalence of 4% and 96%. Then, during the second half of November 2021, the Omicron variant appears in Spain. Due to the higher transmission of this new variant - about 2 times larger than Delta, it quickly spreads through all the cities and becomes the dominant strain in the country. In this work we analyze the propagation of this variant under different mobility restrictions and patient zero scenarios. We first define a baseline scenario which reproduces the existing conditions of the COVID-19 propagation in Spain for our period of study. We then consider alternative scenarios for different starting locations of the propagation. Finally, for each one of these scenarios, we evaluate different transportation intensities - i.e. movement of individuals between the cities. The main conclusion is that, independently of the initial location of the Omicron variant and the existing transportation conditions, the Omicron variant spreads through all the country in a short time interval. The work presented in this paper also implements and evaluates a power monitoring and optimization system aimed at reducing the energy consumption of such massive simulations as the ones performed in EpiGraph.
Miguel Guzmán-Merino, Maria-Cristina V. Marinescu, Alberto Cascajo, Jesús Carretero 0001, David E. Singh
Future Gener. Comput. Syst.3
2022 Improving Congestion Control through Fine-Grain Monitoring of InfiniBand Networks
abstract
Congestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms.
Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001
HOTI1
2021 LIMITLESS - LIght-weight MonItoring Tool for LargE Scale Systems
abstract
This work presents LIMITLESS, a HPC framework that provides new strategies for monitoring clusters. LIMITLESS is a scalable light-weight monitor that is integrated with other HPC runtimes in order to obtain an holistic view of the system that combines both platform and application monitoring. This paper presents a description of the novel components of the architecture, including new approaches for reaching a higher scalability based on a combination of in-transit processing and performance prediction. This work also includes a practical evaluation on simulated and real platforms, that shows significant monitoring scalability, retrieving data capacity and reduced overheads.
Alberto Cascajo, David E. Singh, Jesús Carretero 0001
PDP1