EDBT 2026 Demo / reviewers in the wild / expert
Sandra Méndez
dblp:122/0962
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-5793-1928ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 15+ years of joint parallel application performance analysis/tools training with Scalasca/Score-P and Paraver/Extrae toolsetsabstractThe diverse landscape of distributed heterogeneous computer systems currently available and being created to address computational challenges with the highest performance requirements presents daunting complexity for application developers. They must effectively decompose and distribute their application functionality and data, efficiently orchestrating the associated communication and synchronisation, on multi/manycore CPU processors with multiple attached acceleration devices structured within compute nodes with interconnection networks of various topologies. Sophisticated compilers, runtime systems and libraries are (loosely) matched with debugging, performance measurement and analysis tools, with proprietary versions by integrators/vendors provided exclusively for their systems complemented by portable (primarily) open-source equivalents developed and supported by the international research community over many years. The Scalasca and Paraver toolsets are two widely employed examples of the latter, installed on personal notebook computers through to the largest leadership HPC systems. Over more than fifteen years their developers have worked closely together in numerous collaborative projects culminating in the creation of a universal parallel performance assessment and optimisation methodology focused on application execution efficiency and scalability, and the associated training and coaching of application developers (often in teams) in its productive use, reviewed in this article with lessons learnt therefrom. Brian J. N. Wylie, Judit Giménez, Christian Feld, Markus Geimer, Germán Llort, Sandra Méndez, Estanislao Mercadal, Anke Visser, Marta García-Gasulla |
Future Gener. Comput. Syst. | 6 |
| 2025 | Parallel I/O analysis in distributed deep learning applications on high-performance computingabstractAbstract Distributed deep learning (DDL) applications generate heavy input/output (I/O) workloads that can create bottlenecks in high-performance computing (HPC) systems. Their optimal I/O configuration depends on factors such as access patterns, storage hardware, dataset size, and execution scale. This study proposes a systematic methodology for characterizing and optimizing I/O behavior in DDL applications, represented through the deep learning I/O benchmark (DLIO), and validated with the real DeepGalaxy application. We evaluate access modes, file formats, and Lustre file system configurations, demonstrating that stripe counts optimized for the access pattern and application scale can reduce I/O and execution times, achieving up to 18 GiB/s of bandwidth and a 5X increase in IOPS. HDF5 provides balanced performance, while TFRecord stands out in bandwidth-intensive scenarios. Shared access minimizes contention and improves scalability in multi-node executions. The results are consolidated into configuration guidelines that offer practical recommendations for practitioners to tune DDL applications for efficient execution in HPC environments. Edixon Párraga, Betzabeth León, Sandra Méndez, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 3 |
| 2022 | A model of checkpoint behavior for applications that have I/OabstractAbstract Due to the increase and complexity of computer systems, reducing the overhead of fault tolerance techniques has become important in recent years. One technique in fault tolerance is checkpointing, which saves a snapshot with the information that has been computed up to a specific moment, suspending the execution of the application, consuming I/O resources and network bandwidth. Characterizing the files that are generated when performing the checkpoint of a parallel application is useful to determine the resources consumed and their impact on the I/O system. It is also important to characterize the application that performs checkpoints, and one of these characteristics is whether the application does I/O. In this paper, we present a model of checkpoint behavior for parallel applications that performs I/O; this depends on the application and on other factors such as the number of processes, the mapping of processes and the type of I/O used. These characteristics will also influence scalability, the resources consumed and their impact on the IO system. Our model describes the behavior of the checkpoint size based on the characteristics of the system and the type (or model) of I/O used, such as the number I/O aggregator processes, the buffering size utilized by the two-phase I/O optimization technique and components of collective file I/O operations. The BT benchmark and FLASH I/O are analyzed under different configurations of aggregator processes and buffer size to explain our approach. The model can be useful when selecting what type of checkpoint configuration is more appropriate according to the applications’ characteristics and resources available. Thus, the user will be able to know how much storage space the checkpoint consumes and how much the application consumes, in order to establish policies that help improve the distribution of resources. Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 2 |
| 2022 | Correction to: A model of checkpoint behavior for applications that have I/O
Betzabeth León, Sandra Méndez, Daniel Franco 0002, Dolores Rexachs, Emilio Luque |
J. Supercomput. | 2 |
| 2020 | Analyzing the Efficiency of Hybrid CodesabstractHybrid parallelization may be the only path for most codes to use HPC systems on a very large scale. Even within a small scale, with an increasing number of cores per node, combining MPI with some shared memory thread-based library allows to reduce the application network requirements. Despite the benefits of a hybrid approach, it is not easy to achieve an efficient hybrid execution. This is not only because of the added complexity of combining two different programming models, but also because in many cases the code was initially designed with just one level of parallelization and later extended to a hybrid mode. This paper presents our model to diagnose the efficiency of hybrid applications, distinguishing the contribution of each parallel programming paradigm. The flexibility of the proposed methodology allows us to use it for different paradigms and scenarios, like comparing the MPI+OpenMP and MPI+CUDA versions of the same code. Judit Giménez, Estanislao Mercadal, Germán Llort, Sandra Méndez |
ISPDC | 4 |
| 2017 | Analyzing the Parallel I/O Severity of MPI ApplicationsabstractPerformance evaluation of parallel applications plays an important role in High Performance Computing (HPC). This is also applied to parallel I/O performance evaluation, which requires understanding the I/O pattern of the application and having knowledge about the performance capacity of the HPC I/O system. In this paper, we present a methodology to evaluate the I/O performance of parallel applications based on the I/O severity degree. We define the I/O severity concept taking into account the I/O requirements of a parallel application, the mapping of I/O processes and the configuration of the I/O subsystem. Requirements are expressed in units denominated I/O phases, which are defined using the temporal and spatial pattern of different files of the application. Our approach is applied to the I/O kernels of scientific applications such as S3DIO, FLASH-IO and BT-IO on the SuperMUC supercomputer. Experimental results show that our methodology allows us to identify if a parallel application is limited by the I/O subsystem and identifying possible root causes of the I/O problems. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CCGrid | 1 |
| 2011 | Methodology for Performance Evaluation of the Input/Output System on Computer ClustersabstractThe increase of processing units, speed and computational power, and the complexity of scientific applications that use high performance computing require more efficient Input/Output (I/O) systems. In order to efficiently use the I/O it is necessary to know its performance capacity to determine if it fulfills applications I/O requirements. This paper proposes a methodology to evaluate I/O performance on computer clusters under different I/O configurations. This evaluation is useful to study how different I/O subsystem configurations will affect the application performance. This approach encompasses the characterization of the I/O system at three different levels: application, I/O system and I/O devices. We select different system configuration and/or I/O operation parameters and we evaluate the impact on performance by considering both the application and the I/O architecture. During I/O configuration analysis we identify configurable factors that have an impact on the performance of the I/O system. In addition, we extract information in order to select the most suitable configuration for the application. Sandra Méndez, Dolores Rexachs, Emilio Luque |
CLUSTER | 1 |