VLDB 2026 Research / reviewers in the wild / expert
Sergio Iserte
dblp:118/7483
· DBLP profile ↗
22ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-3654-7924ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 8 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Three ways to share a QPU: Scheduling strategies for hybrid Quantum-HPC applications
Marco Cipollini, Simone Rizzo, Sergio Iserte, Paolo Viviani 0001, Giacomo Vitali, Matteo Barbieri, Gabriella Bettonte, Elisabetta Boella, Fulvio Ganz, Roberto Rocco, Orazio Spina, Antonio J. Peña, Petter Sandås, Iacopo Colonnelli, Alberto Scionti, Chiara Vercellino, Emanuele Dri, Jonathan Frassineti, Sara Marzella, Andrea Muratori, Daniele Ottaviani, Olivier Terzo, Bartolomeo Montrucchio, Daniele Gregori |
Future Gener. Comput. Syst. | 3 |
| 2026 | MPI malleability validation under replayed real-world HPC conditions
Sergio Iserte, Maël Madon, Georges Da Costa, Jean-Marc Pierson, Antonio J. Peña |
Future Gener. Comput. Syst. | 1 |
| 2026 | Resource optimization with MPI process malleability for dynamic workloads in HPC clustersabstractDynamic resource management is essential for optimizing computational efficiency in modern high-performance computing (HPC) environments, particularly as systems scale. While research has demonstrated the benefits of malleability in resource management systems (RMS), the adoption of such techniques in production environments remains limited due to challenges in standardization, interoperability, and usability. Addressing these gaps, this paper extends our prior work on the Dynamic Management of Resources (DMR) framework, which provides a modular and user-friendly approach to dynamic resource allocation. Building upon the original DMRlib reconfiguration runtime, this work integrates new methodology from the Malleability Module (MaM) of the Proteo framework, further enhancing reconfiguration capabilities with new spawning strategies and data redistribution methods. In this paper, we explore new malleability strategies in HPC dynamic workloads, such as merging MPI communicators and asynchronous reconfigurations, which offer new opportunities for dramatically reducing memory overhead. The proposed enhancements are rigorously evaluated on a world-class supercomputer, demonstrating improved resource utilization and workload efficiency. Results show that dynamic resource management can reduce the workload completion time by 40% and increase the resource utilization by over 20%, compared to static resource allocation. Sergio Iserte, Iker Martín-Álvarez, Krzysztof Rojek, José Ignacio Aliaga, María Isabel Castillo, Weronika Folwarska, Antonio J. Peña |
Future Gener. Comput. Syst. | 1 |
| 2026 | High-performance computing heterogeneous systems and subsystems
Sergio Iserte, Pedro Valero-Lara, Kevin A. Brown |
Future Gener. Comput. Syst. | 1 |
| 2026 | Exploring the role of Large Language Models in High-Performance Computing programming: A surveyabstractLarge Language Models (LLMs) are emerging as promising assistants in High-Performance Computing (HPC), where programming remains complex and expertise-intensive. This survey systematically reviews their application across five categories: code generation, parallelization and optimization, frameworks and architectures, evaluation and benchmarking, and broader challenges. The analysis highlights both opportunities and limitations: while general-purpose LLMs perform reasonably well on serial and OpenMP-like tasks, they fall short in distributed paradigms such as MPI, where correctness and scalability are critical. Domain-specialized models (e.g., HPC-Coder, HPC-GPT, chatHPC) achieve higher accuracy through fine-tuning, curated datasets, and retrieval-augmented generation (RAG), yet their scope remains narrow and their evaluations largely limited to benchmarks or micro-kernels. The broader picture is one of dual potential and fragility: LLMs can lower barriers to entry, accelerate prototyping, and support code modernization, but they remain brittle under production-level requirements where correctness, performance portability, and scaling cannot be compromised. We conclude that LLMs are unlikely to replace HPC experts in the near term but are positioned to become powerful collaborators in the software development pipeline. Their effective deployment will require richer datasets, integration with performance analysis and schedulers, rigorous evaluation frameworks, and governance structures that ensure transparency and trust. The convergence of AI and HPC should therefore be understood as a long-term, co-evolutionary process—where each advance uncovers new challenges and opportunities for reshaping scientific software development. Strahinja Ljaljevic, Josep Jorba 0001, Sergio Iserte |
Future Gener. Comput. Syst. | 3 |
| 2025 | Dynamic Resource Management in HPC Systems Using Dynamic Processes with PSetsabstractWith the increasing scale of High-Performance Computing (HPC) systems and a new awareness of the environmental impact of HPC, new strategies are required to improve the efficiency of resource usage on these systems. One such strategy is Dynamic Resource Management (DRM), which allows changing the resources assigned to a job dynamically during its execution. This increased flexibility in resource allocation and job scheduling can lead to improvements in several system efficiency metrics. Despite these benefits, DRM has not yet been established as a ready-to-use technology for production HPC systems. This is caused by the significant changes required in all the layers of the HPC system software stack, which are only achievable with an extensive and holistic co-design process between resource management software and applications. In this work, we demonstrate the applicability of a recently introduced, generic design approach for dynamic resources called Dynamic Processes with PSets (DPP), to enable DRM in realworld systems. To this end, we developed an exemplary, dynamic system software stack implementation following the DPP design principles throughout all layers. Based on this, we assess the applicability and performance of our approach using both synthetic benchmarks and job mixes consisting of several dynamic, real-world applications. On up to$\mathbf{1 0 0}$nodes, we measure moderate overheads for process reconfiguration in applications while significantly improving the system throughput and average job turnaround time compared to static scheduling in crowded system scenarios. Dominik Huber, Keerthi Gaddameedi, Tobias Neckel, Hans-Joachim Bungartz, Martin Schulz 0001, Pierre-François Dutot, Olivier Richard, Martin Schreiber 0001, Sergio Iserte, Antonio J. Peña |
HiPC | 9 |
| 2025 | ODOS-MPI: HPC-Friendly SmartNIC Offloading of Computation/Communication KernelsabstractThe increasing complexity and scale of high-performance computing (HPC) workloads demand innovative approaches to optimize both computation and communication. While OpenMP has been widely adopted for intra-node parallelism and MPI for inter-node communication, emerging SmartNICs introduce new opportunities for offloading communication-intensive tasks. In this work, we extend OpenMP to support MPI kernel offloading to SmartNICs. Our implementation integrates Open MPI communication offloading into the LLVM compiler while utilizing DOCA SDK for efficient interaction with Nvidia BlueField DPUs. Leveraging OpenMP eliminates the need for direct low-level programming, lowering the entry barrier for domain scientists. We demonstrate our framework’s versatility by implementing a SmartNIC-enabled version of the MPI OSU micro-benchmarks and improving the execution time of an atmospheric weather simulation by over 18%, thanks to concurrent computation and communication. Mariano Benito, Sergio Iserte, Antonio J. Peña |
SC | 3 |
| 2024 | Proteo: a framework for the generation and evaluation of malleable MPI applicationsabstractAbstract Applying malleability to HPC systems can increase their productivity without degrading or even improving the performance of running applications. This paper presents Proteo, a configurable framework that allows to design benchmarks to study the effect of malleability on a system, and also incorporates malleability into a real application. Proteo consists of two modules: SAM allows to emulate the computational behavior of iterative scientific MPI applications, and MaM is able to reconfigure a job during execution, adjusting the number of processes, redistributing data, and resuming execution. An in-depth study of all the possibilities shows that Proteo is able to behave like a real malleable or non-malleable application in the range [0.85, 1.15]. Furthermore, the different methods defined in MaM for process management and data redistribution are analyzed, concluding that asynchronous malleability, where reconfiguration and application execution overlap, results in a 1.15 $$\times$$ × speedup. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte |
J. Supercomput. | 4 |
| 2024 | Malleability in Modern HPC Systems: Current Experiences, Challenges, and Future OpportunitiesabstractWith the increase of complex scientific simulations driven by workflows and heterogeneous workload profiles, managing system resources effectively is essential for improving performance and system throughput, especially due to trends like heterogeneous HPC and deeply integrated systems with on-chip accelerators. For optimal resource utilization, dynamic resource allocation can improve productivity across all system and application levels, by adapting the applications' configurations to the system's resources. In this context, malleable jobs, which can change resources at runtime, can increase the system throughput and resource utilization while bringing various advantages for HPC users (e.g., shorter waiting time). Malleability has received much attention recently, even though it has been an active research area for almost two decades [1]. This paper presents the state-of-the-art of malleable implementations in HPC systems, targeting mainly malleability in compute and I/O resources. Based on our experiences, we state our current concerns and list future opportunities for research. Ahmad Tarraf, Martin Schreiber 0001, Alberto Cascajo, Jean-Baptiste Besnard, Marc-Andre Vef, Dominik Huber, Sonja Happ, André Brinkmann, David E. Singh, Hans-Christian Hoppe, Alberto Miranda, Antonio J. Peña, Marta Garcia-Gasulla, Martin Schulz 0001, Paul M. Carpenter, Simon Pickartz, Tiberiu Rotaru, Sergio Iserte, Víctor López 0003, Jorge Ejarque, Heena Sirwani, Jesús Carretero 0001, Felix Wolf 0001 |
IEEE Trans. Parallel Distributed Syst. | 19 |
| 2023 | Configurable synthetic application for studying malleability in HPCabstractNowadays, the throughput improvement in large clusters of computers recommends the development of malleable applications. Thus, during the execution of these applications in a job, the resource management system (RMS) can modify its resource allocation, in order to increase the global throughput. There are different alternatives to complete the different steps in which the reallocation of resources is decomposed. To find the best alternatives, this paper introduces a configurable synthetic iterative MPI malleable application capable of modifying, in execution time, the number of MPI processes according to several parameters. The application includes a performance module to measure stages time within steps, from processes management to data redistribution. In this way, the analysis of different scenarios will allow to conclude how the reconfiguration of application has to be made in different circumstances. At the same time, this tool can be used to create workloads that will allow to analyse the impact of malleability on a system and the work in progress. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte |
PDP | 4 |
| 2023 | Optimizing throughput of Seq2Seq model training on the IPU platform for AI-accelerated CFD simulations
Pawel Rosciszewski, Adam Krzywaniak, Sergio Iserte, Krzysztof Rojek, Pawel Gepner |
Future Gener. Comput. Syst. | 3 |
| 2021 | Malleability Implementation in a MPI Iterative MethodabstractIn this poster is evaluated the data redistribution stage for two malleable versions of the Conjugate Gradient. One version is based on synchronous communications, while the other one uses asynchronous communications to overlap computation and data redistribution. Both improve execution time when adding more processes, but there is not a noticeable difference between them, because the asynchronous method lowers the performance of the iterations due to the method’s own communications. When both versions are compared, the synchronous version is preferred when resizing to more processes, while the asynchronous one achieves better times when resizing to fewer processes. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Rafael Mayo 0002, Sergio Iserte |
CLUSTER | 5 |
| 2021 | A Distributed Mesh Generation Study Case through a Customizable Platform as a Service FrameworkabstractConferencia presentada en 11th International Conference on Simulation and Modeling Methodologies, Technologies and Applications - SIMULTECH 2021 Francesc Costa-Majó, Paloma Barreda, Sergio Iserte |
SIMULTECH | 3 |
| 2021 | Improving the management efficiency of GPU workloads in data centers through GPU virtualizationabstractSummary Graphics processing units (GPUs) are currently used in data centers to reduce the execution time of compute‐intensive applications. However, the use of GPUs presents several side effects, such as increased acquisition costs and larger space requirements. Furthermore, GPUs require a nonnegligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. In a similar way to the use of virtual machines, using virtual GPUs may address the concerns associated with the use of these devices. In this regard, the remote GPU virtualization mechanism could be leveraged to share the GPUs present in the computing facility among the nodes of the cluster. This would increase overall GPU utilization, thus reducing the negative impact of the increased costs mentioned before. Reducing the amount of GPUs installed in the cluster could also be possible. However, in the same way as job schedulers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current job schedulers are not able to deal with virtual GPUs. In this paper, we analyze the performance attained by a cluster using the remote Compute Unified Device Architecture middleware and a modified version of the Slurm scheduler, which is now able to assign remote GPUs to jobs. Results show that cluster throughput, measured as jobs completed per time unit, is doubled at the same time that the total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Leveraging teaching on demand: Approaching HPC to undergrads
Sandra Catalán, Rocío Carratalá-Sáez, Sergio Iserte |
J. Parallel Distributed Comput. | 3 |
| 2021 | DMRlib: Easy-Coding and Efficient Resource Management for Job MalleabilityabstractProcess malleability has proved to have a highly positive impact on the resource utilization and global productivity in data centers compared with the conventional static resource allocation policy. However, the non-negligible additional development effort this solution imposes has constrained its adoption by the scientific programming community. In this work, we present DMRlib, a library designed to offer the global advantages of process malleability while providing a minimalist MPI-like syntax. The library includes a series of predefined communication patterns that greatly ease the development of malleable applications. In addition, we deploy several scenarios to demonstrate the positive impact of process malleability featuring different scalability patterns. Concretely, we study two job submission modes (rigid and moldable) in order to identify the best-case scenarios for malleability using metrics such as resource allocation rate, completed jobs per second, and energy consumption. The experiments prove that our elastic approach may improve global throughput by a factor higher than 3x compared to the traditional workloads of non-malleable jobs. Sergio Iserte, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Antonio J. Peña |
IEEE Trans. Computers | 1 |
| 2020 | An study of the effect of process malleability in the energy efficiency on GPU-based clusters
Sergio Iserte, Krzysztof Rojek |
J. Supercomput. | 1 |
| 2018 | DMR API: Improving cluster productivity by turning applications into malleable
Sergio Iserte, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicenç Beltran 0001, Antonio J. Peña |
Parallel Comput. | 1 |
| 2017 | On the benefits of the remote GPU virtualization mechanism: The rCUDA caseabstractSummary Graphics processing units (GPUs) are being adopted in many computing facilities given their extraordinary computing power, which makes it possible to accelerate many general purpose applications from different domains. However, GPUs also present several side effects, such as increased acquisition costs as well as larger space requirements. They also require more powerful energy supplies. Furthermore, GPUs still consume some amount of energy while idle, and their utilization is usually low for most workloads. In a similar way to virtual machines, the use of virtual GPUs may address the aforementioned concerns. In this regard, the remote GPU virtualization mechanism allows an application being executed in a node of the cluster to transparently use the GPUs installed at other nodes. Moreover, this technique allows to share the GPUs present in the computing facility among the applications being executed in the cluster. In this way, several applications being executed in different (or the same) cluster nodes can share 1 or more GPUs located in other nodes of the cluster. Sharing GPUs should increase overall GPU utilization, thus reducing the negative impact of the side effects mentioned before. Reducing the total amount of GPUs installed in the cluster may also be possible. In this paper, we explore some of the benefits that remote GPU virtualization brings to clusters. For instance, this mechanism allows an application to use all the GPUs present in the computing facility. Another benefit of this technique is that cluster throughput, measured as jobs completed per time unit, is noticeably increased when this technique is used. In this regard, cluster throughput can be doubled for some workloads. Furthermore, in addition to increase overall GPU utilization, total energy consumption can be reduced up to 40%. This may be key in the context of exascale computing facilities, which present an important energy constraint. Other benefits are related to the cloud computing domain, where a GPU can be easily shared among several virtual machines. Finally, GPU migration (and therefore server consolidation) is one more benefit of this novel technique. Federico Silla, Sergio Iserte, Carlos Reaño, Javier Prades |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Increasing the Performance of Data Centers by Combining Remote GPU Virtualization with SlurmabstractThe use of Graphics Processing Units (GPUs) presents several side effects, such as increased acquisition costs as well as larger space requirements. Furthermore, GPUs require a non-negligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. Using the virtual GPUs provided by the remote GPU virtualization mechanism may address the concerns associated with the use of these devices. However, in the same way as workload managers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current workload managers are not able to deal with virtual GPUs. In this paper we analyze the performance attained by a cluster using the rCUDA remote GPU virtualization middleware and a modified version of the Slurm workload manager, which is now able to map remote virtual GPUs to jobs. Results show that cluster throughput is doubled at the same time that total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
CCGrid | 1 |
| 2016 | Enabling GPU Virtualization in Cloud EnvironmentsabstractThe use of accelerators, such as graphics processing units (GPUs), to reduce the execution time of compute-intensive applications has become popular during the past few years. These devices increment the computational power of a node thanks to their parallel architecture. This trend has led cloud service providers as Amazon or middlewares such as OpenStack to add virtual machines (VMs) including GPUs to their facilities instances. To fulfill these needs, the guest hosts must be equipped with GPUs which, unfortunately, will be barely utilized if a non GPU-enabled VM is running in the host. The solution presented in this work is based on GPU virtualization and shareability in order to reach an equilibrium between service supply and the applicationsâ?? demand of accelerators. Concretely, we propose to decouple real GPUs from the nodes by using the virtualization technology rCUDA. With this software configuration, GPUs can be accessed from any VM avoiding the need of placing a physical GPUs in each guest host. Moreover, we study the viability of this approach using a public cloud service configuration, and we develop a module for OpenStack in order to add support for the virtualized devices and the logic to manage them. The results demonstrate this is a viable configuration which adds flexibility to current and well-known cloud solutions. Sergio Iserte, Francisco J. Clemente-Castelló, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
CLOSER (2) | 1 |
| 2014 | SLURM Support for Remote GPU Virtualization: Implementation and Performance StudyabstractSLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler. Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades |
SBAC-PAD | 1 |