EDBT 2026 Demo / reviewers in the wild / expert
María Isabel Castillo
dblp:c/MaribelCastillo · also Maribel Castillo
· DBLP profile ↗
20ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-2826-3086ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resource optimization with MPI process malleability for dynamic workloads in HPC clustersabstractDynamic resource management is essential for optimizing computational efficiency in modern high-performance computing (HPC) environments, particularly as systems scale. While research has demonstrated the benefits of malleability in resource management systems (RMS), the adoption of such techniques in production environments remains limited due to challenges in standardization, interoperability, and usability. Addressing these gaps, this paper extends our prior work on the Dynamic Management of Resources (DMR) framework, which provides a modular and user-friendly approach to dynamic resource allocation. Building upon the original DMRlib reconfiguration runtime, this work integrates new methodology from the Malleability Module (MaM) of the Proteo framework, further enhancing reconfiguration capabilities with new spawning strategies and data redistribution methods. In this paper, we explore new malleability strategies in HPC dynamic workloads, such as merging MPI communicators and asynchronous reconfigurations, which offer new opportunities for dramatically reducing memory overhead. The proposed enhancements are rigorously evaluated on a world-class supercomputer, demonstrating improved resource utilization and workload efficiency. Results show that dynamic resource management can reduce the workload completion time by 40% and increase the resource utilization by over 20%, compared to static resource allocation. Sergio Iserte, Iker Martín-Álvarez, Krzysztof Rojek, José Ignacio Aliaga, María Isabel Castillo, Weronika Folwarska, Antonio J. Peña |
Future Gener. Comput. Syst. | 5 |
| 2025 | Efficient quantum circuit contraction using tensor decision diagramsabstractAbstract Simulating quantum circuits efficiently on classical computers is crucial given the limitations of current noisy intermediate-scale quantum devices. This paper adapts and extends two methods used to contract tensor networks within the fast tensor decision diagram (FTDD) framework. The methods, called iterative pairing and block contraction, exploit the advantages of tensor decision diagrams to reduce both the temporal and spatial cost of quantum circuit simulations. The iterative pairing method minimizes intermediate diagram sizes, while the block contraction algorithm efficiently handles circuits with repetitive structures, such as those found in quantum walks and Grover’s algorithm. Experimental results demonstrate that, in some cases, these methods significantly outperform traditional contraction orders like sequential and cotengra in terms of both memory usage and execution time. Furthermore, simulation tools based on decision diagrams, such as FTDD, show superior performance to matrix-based simulation tools, such as Google tensor networks, enabling the simulation of larger circuits more efficiently. These findings show the potential of decision diagram-based approaches to improve the simulation of quantum circuits on classical platforms. Vicente Lopez-Oliva, José M. Badía, María Isabel Castillo |
J. Supercomput. | 3 |
| 2025 | A community detection-based parallel algorithm for quantum circuit simulation using tensor networksabstractAbstract Quantum computing holds significant promise for solving complex problems, but simulating quantum circuits on classical computers remains essential due to the current limitations of quantum hardware. Efficient simulation is crucial for the development and validation of quantum algorithms and quantum computers. This paper explores and compares various strategies to leverage different levels of parallelism to accelerate the contraction of tensor networks representing large quantum circuits. We propose a new parallel multistage algorithm based on communities. The original tensor network is partitioned into several communities, which are then contracted in parallel. The pairs of tensors of the resulting network can be contracted in parallel using a GPU. We use the Girvan–Newman algorithm to obtain the communities and the contraction plans. We compare the new algorithm with two other parallelisation strategies: one based on contracting all the pairs of tensors in the GPU and another one that uses slicing to cut some indexes of the tensor network and then MPI processes to contract the resulting slices in parallel. The new parallel algorithm gets the best results with different well-known quantum circuits with a high degree of entanglement, including random quantum circuits. In conclusion, the results show that the main factor that limits the simulation is the space cost. However, the parallel multistage algorithm manages to reduce the cost of sequential simulation for circuits with a high number of qubits and allows simulating larger circuits. Alfred M. Pastor, José M. Badía, María Isabel Castillo |
J. Supercomput. | 3 |
| 2024 | Proteo: a framework for the generation and evaluation of malleable MPI applicationsabstractAbstract Applying malleability to HPC systems can increase their productivity without degrading or even improving the performance of running applications. This paper presents Proteo, a configurable framework that allows to design benchmarks to study the effect of malleability on a system, and also incorporates malleability into a real application. Proteo consists of two modules: SAM allows to emulate the computational behavior of iterative scientific MPI applications, and MaM is able to reconfigure a job during execution, adjusting the number of processes, redistributing data, and resuming execution. An in-depth study of all the possibilities shows that Proteo is able to behave like a real malleable or non-malleable application in the range [0.85, 1.15]. Furthermore, the different methods defined in MaM for process management and data redistribution are analyzed, concluding that asynchronous malleability, where reconfiguration and application execution overlap, results in a 1.15 $$\times$$ × speedup. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte |
J. Supercomput. | 3 |
| 2023 | Configurable synthetic application for studying malleability in HPCabstractNowadays, the throughput improvement in large clusters of computers recommends the development of malleable applications. Thus, during the execution of these applications in a job, the resource management system (RMS) can modify its resource allocation, in order to increase the global throughput. There are different alternatives to complete the different steps in which the reallocation of resources is decomposed. To find the best alternatives, this paper introduces a configurable synthetic iterative MPI malleable application capable of modifying, in execution time, the number of MPI processes according to several parameters. The application includes a performance module to measure stages time within steps, from processes management to data redistribution. In this way, the analysis of different scenarios will allow to conclude how the reconfiguration of application has to be made in different circumstances. At the same time, this tool can be used to create workloads that will allow to analyse the impact of malleability on a system and the work in progress. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte |
PDP | 3 |
| 2021 | Malleability Implementation in a MPI Iterative MethodabstractIn this poster is evaluated the data redistribution stage for two malleable versions of the Conjugate Gradient. One version is based on synchronous communications, while the other one uses asynchronous communications to overlap computation and data redistribution. Both improve execution time when adding more processes, but there is not a noticeable difference between them, because the asynchronous method lowers the performance of the iterations due to the method’s own communications. When both versions are compared, the synchronous version is preferred when resizing to more processes, while the asynchronous one achieves better times when resizing to fewer processes. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Rafael Mayo 0002, Sergio Iserte |
CLUSTER | 3 |
| 2017 | Accelerating FaST-LMM for Epistasis Tests
Héctor Martínez 0002, Sergio Barrachina 0001, María Isabel Castillo, Enrique S. Quintana-Ortí, Jordi Rambla De Argila, Xavier Farré, Arcadi Navarro |
ICA3PP | 3 |
| 2015 | Unveiling the performance-energy trade-off in iterative linear system solvers for multithreaded processorsabstractSummary In this paper, we analyze the interactions occurring in the triangle performance‐power‐energy for the execution of a pivotal numerical algorithm, the iterative conjugate gradient (CG) method, on a diverse collection of parallel multithreaded architectures. This analysis is especially timely in a decade where the power wall has arisen as a major obstacle to build faster processors. Moreover, the CG method has recently been proposed as a complement to the LINPACK benchmark, as this iterative method is argued to be more archetypical of the performance of today's scientific and engineering applications. To gain insights about the benefits of hands‐on optimizations we include runtime and energy efficiency results for both out‐of‐the‐box usage relying exclusively on compiler optimizations, and implementations manually optimized for target architectures, that range from general‐purpose and digital signal multicore processors to manycore graphics processing units, all representative of current multithreaded systems. Copyright © 2014 John Wiley & Sons, Ltd. José Ignacio Aliaga, Hartwig Anzt, María Isabel Castillo, Juan Carlos Fernández 0002, German Leon, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Out-of-core macromolecular simulations on multithreaded architecturesabstractSummary We address the solution of large‐scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multithreaded platforms, consisting of multicore processors and possibly a graphics processor (graphics processing unit). In particular, we compare specialized implementations of several high‐performance eigensolvers that, by relying on disk storage and out‐of‐core techniques, can in principle tackle the large memory requirements of these biological problems, which in general do not fit into the main memory of current desktop machines. All these out‐of‐core eigensolvers, except for one, are composed of compute‐bound (i.e., arithmetically intensive) operations, which we accelerate by exploiting the performance of current multicore processors and, in some cases, by additionally off‐loading certain parts of the computation to a graphics processing unit accelerator. One of the eigensolvers is a memory‐bound algorithm, which strongly constrains its performance when the data is on disk. However, this method exhibits a much lower arithmetic cost compared with its compute‐bound alternatives for this particular application. Experimental results on a desktop platform, representative of current server technology, illustrate the potential of these methods to address the simulation of biological activity. Copyright © 2014 John Wiley & Sons, Ltd. José Ignacio Aliaga, José M. Badía, María Isabel Castillo, Davor Davidovic, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Concurrent and Accurate Short Read Mapping on Multicore ProcessorsabstractWe introduce a parallel aligner with a work-flow organization for fast and accurate mapping of RNA sequences on servers equipped with multicore processors. Our software, HPG Aligner SA (HPG Aligner SA is an open-source application. The software is available at http://www.opencb.org, exploits a suffix array to rapidly map a large fraction of the RNA fragments (reads), as well as leverages the accuracy of the Smith-Waterman algorithm to deal with conflictive reads. The aligner is enhanced with a careful strategy to detect splice junctions based on an adaptive division of RNA reads into small segments (or seeds), which are then mapped onto a number of candidate alignment locations, providing crucial information for the successful alignment of the complete reads. The experimental results on a platform with Intel multicore technology report the parallel performance of HPG Aligner SA, on RNA reads of 100-400 nucleotides, which excels in execution time/sensitivity to state-of-the-art aligners such as TopHat 2+Bowtie 2, MapSplice, and STAR. Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2013 | A dynamic pipeline for RNA sequencing on multicore processorsabstractWe present a concurrent algorithm for mapping short and long RNA sequences on multicore processors. Our solution processes the data, initially stored on disk, in batches of reads which are passed between the consecutive stages of a pipeline. A major operational reorganization of the original static pipeline, combined with a complete reimplementation based on POSIX threads, renders a dissociated execution between threads and stages/task types, so that threads can compute any type of pending task resulting in a dynamic pipeline. The experiments on a multicore platform reveal that this reorganization yields significantly higher performance, specially for architectures equipped with a small to moderate number of cores. Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí |
EuroMPI | 5 |
| 2012 | Analysis of Strategies to Save Energy for Message-Passing Dense Linear Algebra KernelsabstractIn this paper we analyze the impact that energy-saving strategies, like the application of DVFS via Linux governors and the MPI communication mode, have on the performance and energy consumption of message-passing dense linear algebra operations. In the study, we employ codes from ScaLAPACK for three matrix kernels, the matrix-matrix and matrix-vector products and the Cholesky factorization, which exhibit different levels of concurrency and CPU/memory activity. Following a recent trend, we also include an accelerated version of the matrix-matrix product that off-loads all computation to a graphics processor and study the energy gains of this hybrid solver when the general-purpose cores of the system are promoted to a low consuming mode. Experimental results on a cluster equipped with state-of-the-art computation and communication hardware illustrate the results of this study. María Isabel Castillo, Juan Carlos Fernández 0002, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicente Roca |
PDP | 1 |
| 2009 | Exploiting the capabilities of modern GPUs for dense matrix computationsabstractAbstract We present several algorithms to compute the solution of a linear system of equations on a graphics processor (GPU), as well as general techniques to improve their performance, such as padding and hybrid GPU‐CPU computation. We compare single and double precision performance of a modern GPU with unified architecture, and show how iterative refinement with mixed precision can be used to regain full accuracy in the solution of linear systems, exploiting the potential of the processor for single precision arithmetic. Experimental results on a GTX280 using CUBLAS 2.0, the implementation of BLAS for NVIDIA® GPUs with unified architecture, illustrate the performance of the different algorithms and techniques proposed. Copyright © 2009 John Wiley & Sons, Ltd. Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
J. Supercomput. | 6 |
| 2008 | Solving Dense Linear Systems on Graphics Processors
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 2 |
| 2008 | Evaluation and tuning of the Level 3 CUBLAS for graphics processorsabstractThe increase in performance of the last generations of graphics processors (GPUs) has made this class of platform a coprocessing tool with remarkable success in certain types of operations. In this paper we evaluate the performance of the Level 3 operations in CUBLAS, the implementation of BIAS for NVIDIAreg GPUs with unified architecture. From this study, we gain insights on the quality of the kernels in the library and we propose several alternative implementations that are competitive with those in CUBLAS. Experimental results on a GeForce 8800 Ultra compare the performance of CUBLAS and the new variants. Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
IPDPS | 2 |
| 2007 | Stabilizing large-scale generalized systems on parallel computers using multithreading and message-passingabstractAbstract We discuss the parallelization of an efficient algorithm for the partial stabilization of large‐scale linear control systems in generalized state‐space form. The algorithm is composed of highly parallel iterative schemes that appear in the computation of certain matrix functions. Here we evaluate different approaches to exploit parallelism at two levels, based on threads and processes. Our experimental results on a cluster of symmetric multiprocessors and a CC‐NUMA platform show that the efficiency of the matrix operations underlying the iterative schemes carry over to the parallel implementation of the stabilization algorithm. Copyright © 2006 John Wiley & Sons, Ltd. Peter Benner, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | Parallelization of GSL: The Web Service InterfaceabstractWe present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements. José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
PDP | 4 |
| 2001 | Efficient Algorithms for the Block Hessenberg Form
Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, María Isabel Castillo, Vicente Hernández |
J. Supercomput. | 3 |
| 2000 | Parallel Partial Stabilizing Algorithms for Large Linear Control Systems
Peter Benner, María Isabel Castillo, Enrique S. Quintana-Ortí, Vicente Hernández |
J. Supercomput. | 2 |