EDBT 2026 Demo / reviewers in the wild / expert
Ester M. Garzón
dblp:09/2896 · also G. Ester Martín Garzón, Gracia Ester Martín Garzón
· DBLP profile ↗
52ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-0568-5470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Theory of computation · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A hybrid quantum-classical approach for liver disease detection using quantum machine learningabstractQuantum Machine Learning (QML) combines principles of quantum computing with traditional Machine Learning (ML) to explore computational advantages in data processing and model efficiency. With the rise of Noisy Intermediate-Scale Quantum (NISQ) devices, hybrid quantum–classical approaches are gaining momentum, especially in domains requiring high precision such as healthcare. In this work, we investigate whether hybrid quantum computing can enhance certain aspects of classical ML, specifically in dataset balancing and the complexity of the neural network involved in training. To this end, we use the Indian Liver Patient Dataset as a case study to determine the presence of liver disease. We present the methodology for developing ‘QML-Liver’, a hybrid approach that seamlessly integrates classical and QML techniques. This includes data preprocessing, model design, and optimal configuration. Our results demonstrate that ‘QML-Liver’ improves key performance metrics, such as accuracy and F1-Score. Additionally, we successfully reduce the number of required qubits to just two, making practical deployment more feasible. These findings underscore the potential of QML for medical diagnostics, particularly in the NISQ era. Laura María Donaire, Gloria Ortega, Francisco José Orts Gómez, Ester M. Garzón, Ernestas Filatovas |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | A quantum-classical hybrid neural network for hate speech detection in Spanishabstract• Hybrid quantum-classical model for Spanish hate speech detection. • Two-phase training stabilizes quantum circuit optimization. • Competitive with transformers; best results on HaterNet dataset. • Consistently outperforms classical and recurrent baselines. • Demonstrates viability of quantum NLP in real-world tasks. Hate speech detection in social media remains a pressing challenge in natural language processing, particularly for languages such as Spanish where annotated resources are limited. This work proposes a hybrid quantum-classical neural architecture that combines bidirectional gated recurrent units with attention and a variational quantum circuit used as a non-linear classifier. The model is trained in two phases: first the recurrent and attention-based layers are optimized to produce stable representations, then these are frozen and a quantum circuit is fine-tuned for classification. Evaluation on two benchmark corpora, HatEval and HaterNet, shows that the proposed hybrid approach achieves competitive performance with strong transformer baselines such as BETO and XLM-R, while consistently outperforming traditional machine learning and recurrent neural models. On HaterNet, the proposed model performs on par with, and in some metrics slightly better than, the transformer baselines, whereas on HatEval it attains slightly lower scores. Its strength lies in detecting hate speech under class imbalance, as reflected in solid F1 scores for the hate speech class. These findings provide an initial empirical assessment of quantum-enhanced NLP in a realistic hate speech detection scenario and suggest promising directions for further study as quantum hardware matures, without constituting evidence of quantum advantage. Francisco José Orts Gómez, Laura María Donaire, Gloria Ortega, Ester M. Garzón |
Expert Syst. Appl. | 4 |
| 2026 | Fast automatic radiotherapy planning via algorithmic improvements and computational accelerationabstractIntensity-Modulated Radiation Therapy enhances dose delivery by dynamically adjusting beam intensities to target tumorous tissues while preserving healthy organs. One of the most effective planning approaches uses the Generalized Equivalent Uniform Dose metric, which ensures high-quality treatment plans but requires tuning several hyperparameters for each anatomical structure. Traditionally, this process is performed manually by clinical experts, making it time-consuming and dependent on human expertise. To address these challenges, a previous method combined multi-objective evolutionary search with gradient-based optimization to automate the tuning process. However, this hybrid strategy incurs high computational cost, as each candidate solution must undergo a complete gradient-based optimization step, repeated thousands of times throughout the process. This study introduces two complementary strategies to improve the efficiency of this framework. First, we analyze alternative multi-objective evolutionary algorithms that converge more rapidly, thereby reducing the number of required function evaluations, and we compare three gradient-based optimization methods to identify the one that accelerates convergence without compromising plan quality. Second, we implement a parallel computing framework that distributes the function evaluations across heterogeneous multicore computing clusters using a static batch scheduling strategy adapted to each node’s computational capacity. Combined, these algorithmic and computational enhancements yield an acceleration factor of 4049 compared to the original implementation. As a result, high-quality radiotherapy treatment plans can be automatically generated in approximately one hour, making this approach viable for integration into time-constrained clinical workflows. Juan José Moreno, Savíns Puertas-Martín, Nelson Garcia Roman, Juana López Redondo, Ester M. Garzón |
Future Gener. Comput. Syst. | 5 |
| 2026 | Low-qubit quantum circuits for efficient integer squaringabstractAbstract Quantum squaring circuits play a critical role in many quantum algorithms; however, most existing designs incur a significant qubit overhead due to the loss of input states and excessive use of ancillary qubits. In this work, we introduce a qubit-efficient quantum circuit for integer squaring that achieves a linear qubit cost of only 3 N qubits for an N -bit input, significantly outperforming state-of-the-art designs that scale quadratically in terms of qubits. Our approach reintegrates the input operand after computation, enabling the uncomputation of intermediate results and efficient recycling of ancilla qubits. This reversible strategy prevents the retention of redundant information, which is a common limitation of prior works. The comparative analysis confirms the scalability and practicality of our design for qubit-constrained quantum hardware, offering a promising solution for arithmetic operations in resource-limited quantum environments. Laura María Donaire, Gloria Ortega, Ester M. Garzón, Ernestas Filatovas, Francisco José Orts Gómez |
J. Supercomput. | 3 |
| 2025 | Problem-Based Learning by Building an Incremental Web ApplicationabstractThis work aims to present an innovative methodology to be followed in the subject “Advanced Computing”, part of the Master's program in Computer Science at the University of Almeria. The methodology seeks to enhance students' engagement with their learning process through modern and effective approaches. Simultaneously, it aims to expand their practical experience through the use of cutting-edge tools for rapid application development, such as Spring Boot and Angular, and High Performance Computing techniques applied using CUDA. The primary methodological approach adopted in this course is based on the flipped classroom methodology. This approach will be implemented in conjunction with a real scientific case involving a physical application of microrheology, providing students with a practical and engaging learning experience. Subsequently, the course delves into tools for developing a web application that serves as a visual interface for the generated data, employing rapid application development techniques. Throughout the course, brief theory blocks will be provided as video lectures uploaded by the professor to explain computational tools and the scientific case. However, classes will primarily focus on practical development, where students are provided with a base project from the outset, allowing them sufficient time to complete it independently. As part of the flipped classroom model, the professor will offer support if needed. Additionally, students will be assigned tasks to implement distinct and straightforward features, fostering both their confidence and creativity as independent computer engineers. The course also incorporates AI-driven programming assistance to promote modern productivity methodologies, avoiding common pitfalls that may limit students' programming skills. J. Navarro-Lázaro, Gloria Ortega, Ester M. Garzón, Francisco José Orts Gómez, Antonio Manuel Puertas |
EDUCON | 3 |
| 2025 | Where High-Performance Computing Meets Radiotherapy for Enhanced Intensity-Modulated Radiation Therapy PlanningabstractABSTRACT Intensity Modulated Radiotherapy (IMRT) employs radiation beams with varying angles and intensities to precisely target cancerous tissues while sparing healthy organs. Planning methods based on the generalized Equivalent Uniform Dose (gEUD) metric achieve excellent Planning Target Volume coverage. However, computing these plans requires extensive parameter adjustments and multiple model evaluations, making the process resource‐intensive and time‐consuming. This study aims to enhance the computational efficiency of radiotherapy plans by automating the adjustment of gEUD parameters, reducing solution times, and facilitating clinical integration. We introduced a novel approach that combines Gradient Descent algorithms with evolutionary optimization to explore the gEUD parameter space. This hybrid methodology generates radiation plans that meet clinical constraints. To address the high computational costs, we implemented parallelization and batching strategies, leveraging multicore servers to accelerate the optimization process and enable real‐time clinical applications. Benchmarking was conducted on three multicore platforms with distinct micro‐architectures, testing various batch sizes and thread configurations. Using a dataset of three Head and Neck IMRT patients treated with nine beams, our approach demonstrated substantial computational speed‐ups. Results confirmed the ability of the method to consistently produce high‐quality radiation therapy plans that meet clinical constraints. By effectively exploiting multicore servers, this approach overcomes the computational challenges of gEUD parameter tuning, enabling its integration into clinical practice. This advancement reduces planning times, supports medical physicists, and ultimately enhances patient care in radiotherapy. Juan José Moreno, Savíns Puertas-Martín, Juana López Redondo, Pilar Martínez Ortigosa, Ester M. Garzón |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | Lowering the cost of quantum comparator circuitsabstractAbstract Quantum comparators hold substantial significance in the scientific community as fundamental components in a wide array of algorithms. In this research, we present an innovative approach where we explore the realm of comparator circuits, specifically focussing on three distinct circuit designs present in the literature. These circuits are notable for their use of T-gates, which have gained significant attention in circuit design due to their ability to enable the utilisation of error-correcting codes. However, it is important to note that T-gates come at a considerable computational cost. One of the key contributions of our work is the optimisation of the quantum gates used within these circuits. We articulate the proposed circuits employing Clifford+T gates, facilitating error correction code implementation. Additionally, we minimise T-gate usage, thereby reducing computational costs and fortifying circuit robustness against errors and environmental disturbances-essential for mitigating the effects of internal and external noise. Our methodology employs a bottom-up examination of comparator circuits, initiating with a detailed study of their gates. Subsequently, we systematically dissect the functions of these gates, thereby advancing towards a comprehensive understanding of the circuit’s overall functionality. This meticulous examination forms the foundation of our research, enabling us to identify areas where optimisations can be made to improve their performance. Laura María Donaire, Gloria Ortega, Ester M. Garzón, Francisco José Orts Gómez |
J. Supercomput. | 3 |
| 2024 | Quantum circuits for computing Hamming distance requiring fewer T gates
Francisco José Orts Gómez, Gloria Ortega, Elías F. Combarro, Ignacio F. Rúa, Ester M. Garzón |
J. Supercomput. | 5 |
| 2023 | Quantum annealing solution for the unrelated parallel machine scheduling with priorities and delay of task switching on machines
Francisco José Orts Gómez, Antonio Manuel Puertas, Gloria Ortega, Ester M. Garzón |
Future Gener. Comput. Syst. | 4 |
| 2023 | Fault-tolerant quantum algorithm for dual-threshold image segmentationabstractAbstract The intrinsic high parallelism and entanglement characteristics of quantum computing have made quantum image processing techniques a focus of great interest. One of the most widely used techniques in image processing is segmentation, which in one of their most basic forms can be carried out using thresholding algorithms. In this paper, a fault-tolerant quantum dual-threshold algorithm has been proposed. This algorithm has been built using only Clifford+T gates for compatibility with error detection and correction codes. Because fault-tolerant implementation of T gates has a much higher cost than other quantum gates, our focus has been on reducing the number of these gates. This has allowed adding noise tolerance, computational cost reduction, and fault tolerance to the state-of-the-art dual-threshold segmentation circuits. Since the dual-threshold image segmentation involves the comparison operation, as part of this work we have implemented two full comparator circuits. These circuits optimize the metrics T-count and T-depth with respect to the best circuit comparators currently available in the literature. Luis O. López, Francisco José Orts Gómez, Gloria Ortega, Vicente González Ruiz, Ester M. Garzón |
J. Supercomput. | 5 |
| 2023 | Efficient design of a quantum absolute-value circuit using Clifford+T gatesabstractAbstract Current quantum computers have a limited number of resources and are heavily affected by internal and external noise. Therefore, small, noise-tolerant circuits are of great interest. With regard to circuit size, it is especially important to reduce the number of required qubits. Concerning to fault-tolerance, circuits entirely built with Clifford+T gates allow the use of error correction codes. However, the T-gate has an excessive cost, so circuits with a high number of T-gates should be avoided. This work focuses on optimising in such terms an operation that is widely used in larger circuits and algorithms: the calculation of the absolute-value of two’s complement encoded integers. The proposed circuit halves the number of required T gates with respect to the best circuit currently available in the literature. Moreover, our circuit requires at least 2 qubits less than the other circuits for such an operation. Francisco José Orts Gómez, Gloria Ortega, Elías F. Combarro, Ignacio F. Rúa, Antonio Manuel Puertas, Ester M. Garzón |
J. Supercomput. | 6 |
| 2022 | On the 2-domination Number of Cylinders with Small CyclesabstractDomination-type parameters are difficult to manage in Cartesian product graphs and there is usually no general relationship between the parameter in both factors and in the product graph. This is the situation of the domination number, the Roman domination number or the $2$-domination number, among others. Contrary to what happens with the domination number and the Roman domination number, the $2$-domination number remains unknown in cylinders, that is, the Cartesian product of a cycle and a path and in this paper, we will compute this parameter in the cylinders with small cycles. We will develop two algorithms involving the $(\min,+)$ matrix product that will allow us to compute the desired values of $\gamma_2(C_n\Box P_m)$, with $3\leq n\leq 15$ and $m\geq 2$. We will also pose a conjecture about the general formulae for the $2$-domination number in this graph class. Comment: 15 pages, 1 figure Ester M. Garzón, José Antonio Martínez, Juan José Moreno, María Luz Puertas |
Fundam. Informaticae | 1 |
| 2022 | HPC acceleration of large (min, +) matrix products to compute domination-type parameters in graphsabstractAbstract The computation of the domination-type parameters is a challenging problem in Cartesian product graphs. We present an algorithmic method to compute the 2-domination number of the Cartesian product of a path with small order and any cycle, involving the $$(\min ,+)$$ ( min , + ) matrix product. We establish some theoretical results that provide the algorithms necessary to compute that parameter, and the main challenge to run such algorithms comes from the large size of the matrices used, which makes it necessary to improve the techniques to handle these objects. We analyze the performance of the algorithms on modern multicore CPUs and on GPUs and we show the advantages over the sequential implementation. The use of these platforms allows us to compute the 2-domination number of cylinders such that their paths have at most 12 vertices. Ester M. Garzón, José Antonio Martínez, Juan José Moreno, María Luz Puertas |
J. Supercomput. | 1 |
| 2022 | HPC enables efficient 3D membrane segmentation in electron tomography
Juan José Moreno, Ester M. Garzón, José-Jesús Fernández, Antonio Martínez-Sánchez |
J. Supercomput. | 2 |
| 2022 | Implementation of three efficient 4-digit fault-tolerant quantum carry lookahead addersabstractAbstract Adders are one of the most interesting circuits in quantum computing due to their use in major algorithms that benefit from the special characteristics of this type of computation. Among these algorithms, Shor’s algorithm stands out, which allows decomposing numbers in a time exponentially lower than the time needed to do it with classical computation. In this work, we propose three fault-tolerant carry lookahead adders that improve the cost in terms of quantum gates and qubits with respect to the rest of quantum circuits available in the literature. Their optimal implementation in a real quantum computer is also presented. Finally, the work ends with a rigorous comparison where the advantages and disadvantages of the proposed circuits against the rest of the circuits of the state of the art are exposed. Moreover, the information obtained from such a comparison is summarized in tables that allow a quick consultation to interested researchers. Francisco José Orts Gómez, Gloria Ortega, Ernestas Filatovas, Ester M. Garzón |
J. Supercomput. | 4 |
| 2021 | Parallel radiation dose computations with GENOCOP III on GPUs
Juan José Moreno, Janusz Miroforidis, Ernestas Filatovas, Ignacy Kaliszewski, Ester M. Garzón |
J. Supercomput. | 5 |
| 2021 | Optimal fault-tolerant quantum comparators for image binarization
Francisco José Orts Gómez, Gloria Ortega, A. C. Cucura, Ernestas Filatovas, Ester M. Garzón |
J. Supercomput. | 5 |
| 2020 | A review on reversible quantum adders
Francisco José Orts Gómez, Gloria Ortega, Elías F. Combarro, Ester M. Garzón |
J. Netw. Comput. Appl. | 4 |
| 2020 | On solving the unrelated parallel machine scheduling problem: active microrheology as a case study
Francisco José Orts Gómez, Gloria Ortega, Antonio Manuel Puertas, Inmaculada García, Ester M. Garzón |
J. Supercomput. | 5 |
| 2019 | Improving the energy efficiency of SMACOF for multidimensional scaling on modern architectures
Francisco José Orts Gómez, Ernestas Filatovas, Gloria Ortega, Olga Kurasova, Ester M. Garzón |
J. Supercomput. | 5 |
| 2018 | TomoEED: fast edge-enhancing denoising of tomographic volumesabstractSummary: TomoEED is an optimized software tool for fast feature-preserving noise filtering of large 3D tomographic volumes on CPUs and GPUs. The tool is based on the anisotropic nonlinear diffusion method. It has been developed with special emphasis in the reduction of the computational demands by using different strategies, from the algorithmic to the high performance computing perspectives. TomoEED manages to filter large volumes in a matter of minutes in standard computers. Availability and implementation: TomoEED has been developed in C. It is available for Linux platforms at http://www.cnb.csic.es/%7ejjfernandez/tomoeed. Supplementary information: Supplementary data are available at Bioinformatics online. Juan José Moreno, Antonio Martínez-Sánchez, José Antonio Martínez, Ester M. Garzón, José-Jesús Fernández |
Bioinform. | 4 |
| 2018 | Improving the performance and energy of Non-Dominated Sorting for evolutionary multiobjective optimization on GPU/CPU platforms
Juan José Moreno, Gloria Ortega, Ernestas Filatovas, José Antonio Martínez, Ester M. Garzón |
J. Glob. Optim. | 5 |
| 2017 | Non-dominated sorting procedure for Pareto dominance ranking on multicore CPU and/or GPU
Gloria Ortega, Ernestas Filatovas, Ester M. Garzón, Leocadio G. Casado |
J. Glob. Optim. | 3 |
| 2017 | An approach to optimise the energy efficiency of iterative computation on integrated GPU-CPU systems
Ester M. Garzón, Juan José Moreno, José Antonio Martínez |
J. Supercomput. | 1 |
| 2017 | Using low-power platforms for Evolutionary Multi-Objective Optimization algorithms
Juan José Moreno, Gloria Ortega, Ernestas Filatovas, José Antonio Martínez, Ester M. Garzón |
J. Supercomput. | 5 |
| 2017 | Accelerating the problem of microrheology in colloidal systems on a GPU
Gloria Ortega, Antonio Manuel Puertas, Ester M. Garzón |
J. Supercomput. | 3 |
| 2016 | GPU Computing to Speed-Up the Resolution of Microrheology Models
Gloria Ortega, Antonio Manuel Puertas, Francisco Javier de las Nieves, Ester M. Garzón |
ICA3PP | 4 |
| 2015 | Parallel resolution of the 3D Helmholtz equation based on multi-graphics processing unit clustersabstractSummary The resolution of the 3D Helmholtz equation is required in the development of models related to a wide range of scientific and technological applications. For solving this equation in complex arithmetic, the biconjugate gradient (BCG) method is one of the most relevant solvers. However, this iterative method has a high computational cost because of the large sparse matrix and the vector operations involved. In this paper, a specific BCG method, adapted for the regularities of the Helmholtz equation is presented. This BCG is based on the implementation of a novel format (named ‘Regular Format’) that allows the storage of the large sparse matrix involved in the sparse matrix vector product in a compact form. The contribution of this work is twofold: (1) decreasing the memory requirements of the 3D Helmholtz equation using the ‘Regular Format’ and (2) speeding up the resolution of the equation using high performance computing resources. A hybrid Message Passing Interface (MPI)‐graphics processing unit CUDA GPU parallelization that is capable of solving complex problems in short time has carried out (Fast‐Helmholtz). Fast‐Helmholtz combines optimizations at Message Passing Interface and GPU levels to reduce communications costs and to improve the exploitation of GPU architecture. This strategy makes it possible to extend the dimension of the Helmholtz problem to be solved, thanks to the relevant reduction of memory requirements and runtime. Copyright © 2014 John Wiley & Sons, Ltd. Gloria Ortega, Julia Lobera, Inmaculada García, María del Pilar Arroyo, Ester M. Garzón |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Exploring the performance-power-energy balance of low-power multicore and manycore architectures for anomaly detection in remote sensing
German Leon, José M. Molero, Ester M. Garzón, Inmaculada García, Antonio Plaza, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2014 | FastSpMM: An Efficient Library for Sparse Matrix Matrix Product on GPUsabstractSparse matrix matrix (SpMM) multiplication is involved in a wide range of scientific and technical applications. The computational requirements for this kind of operation are enormous, especially for large matrices. This paper analyzes and evaluates a method to efficiently compute the SpMM product in a computing environment that includes graphics processing units (GPUs). Some libraries to compute this matricial operation can be found in the literature. However, our strategy (FastSpMM) outperforms the existing approaches because it combines the use of the ELLPACK-R storage format with the exploitation of the high ratio computation/memory access of the SpMM operation and the overlapping of CPU–GPU communications/computations by Compute Unified Device Architecture streaming computation. In this work, FastSpMM is described and its performance evaluated with regard to the CUSPARSE library (supplied by NVIDIA), which also includes routines to compute SpMM on GPUs. Experimental evaluations based on a representative set of test matrices show that, in terms of performance, FastSpMM outperforms the CUSPARSE routine as well as the implementation of the SpMM as a set of sparse matrix vector products. Gloria Ortega, Francisco Vázquez, Inmaculada García, Ester M. Garzón |
Comput. J. | 4 |
| 2014 | A GPU implementation of a hybrid evolutionary algorithm: GPuEGO
J. M. García-Martínez, Ester M. Garzón, Pilar Martínez Ortigosa |
J. Supercomput. | 2 |
| 2014 | High performance computing: an essential tool for science and engineering breakthroughs
José Ranilla, Ester M. Garzón, Jesús Vigo-Aguiar |
J. Supercomput. | 2 |
| 2014 | Performance evaluation of kernel fusion BLAS routines on the GPU: iterative solvers as case study
Siham Tabik, Gloria Ortega, Ester M. Garzón |
J. Supercomput. | 3 |
| 2013 | The BiConjugate gradient method on GPUs
Gloria Ortega, Ester M. Garzón, Francisco Vázquez, Inmaculada García |
J. Supercomput. | 2 |
| 2012 | Dynamic Load Scheduling on CPU-GPU for Iterative Tomographic ReconstructionabstractThis work presents a hybrid computing approach which combines GPUs and multicore processors to fully take advantage of the computing power latent in modern computers. It also presents its application to the problem of tomographic reconstruction. One inherent characteristic of these modern platforms is their heterogeneity, which raises the issue of workload distribution among the different processing elements. Adaptive load balancing techniques are thus necessary to properly adjust the amount of work to be done by each computing element. Here, we have chosen the 'on-demand' strategy, a well-known technique in the HPC field by which the different elements asynchronously request a piece of work when they become idle, thereby keeping the system fairly well balanced. The results show that our scheme accommodates to the heterogeneous platform where it runs as it assigns more work to the faster processing elements automatically, which allows to correctly exploit all the resources available and to get complete reconstructions in less time than pure CPU or GPU approaches. Jose Ignacio Agulleiro Baldo, Francisco Miguel Vázquez López, Ester M. Garzón, José-Jesús Fernández |
ISPA | 3 |
| 2012 | Fast Sparse Matrix Matrix Product Based on ELLR-T and GPU ComputingabstractA wide range of applications in engineering and scientific computing are based on the computation of matrices products, where one of them is sparse. The computational requirements of these operations are very high when dimensions of the matrices increase. The goal of this work is the acceleration of the sparse matrix matrix product (SpMM) on Graphics Processing Units (GPUs). The operation SpMM can be computed by a set of sparse matrix vector operations (SpMV). However, this approach does not reach optimal performance because it cannot benefit from the large value of the ratio computation/memory access associated to the SpMM operation. In this work a routine called FastSpMM is described and its performance evaluated. FastSpMM can be considered as an extension of the ELLRT routine to compute SpMV on GPUs which is based on the ELLPACK-R storage format for sparse matrices. FastSpMM combines the high ratio computation/memory access with the advantages of ELLR-T to exploit the GPU architecture. The CUSPARSE library, supplied by NVIDIA, which also includes routines to compute SpMM on GPUs is used in this work as a reference for performance comparison. Experimental evaluations based on a representative set of test matrices show that FastSpMM outperforms the corresponding CUSPARSE routine in terms of performance. Francisco Vázquez, Gloria Ortega, José-Jesús Fernández, Inmaculada García, Ester M. Garzón |
ISPA | 5 |
| 2012 | Automatic tuning of the sparse matrix vector product on GPUs based on the ELLR-T approach
Francisco Vázquez, José-Jesús Fernández, Ester M. Garzón |
Parallel Comput. | 3 |
| 2011 | Multi-core Desktop Processors Make Possible Real-Time Electron TomographyabstractElectron tomography (ET) allows elucidation of the three-dimensional (3D) structure of large complex biological specimens at molecular resolution. In order to achieve such resolution levels, large projection images have to be used to compute the 3D reconstructions. Tomographic reconstruction on this scale requires a tremendous use of computational resources and considerable processing time. Traditionally, parallel and distributed systems, and more recently GPUs, have been the key to cope with this demanding procedure. This work demonstrates that full exploitation of the impressive processing power within modern multi-core processors make them a feasible alternative. The use of parallel computing, vectorization and code optimization allows ultra-fast tomographic reconstructions on standard computers, even outperforming GPUs. Our results confirm that modern processors succeed in providing reconstructed volumes in very little time, which enables them for real-time ET. Jose Ignacio Agulleiro Baldo, Ester M. Garzón, Inmaculada García, José-Jesús Fernández |
PDP | 2 |
| 2011 | Matrix Implementation of Simultaneous Iterative Reconstruction Technique (SIRT) on GPUsabstractElectron tomography (ET) is an important technique in biosciences that is providing new insights into the cellular ultrastructure. Iterative reconstruction methods have been shown to be robust against the noise and limited-tilt range conditions present in ET. Nevertheless, these methods are not extensively used due to their computational demands. Instead, the simpler method weighted backprojection (WBP) remains prevalent. Recently, we have demonstrated that a matrix approach to WBP allows a significant reduction in processing time both on central processing units and on graphics processing units (GPUs). In this work, we extend that matrix approach to one of the most common iterative methods in ET, simultaneous iterative reconstruction technique (SIRT). We show that it is possible to implement this method targeted at GPU directly, using sparse algebra. We also analyse this approach on different GPU platforms and confirm that these implementations exhibit high performance. This may thus help to the widespread use of SIRT. Francisco Vázquez, Ester M. Garzón, José-Jesús Fernández |
Comput. J. | 2 |
| 2011 | A new approach for sparse matrix vector product on NVIDIA GPUsabstractAbstract The sparse matrix vector product (SpMV) is a key operation in engineering and scientific computing and, hence, it has been subjected to intense research for a long time. The irregular computations involved in SpMV make its optimization challenging. Therefore, enormous effort has been devoted to devise data formats to store the sparse matrix with the ultimate aim of maximizing the performance. Graphics Processing Units (GPUs) have recently emerged as platforms that yield outstanding acceleration factors. SpMV implementations for NVIDIA GPUs have already appeared on the scene. This work proposes and evaluates a new implementation of SpMV for NVIDIA GPUs based on a new format, ELLPACK‐R, that allows storage of the sparse matrix in a regular manner. A comparative evaluation against a variety of storage formats previously proposed has been carried out based on a representative set of test matrices. The results show that, although the performance strongly depends on the specific pattern of the matrix, the implementation based on ELLPACK‐R achieves higher overall performance. Moreover, a comparison with standard state‐of‐the‐art superscalar processors reveals that significant speedup factors are achieved with GPUs. Copyright © 2010 John Wiley & Sons, Ltd. Francisco Vázquez, José-Jesús Fernández, Ester M. Garzón |
Concurr. Comput. Pract. Exp. | 3 |
| 2011 | Adaptive load balancing of iterative computation on heterogeneous nondedicated systems
José Antonio Martínez, Francisco Almeida, Ester M. Garzón, Alejandro Acosta, Vicente Blanco 0001 |
J. Supercomput. | 3 |
| 2011 | Automatic tuning of iterative computation on heterogeneous multiprocessors with ADITHE
José Antonio Martínez, Ester M. Garzón, Antonio Plaza, Inmaculada García |
J. Supercomput. | 2 |
| 2011 | Fast anomaly detection in hyperspectral images with RX method on heterogeneous clusters
José M. Molero, Abel Paz, Ester M. Garzón, José Antonio Martínez, Antonio Plaza, Inmaculada García |
J. Supercomput. | 3 |
| 2010 | Ultra-fast Tomographic Reconstruction with a Highly Optimized Weighted Back-Projection AlgorithmabstractElectron tomography (ET) allows elucidation of the three-dimensional (3D) structure of large complex biological specimens at molecular resolution. In order to achieve such resolution levels, large projection images have to be used to compute the 3D reconstructions. Tomographic reconstruction on this scale requires a tremendous use of computational resources and a considerable processing time. In this work, we present and evaluate a highly optimized implementation of the Weighted Back-Projection reconstruction algorithm. Briefly, optimizations made to the code comprise (1) vector processing with SSE (Streaming SIMD Extensions) instructions, (2) an efficient use of cache memory, (3) to take advantage of the inherent image symmetry, (4) to use the FFTW (Fastest Fourier Transform in the West) library for image filtering, (5) to use regions of interest and last, but not least, (6) a wide range of minor optimizations like some data pre-calculations or an instruction level parallelism improvement. We have evaluated the method on tomographic reconstructions of several datasets and on two computing platforms. The results show that our version speeds up the method by a factor around 14 or 16, depending on the platform. Jose Ignacio Agulleiro Baldo, Ester M. Garzón, Inmaculada García, José-Jesús Fernández |
PDP | 2 |
| 2008 | Fast Tomographic Reconstruction with Vectorized BackprojectionabstractElectron tomography allows elucidation of the three-dimensional (3D) structure of large complex biological specimens at molecular resolution. In order to achieve such resolution levels, large projection images have to be used to compute the 3D reconstructions. Tomographic reconstruction on this scale requires a tremendous use of computational resources and considerable processing time. In this work, we present and evaluate a vector approach for fast 3D reconstruction that takes advantage of the multimedia extensions in modern processors. We have implemented the standard 3D reconstruction method, weighted backprojection, using the Streaming SIMD Extensions (SSE). We have evaluated the method on tomographic reconstruction of several datasets of various sizes on a computing platform based on Intel Xeon processor. The results show that our approach speeds up the method by a factor of 3. Jose Ignacio Agulleiro Baldo, Ester M. Garzón, Inmaculada García, José-Jesús Fernández |
PDP | 2 |
| 2007 | Three-Dimensional Bursting Simulation on Two Parallel Systems
Siham Tabik, Luis F. Romero, Ester M. Garzón, Inmaculada García, Juan I. Ramos 0001 |
ICCSA (3) | 3 |
| 2006 | Evaluation of Parallel Paradigms on Anisotropic Nonlinear Diffusion
Siham Tabik, Ester M. Garzón, Inmaculada García, José-Jesús Fernández |
Euro-Par | 2 |
| 2006 | Multiprocessing of Anisotropic Nonlinear Diffusion for Filtering 3D ImagesabstractThis article describes and analyzes the parallelization of the Anisotropic Nonlinear Diffusion (AND) for filtering 3D images. AND is one of the most powerful denoising techniques in the field of computer vision. This technique consists in resolving the equation of diffusion tightly coupled with a massive set of eigensystems. Denoising large 3D images in biomedicine and structural cellular biology by AND has a high computational cost. In this work, we propose a portable and efficient parallel implementation of AND based on a hybrid paradigm that combines (1) the message passing model and (2)the shared address space model. The proposed parallel implementation has been evaluated on a cluster of SMPs based on a UMA Uniform Memory access. The evaluation results show that the hybrid model is more suitable for this kind of platforms. Siham Tabik, Ester M. Garzón, Inmaculada García, José-Jesús Fernández |
PDP | 2 |
| 2005 | Approaches Based on Permutations for Partitioning Sparse Matrices on Multiprocessors
Ester M. Garzón, Inmaculada García |
J. Supercomput. | 1 |
| 2003 | A VHDL Library to Analyse Fault Tolerant Techniques
Pilar Martínez Ortigosa, O. López, R. Estrada, Inmaculada García, Ester M. Garzón |
FPL | 5 |
| 2003 | Floating point arithmetic teaching for computational science
José-Jesús Fernández, Inmaculada García, Ester M. Garzón |
Future Gener. Comput. Syst. | 3 |
| 1996 | Parallel Implementation of the Lanczos Method for Sparse Matrices: Analysis of Data DistributionsabstractIn this paper an efficient parallel implementation of the Lanczos method for matrix tridlagonalization is described.This work is restricted to sparse matrices where the work load unbalance problem must be solved in order to optimize the efficiency of the parallel implementation.A new strategy for data dktribution (called Pivoting Block) is purposed in this work. Efficiency of the parallelLanczos method for sparse matrices using Pivoting Block is slightly improved compared to others data dktribution strategies. Ester M. Garzón, Inmaculada García |
International Conference on Supercomputing | 1 |