EDBT 2026 Demo / reviewers in the wild / expert
Juan A. Rico-Gallego
dblp:41/1939 · also Juan-Antonio Rico-Gallego
· DBLP profile ↗
20ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0002-4264-7473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Energy Efficiency in a Data Center: PUE Analyzing and TuningabstractIn the digital era, energy efficiency in data centers is crucial due to the exponential growth of data and the increasing demand for technological infrastructures. Power Usage Effectiveness (PUE) is a key indicator for evaluating this energy efficiency, measuring the ratio between the total energy consumption of a data center and the energy used by information technology (IT) equipment. Integrating sensors to monitor key variables provides a detailed and comprehensive view of data center operations. However, analyzing these large volumes of complex data requires advanced processing approaches. In this context, machine learning technologies play a decisive role, as learning algorithms can identify hidden patterns and correlations that would be difficult to detect with traditional methods. This research presents a step-by-step methodology to optimize data center operations by combining advanced sensor technologies and machine learning. It identifies key variables, integrates sensors to monitor them, and analyzes the data to reveal hidden patterns that traditional methods may miss. This approach enables realtime, data-driven decisions, improving efficiency, reducing energy consumption, and optimizing PUE. Validated through a real use case, the methodology demonstrates its potential to enhance energy management and promote sustainability in data centers. Daniel Flores-Martin, Miguel Mahillo, Felipe Lemus-Prieto, Javier Corral-García, Juan A. Rico-Gallego |
CCGrid | 5 |
| 2025 | Spiking Neuron-Astrocyte Networks for Image RecognitionabstractFrom biological and artificial network perspectives, researchers have started acknowledging astrocytes as computational units mediating neural processes. Here, we propose a novel biologically inspired neuron-astrocyte network model for image recognition, one of the first attempts at implementing astrocytes in spiking neuron networks (SNNs) using a standard data set. The architecture for image recognition has three primary units: the preprocessing unit for converting the image pixels into spiking patterns, the neuron-astrocyte network forming bipartite (neural connections) and tripartite synapses (neural and astrocytic connections), and the classifier unit. In the astrocyte-mediated SNNs, an astrocyte integrates neural signals following the simplified Postnov model. It then modulates the integrate-and-fire (IF) neurons via gliotransmission, thereby strengthening the synaptic connections of the neurons within the astrocytic territory. We develop an architecture derived from a baseline SNN model for unsupervised digit classification. The spiking neuron-astrocyte networks (SNANs) display better network performance with an optimal variance-bias trade-off than SNN alone. We demonstrate that astrocytes promote faster learning, support memory formation and recognition, and provide a simplified network architecture. Our proposed SNAN can serve as a benchmark for future researchers on astrocyte implementation in artificial networks, particularly in neuromorphic systems, for its simplified design. Jhunlyn Lorenzo, Juan A. Rico-Gallego, Stéphane Binczak, Sabir Jacquir |
Neural Comput. | 2 |
| 2024 | Federated learning meets remote sensingabstractRemote sensing (RS) imagery provides invaluable insights into characterizing the Earth’s land surface within the scope of Earth observation (EO). Technological advances in capture instrumentation, coupled with the rise in the number of EO missions aimed at data acquisition, have significantly increased the volume of accessible RS data. This abundance of information has alleviated the challenge of insufficient training samples, a common issue in the application of machine learning (ML) techniques. In this context, crowd-sourced data play a crucial role in gathering diverse information from multiple sources, resulting in heterogeneous datasets that enable applications to harness a more comprehensive spatial coverage of the surface. However, the sensitive nature of RS data requires ensuring the privacy of the complete collection. Consequently, federated learning (FL) emerges as a privacy-preserving solution, allowing collaborators to combine such information from decentralized private data collections to build efficient global models. This paper explores the convergence between the FL and RS domains, specifically in developing data classifiers. To this aim, an extensive set of experiments is conducted to analyze the properties and performance of novel FL methodologies. The main emphasis is on evaluating the influence of such heterogeneous and disjoint data among collaborating clients. Moreover, scalability is evaluated for a growing number of clients, and resilience is assessed against Byzantine attacks. Finally, the work concludes with future directions and serves as the opening of a new research avenue for developing efficient RS applications under the FL paradigm. The source code is publicly available at https://github.com/hpc-unex/FLmeetsRS. Sergio Moreno-Álvarez, Mercedes Eugenia Paoletti, Andres Jesus Sanchez, Juan A. Rico-Gallego, Lirong Han, Juan Mario Haut |
Expert Syst. Appl. | 4 |
| 2022 | Optimizing Distributed Deep Learning in Heterogeneous Computing Platforms for Remote Sensing Data ClassificationabstractApplications from Remote Sensing (RS) unveiled unique challenges to Deep Learning (DL) due to the high volume and complexity of their data. On the one hand, deep neural network architectures have the capability to automatically ex-tract informative features from RS data. On the other hand, these models have massive amounts of tunable parameters, re-quiring high computational capabilities. Distributed DL with data parallelism on High-Performance Computing (HPC) sys-tems have proved necessary in dealing with the demands of DL models. Nevertheless, a single HPC system can be al-ready highly heterogeneous and include different computing resources with uneven processing power. In this context, a standard data parallelism strategy does not partition the data efficiently according to the available computing resources. This paper proposes an alternative approach to compute the gradient, which guarantees that the contribution to the gradi-ent calculation is proportional to the processing speed of each DL model's replica. The experimental results are obtained in a heterogeneous HPC system with RS data and demon-strate that the proposed approach provides a significant training speed up and gain in the global accuracy compared to one of the state-of-the-art distributed DL framework. Sergio Moreno-Álvarez, Mercedes Eugenia Paoletti, Juan A. Rico-Gallego, Gabriele Cavallaro, Juan Mario Haut |
IGARSS | 3 |
| 2022 | Model-based selection of optimal MPI broadcast algorithms for multi-core clustersabstractThe performance of collective communication operations determines the overall performance of MPI applications. Different algorithms have been developed and implemented for each MPI collective operation, but none proved superior in all situations. Therefore, MPI implementations have to solve the problem of selecting the optimal algorithm for the collective operation depending on the platform, the number of processes involved, the message size(s), etc. The current solution method is purely empirical. Recently, an alternative solution method using analytical performance models of collective algorithms has been proposed and proved both accurate and efficient for one-process-per-CPU configurations. The method derives the analytical performance models of algorithms from their code implementation rather than from high-level mathematical definitions, and estimates the parameters of the models separately for each algorithm. The method is network and topology oblivious and uses the Hockney model for point-to-point communications. In this paper, we extend that selection method to the case of clusters of multi-core processors, where each core of the platform runs a process of the MPI application. We present the proposed approach using Open MPI broadcast algorithms, and experimentally validate it on three different clusters of multi-core processors, Grisou, Gros and MareNostrum4. Emin Nuriyev, Juan A. Rico-Gallego, Alexey L. Lastovetsky |
J. Parallel Distributed Comput. | 2 |
| 2022 | Remote Sensing Image Classification Using CNNs With Balanced Gradient for Distributed Heterogeneous ComputingabstractLand-cover classification methods are based on the processing of large image volumes to accurately extract representative features. Particularly, convolutional models provide notable characterization properties for image classification tasks. Distributed learning mechanisms on high performance computing platforms have been proposed to speed up the processing, whilst achieving an efficient feature extraction. High performance computing platforms are commonly composed of a combination of CPUs and GPUs, with different computational capabilities. As a result, current homogeneous workload distribution techniques for deep learning become obsolete due to their inefficient use of computational resources. To address this, new computational balancing proposals, such as heterogeneous data parallelism, have been implemented. Nevertheless, these techniques should be improved to handle the peculiarities of working with heterogeneous data workloads in the training of distributed deep learning models. The objective of handling heterogeneous workloads for current platforms motivates the development of this work. This paper proposes an innovative heterogeneous gradient calculation applied to land-cover classification tasks through convolutional models, considering the data amount assigned to each device in the platform whilst maintaining the acceleration. Extensive experimentation has been conducted on multiple datasets, considering different deep models on heterogeneous platforms to demonstrate the performance of the proposed methodology. Sergio Moreno-Álvarez, Mercedes Eugenia Paoletti, Gabriele Cavallaro, Juan A. Rico-Gallego, Juan Mario Haut |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Heterogeneous gradient computing optimization for scalable deep neural networksabstractAbstract Nowadays, data processing applications based on neural networks cope with the growth in the amount of data to be processed and with the increase in both the depth and complexity of the neural networks architectures, and hence in the number of parameters to be learned. High-performance computing platforms are provided with fast computing resources, including multi-core processors and graphical processing units, to manage such computational burden of deep neural network applications. A common optimization technique is to distribute the workload between the processes deployed on the resources of the platform. This approach is known as data-parallelism. Each process, known as replica, trains its own copy of the model on a disjoint data partition. Nevertheless, the heterogeneity of the computational resources composing the platform requires to unevenly distribute the workload between the replicas according to its computational capabilities, to optimize the overall execution performance. Since the amount of data to be processed is different in each replica, the influence of the gradients computed by the replicas in the global parameter updating should be different. This work proposes a modification of the gradient computation method that considers the different speeds of the replicas, and hence, its amount of data assigned. The experimental results have been conducted on heterogeneous high-performance computing platforms for a wide range of models and datasets, showing an improvement in the final accuracy with respect to current techniques, with a comparable performance. Sergio Moreno-Álvarez, Mercedes Eugenia Paoletti, Juan A. Rico-Gallego, Juan Mario Haut |
J. Supercomput. | 3 |
| 2021 | Heterogeneous model parallelism for deep neural networks
Sergio Moreno-Álvarez, Juan Mario Haut, Mercedes Eugenia Paoletti, Juan A. Rico-Gallego |
Neurocomputing | 4 |
| 2021 | Distributed Deep Learning for Remote Sensing Data InterpretationabstractAs a newly emerging technology, deep learning (DL) is a very promising field in big data applications. Remote sensing often involves huge data volumes obtained daily by numerous in-orbit satellites. This makes it a perfect target area for data-driven applications. Nowadays, technological advances in terms of software and hardware have a noticeable impact on Earth observation applications, more specifically in remote sensing techniques and procedures, allowing for the acquisition of data sets with greater quality at higher acquisition ratios. This results in the collection of huge amounts of remotely sensed data, characterized by their large spatial resolution (in terms of the number of pixels per scene), and very high spectral dimensionality, with hundreds or even thousands of spectral bands. As a result, remote sensing instruments on spaceborne and airborne platforms are now generating data cubes with extremely high dimensionality, imposing several restrictions in terms of both processing runtimes and storage capacity. In this article, we provide a comprehensive review of the state of the art in DL for remote sensing data interpretation, analyzing the strengths and weaknesses of the most widely used techniques in the literature, as well as an exhaustive description of their parallel and distributed implementations (with a particular focus on those conducted using cloud computing systems). We also provide quantitative results, offering an assessment of a DL technique in a specific case study (source code available: https://github.com/mhaut/cloud-dnn-HSI). This article concludes with some remarks and hints about future challenges in the application of DL techniques to distributed remote sensing data interpretation problems. We emphasize the role of the cloud in providing a powerful architecture that is now able to manage vast amounts of remotely sensed data due to its implementation simplicity, low cost, and high efficiency compared to other parallel and distributed architectures, such as grid computing or dedicated clusters. Juan Mario Haut, Mercedes Eugenia Paoletti, Sergio Moreno-Álvarez, Javier Plaza, Juan A. Rico-Gallego, Antonio Plaza |
Proc. IEEE | 5 |
| 2020 | Training deep neural networks: a static load balancing approach
Sergio Moreno-Álvarez, Juan Mario Haut, Mercedes Eugenia Paoletti, Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Javier Plaza |
J. Supercomput. | 4 |
| 2020 | A tool to assess the communication cost of parallel kernels on heterogeneous platforms
Juan A. Rico-Gallego, Sergio Moreno-Álvarez, Juan Carlos Díaz Martín, Alexey L. Lastovetsky |
J. Supercomput. | 1 |
| 2019 | Analytical Communication Performance Models as a metric in the partitioning of data-parallel kernels on heterogeneous platforms
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Carmen Calvo-Jurado, Sergio Moreno-Álvarez, Juan-Luis García Zapata |
J. Supercomput. | 1 |
| 2017 | Formal modeling and performance evaluation of a run-time rank remapping technique in Broadcast, Allgather and Allreduce MPI collective operationsabstractMPI collective operations are implemented using a variety of algorithms which define different communication patterns between the ranks involved in the operation. The performance of these algorithms in multi-core clusters highly depends on the mapping of the ranks to the system processors due to the uneven capabilities of shared memory and network channels. The hierarchical design of these algorithms contributes to use optimally the communication channels. Nevertheless, common hierarchical algorithms have shown themselves, for some collectives as allgather, inefficient and even impracticable. This paper analyzed the reasons for that and works out an alternate approach through performance modeling. Such approach, departing from the a priori knowledge of a regular mapping as round-robin or sequential, and keeping the original algorithm unmodified, switches at run time the rank to process mapping into another regular mapping that reduces network traffic. The methodology is evaluated with three collectives and their underlying algorithms, showing speedups of up to 5x in the Binomial Tree or 3x in Ring algorithms compared to unfavorable mappings. Jesús M. Álvarez-Llorente, Juan Carlos Díaz Martín, Juan A. Rico-Gallego |
CCGrid | 3 |
| 2017 | Model-Based Estimation of the Communication Cost of Hybrid Data-Parallel Applications on Heterogeneous ClustersabstractHeterogeneous systems composed of CPUs and accelerators sharing communication channels of different performance are getting mainstream in HPC but, at the same time, they show a complexity that makes it difficult to optimize the deployment of a data parallel application. Recent analytical tools such as Functional Performance Models, combined with advanced partitioning algorithms, manage to achieve a balanced configuration by distributing the workload unevenly, according to the performance of the different processing units. Unfortunately, such uneven distribution of the computation load leads to communication unbalances that, very often, render worthless the previous workload balancing efforts. Finding the optimal communication scheme without expensive testing on the executing platform requires an analytical approach to the estimation of the communication cost of different configurations of the application. With this goal in mind, we propose and discuss an extension of the t-Lop communication performance model to cover heterogeneous architectures. In order to provide a quantitative assessment of this extended model, we conduct experiments with two representative computational kernels, the SUMMA algorithm and the 2D wave equation solver. The t-Lop predictions are compared against the HLogGP model and the observed costs for a variety of configurations, hardware resources and problem sizes. Juan A. Rico-Gallego, Alexey L. Lastovetsky, Juan Carlos Díaz Martín |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Extending τ-Lop to model concurrent MPI communications in multicore clusters
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Alexey L. Lastovetsky |
Future Gener. Comput. Syst. | 1 |
| 2015 | τ-Lop: Modeling performance of shared memory MPI
Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
Parallel Comput. | 1 |
| 2013 | On the performance of concurrent transfers in collective algorithmsabstractInter- and intra-machine MPI collective operations in current multicore clusters are essentially different, and therefore their performance modelling ask for different approaches. Inside a multicore each individual message transmission in a collective operation flow in parallel with others, but sharing the bandwidth of the main memory channel. Current models ignore this issue, making errors like giving the same cost estimation to quite different collective algorithms. We outline a new cost model focused on shared channels. Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
EuroMPI | 1 |
| 2012 | Improving Collectives by User Buffer Relocation
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Carolina Gómes-Tostón Gutierrez, Álvaro Cortés Fácila |
EuroMPI | 1 |
| 2011 | Performance Evaluation of Thread-Based MPI in Shared Memory
Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
EuroMPI | 1 |
| 2008 | A Network Service for DSP Multicomputers
Juan A. Rico-Gallego, Jesús M. Álvarez-Llorente, Juan Carlos Díaz Martín, Francisco J. Perogil-Duque |
ICA3PP | 1 |