VLDB 2026 Research / reviewers in the wild / expert
José M. Badía
dblp:b/JoseMBadia · also José Manuel Badía-Contelles
· DBLP profile ↗
31ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-5927-0449ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 5Graphics, computer vision, multimedia, augmented reality and games · 5Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy and robustness trade-offs in adaptive neural mmWave channel estimation on edge devicesabstractAbstract The evolution toward 6G will continue to leverage massive multiple-input multiple-output and millimeter-wave systems, which demand accurate angle-of-arrival (AoA) and angle-of-departure (AoD) estimation. While several deep learning models have demonstrated strong performance for this task, their accuracy, like that of most estimation methods, is often degraded by hardware non-idealities, which can be further exacerbated by time-varying operational factors such as component aging and adverse weather, among others. Building on a pre-trained U-Net architecture with demonstrated competitive performance for AoA/AoD estimation, we first propose an adaptation mechanism based on fine-tuning with impairment-augmented data. Specifically, we simulate hardware imperfections by introducing random phase errors in the antenna elements, ranging from mild fluctuations to severe signal distortions. The U-Net model with adaptation capabilities is then implemented on an NVIDIA Jetson Orin Nano device, a compact edge platform with heterogeneous computing resources. To this end, we design a co-execution strategy that performs AoA/AoD estimation (inference) on the CPU while simultaneously fine-tuning the model on the GPU, thus enabling continuous model adaptation to changing environmental or hardware conditions while preserving real-time inference performance. Experimental results show that impairment-aware fine-tuning effectively counters hardware degradation, particularly under significant phase impairments. In such scenarios, the fine-tuned model consistently preserves or even improves estimation accuracy, reducing the Root Mean Square Error (RMSE) by approximately 3.6% and increasing the Probability of Detection ( $$P_D$$ P D ) by up to 1 percentage point compared to the base model. Furthermore, a detailed energy-performance analysis demonstrates that while maximum frequency settings reduce training time by over 11 $$\times $$ × , they also increase power consumption by more than 5 $$\times $$ × , with optimal energy efficiency achieved at mid-range CPU and high GPU frequencies. This work establishes the feasibility of concurrent training and inference on resource-constrained heterogeneous hardware, paving the way for resilient and autonomous edge intelligence in future 6G systems. Eric Meneses Albalá, Saúl Villaescusa, José M. Badía, German Leon, Carmen Botella-Mascarell, Sandra Roger 0002 |
J. Supercomput. | 3 |
| 2026 | Dependability analysis and hardening of vision transformers against soft errorsabstractAbstract The deployment of Vision Transformers (ViTs) in safety-critical domains needs a clear understanding of their resilience to soft errors, since their specific layer-level vulnerabilities are currently insufficiently characterized. This work presents a dependability analysis of the ViT-Base architecture against injection-induced soft errors. Using a high-fidelity, software-level fault injection methodology with custom CUDA kernels, the study injects random bit-flips directly into the IEEE 754 binary32 floating-point representation of the intermediate data tensors resulting from the Transformer modules to quantify model accuracy degradation across increasing bit error rates. As a primary result, a vulnerability map across ViT layers is presented, confirming that the results of normalization and fully connected layers exhibit critical sensitivity to soft errors. To address these vulnerabilities, the work evaluates targeted hardening strategies. These include Fault-Aware Training (FAT), applied both globally and selectively to linear layers, as well as practical runtime mitigations such as range-based value clipping and filtering of non-numeric values. The findings demonstrate that these software-only approaches can significantly protect model accuracy. Lester Frias-Dominguez, José M. Badía, German Leon, Adrian Amor-Martin, Jose A. Belloch |
J. Supercomput. | 2 |
| 2025 | Optimizing Millimeter Wave MIMO Channel Estimation Through GPU-Based Edge Artificial IntelligenceabstractIn the context of upcoming sixth-generation (6G) wireless communication systems, the use of millimeter wave (mmWave) frequencies is a key technology for achieving high-throughput communications. Accurate parametric estimation of mmWave channels is critical for effective beamforming design and configuration, requiring sophisticated models to capture the directional characteristics of these channels. This work considers an innovative artificial intelligence (AI) approach for accurate estimation of angle-of-arrival (AoA) and angle-of-departure (AoD) parameters from frequency-domain channel observations. Our approach is based on the implementation of two convolutional neural networks (CNNs): a residual CNN (ResNet) and a U-Net CNN. Specifically, this work focuses on the efficient implementation of both schemes in an embedded system suitable for edge AI. We performed the experiments in a low-power NVIDIA Jetson Orin Nano platform and evaluated the effect of modifying the frequencies of its CPU and GPU on the performance of the inference process, both in terms of execution time and energy consumption. Experimental results showed that the U-Net model is more power consuming, but as it is faster, it consumes less energy per channel. Diego Lloria, Sandra Roger 0002, German Leon, José M. Badía, Carmen Botella-Mascarell, Jose A. Belloch |
J. Supercomput. | 4 |
| 2025 | Efficient quantum circuit contraction using tensor decision diagramsabstractAbstract Simulating quantum circuits efficiently on classical computers is crucial given the limitations of current noisy intermediate-scale quantum devices. This paper adapts and extends two methods used to contract tensor networks within the fast tensor decision diagram (FTDD) framework. The methods, called iterative pairing and block contraction, exploit the advantages of tensor decision diagrams to reduce both the temporal and spatial cost of quantum circuit simulations. The iterative pairing method minimizes intermediate diagram sizes, while the block contraction algorithm efficiently handles circuits with repetitive structures, such as those found in quantum walks and Grover’s algorithm. Experimental results demonstrate that, in some cases, these methods significantly outperform traditional contraction orders like sequential and cotengra in terms of both memory usage and execution time. Furthermore, simulation tools based on decision diagrams, such as FTDD, show superior performance to matrix-based simulation tools, such as Google tensor networks, enabling the simulation of larger circuits more efficiently. These findings show the potential of decision diagram-based approaches to improve the simulation of quantum circuits on classical platforms. Vicente Lopez-Oliva, José M. Badía, María Isabel Castillo |
J. Supercomput. | 2 |
| 2025 | Evaluating and accelerating vision transformers on GPU-based embedded edge AI systemsabstractAbstract Many current embedded systems comprise heterogeneous computing components including quite powerful GPUs, which enables their application across diverse sectors. This study demonstrates the efficient execution of a medium-sized self-supervised audio spectrogram transformer (SSAST) model on a low-power system-on-chip (SoC). Through comprehensive evaluation, including real time inference scenarios, we show that GPUs outperform multi-core CPUs in inference processes. Optimization techniques such as adjusting batch size, model compilation with TensorRT, and reducing data precision significantly enhance inference time, energy consumption, and memory usage. In particular, negligible accuracy degradation is observed, with post-training quantization to 8-bit integers showing less than 1% loss. This research underscores the feasibility of deploying transformer neural networks on low-power embedded devices, ensuring efficiency in time, energy, and memory, while maintaining the accuracy of the results. Ignacio Martin-Salinas, José M. Badía, Óscar Valls, German Leon, Rocío del Amor, Jose A. Belloch, Adrian Amor-Martin, Valery Naranjo |
J. Supercomput. | 2 |
| 2025 | A community detection-based parallel algorithm for quantum circuit simulation using tensor networksabstractAbstract Quantum computing holds significant promise for solving complex problems, but simulating quantum circuits on classical computers remains essential due to the current limitations of quantum hardware. Efficient simulation is crucial for the development and validation of quantum algorithms and quantum computers. This paper explores and compares various strategies to leverage different levels of parallelism to accelerate the contraction of tensor networks representing large quantum circuits. We propose a new parallel multistage algorithm based on communities. The original tensor network is partitioned into several communities, which are then contracted in parallel. The pairs of tensors of the resulting network can be contracted in parallel using a GPU. We use the Girvan–Newman algorithm to obtain the communities and the contraction plans. We compare the new algorithm with two other parallelisation strategies: one based on contracting all the pairs of tensors in the GPU and another one that uses slicing to cut some indexes of the tensor network and then MPI processes to contract the resulting slices in parallel. The new parallel algorithm gets the best results with different well-known quantum circuits with a high degree of entanglement, including random quantum circuits. In conclusion, the results show that the main factor that limits the simulation is the space cost. However, the parallel multistage algorithm manages to reduce the cost of sequential simulation for circuits with a high number of qubits and allows simulating larger circuits. Alfred M. Pastor, José M. Badía, María Isabel Castillo |
J. Supercomput. | 2 |
| 2024 | Comparative analysis of soft-error sensitivity in LU decomposition algorithms on diverse GPUsabstractAbstract Graphics processing units (GPUs) have become integral to embedded systems and supercomputing centres due to their large memory, cutting-edge technology and high performance per watt. However, their susceptibility to transient errors requires a comprehensive analysis of error sensitivity, as well as the development of error mitigation techniques and fault-tolerant algorithms. This study focuses on evaluating the soft-error sensitivity of two distinct versions of LU decomposition algorithms implemented on two very different GPUs—a low-power SoC embedded GPU and a high-performance massively parallel GPU. Through extensive fault injection campaigns on both GPUs, we examine the vulnerability of the algorithms, identify error causes, and determine critical code components requiring enhanced protection. The experiments reveal that most single bit flip fault injections in the instruction results lead to erroneous outcomes or unrecoverable errors. Notably, efficient GPU resource utilisation can increase the number of masked errors, thereby enhancing error resilience. Additionally, while different parts of the code exhibit similar error occurrence types and rates, the propagation of errors to elements within the result matrix differs significantly. German Leon, José M. Badía, Jose A. Belloch, Almudena Lindoso, Luis Entrena |
J. Supercomput. | 2 |
| 2023 | Strategies to parallelize a finite element mesh truncation technique on multi-core and many-core architecturesabstractAbstract Achieving maximum parallel performance on multi-core CPUs and many-core GPUs is a challenging task depending on multiple factors. These include, for example, the number and granularity of the computations or the use of the memories of the devices. In this paper, we assess those factors by evaluating and comparing different parallelizations of the same problem on a multiprocessor containing a CPU with 40 cores and four P100 GPUs with Pascal architecture. We use, as study case, the convolutional operation behind a non-standard finite element mesh truncation technique in the context of open region electromagnetic wave propagation problems. A total of six parallel algorithms implemented using OpenMP and CUDA have been used to carry out the comparison by leveraging the same levels of parallelism on both types of platforms. Three of the algorithms are presented for the first time in this paper, including a multi-GPU method, and two others are improved versions of algorithms previously developed by some of the authors. This paper presents a thorough experimental evaluation of the parallel algorithms on a radar cross-sectional prediction problem. Results show that performance obtained on the GPU clearly overcomes those obtained in the CPU, much more so if we use multiple GPUs to distribute both data and computations. Accelerations close to 30 have been obtained on the CPU, while with the multi-GPU version accelerations larger than 250 have been achieved. José M. Badía, Adrian Amor-Martin, Jose A. Belloch, L. E. García-Castillo |
J. Supercomput. | 1 |
| 2022 | Multicore implementation of a multichannel parallel graphic equalizerabstractAbstract Numerous signal processing applications are emerging on mobile computing systems. These applications are subject to responsiveness constraints for user interactivity and, at the same time, must be optimized for energy efficiency. Many current embedded devices are composed of low-power multicore processors that offer a good trade-off between computational capacity and low power consumption. In this context, equalizers are widely used in multiple mobile-based applications such as “Music streaming” to adjust the levels of bass and treble in sound reproduction. In this study, we evaluate a graphic equalizer from audio, computational capacity, and energy efficiency perspectives, as well as the execution of multiple real-time equalizers running on an embedded quad-core processor of a mobile device. To this end, we experiment with the working frequencies as well as the parallelism that can be extracted from a quad-core ARM Cortex-A57. Results show that using high CPU frequencies and three or four cores, our parallel algorithm is able to equalize more than five channels per watt in real time with an audio buffer of 4096 samples, which implies a latency of 92.8 ms at the standard sample rate of 44.1 kHz. Jose A. Belloch, José M. Badía, German Leon, Balázs Bank, Vesa Välimäki |
J. Supercomput. | 2 |
| 2021 | On the performance of a GPU-based SoC in a distributed spatial audio system
Jose A. Belloch, José M. Badía, Diego Francisco Larios Marín, Enrique Personal, Miguel Ferrer 0001, Laura Fuster, Mihaita Lupoiu, Alberto González 0001, Carlos León 0001, Antonio M. Vidal, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2021 | Evaluating the computational performance of the Xilinx Ultrascale+ EG Heterogeneous MPSoC
Jose A. Belloch, German Leon, José M. Badía, Almudena Lindoso, Enrique San Millán |
J. Supercomput. | 3 |
| 2019 | Practical Considerations for Acoustic Source Localization in the IoT Era: Platforms, Energy Efficiency, and PerformanceabstractThe rapid development of the Internet of Things (IoT) has posed important changes in the way emerging acoustic signal processing applications are conceived. While traditional acoustic processing applications have been developed taking into account high-throughput computing platforms equipped with expensive multichannel audio interfaces, the IoT paradigm is demanding the use of more flexible and energy-efficient systems. In this context, algorithms for source localization and ranging in wireless acoustic sensor networks can be considered an enabling technology for many IoT-based environments, including security, industrial, and health-care applications. This paper is aimed at evaluating important aspects dealing with the practical deployment of IoT systems for acoustic source localization. Recent systems-on-chip composed of low-power multicore processors, combined with a small graphics accelerator (or GPU), yield a notable increment of the computational capacity needed in intensive signal processing algorithms while partially retaining the appealing low power consumption of embedded systems. Different algorithms and implementations over several state-of-the-art platforms are discussed, analyzing important aspects, such as the tradeoffs between performance, energy efficiency, and exploitation of parallelism by taking into account real-time constraints. Jose A. Belloch, José M. Badía, Francisco D. Igual, Maximo Cobos |
IEEE Internet Things J. | 2 |
| 2019 | Accelerating the SRP-PHAT algorithm on multi- and many-core platforms using OpenCL
José M. Badía, Jose A. Belloch, Maximo Cobos, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2017 | Accessing very high dimensional spaces in parallel
Fernando José Artigas-Fuentes, José M. Badía |
J. Supercomput. | 2 |
| 2015 | Out-of-core macromolecular simulations on multithreaded architecturesabstractSummary We address the solution of large‐scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multithreaded platforms, consisting of multicore processors and possibly a graphics processor (graphics processing unit). In particular, we compare specialized implementations of several high‐performance eigensolvers that, by relying on disk storage and out‐of‐core techniques, can in principle tackle the large memory requirements of these biological problems, which in general do not fit into the main memory of current desktop machines. All these out‐of‐core eigensolvers, except for one, are composed of compute‐bound (i.e., arithmetically intensive) operations, which we accelerate by exploiting the performance of current multicore processors and, in some cases, by additionally off‐loading certain parts of the computation to a graphics processing unit accelerator. One of the eigensolvers is a memory‐bound algorithm, which strongly constrains its performance when the data is on disk. However, this method exhibits a much lower arithmetic cost compared with its compute‐bound alternatives for this particular application. Experimental results on a desktop platform, representative of current server technology, illustrate the potential of these methods to address the simulation of biological activity. Copyright © 2014 John Wiley & Sons, Ltd. José Ignacio Aliaga, José M. Badía, María Isabel Castillo, Davor Davidovic, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Fast k-NN Classifier for Documents Based on a Graph Structure
Fernando José Artigas-Fuentes, Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
CIARP | 3 |
| 2010 | A High-Dimensional Access Method for Approximated Similarity Search in Text MiningabstractIn this paper, a new access method for very high-dimensional data space is proposed. The method uses a graph structure and pivots for indexing objects, such as documents in text mining. It also applies a simple search algorithm that uses distance or similarity based functions in order to obtain the k-nearest neighbors for novel query objects. This method shows a good selectivity over very-high dimensional data spaces, and a better performance than other state-of-the-art methods. Although it is a probabilistic method, it shows a low error rate. The method is evaluated on data sets from the well-known collection Reuters corpus version 1 (RCV1-v2) and dealing with thousands of dimensions. Fernando José Artigas-Fuentes, Reynaldo Gil-García, José M. Badía |
ICPR | 3 |
| 2009 | Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
J. Supercomput. | 3 |
| 2007 | Parallel Nearest Neighbour Algorithms for Text Categorization
Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
Euro-Par | 2 |
| 2007 | Performance analysis for clusters of symmetric multiprocessorsabstractIn this article we analyze and model the performance of a symmetrical multiprocessor cluster. The obtained model takes into account the heterogeneity of the architecture's communications, which allows for a better adjustment and predictive ability of the performance. We used diverse applications with different parallel schemes to check the quality of the adjustments. The adjustment of the theoretical model and the experimental results is clearly improved when we take into account the combination of various mechanisms of communication Francisco Almeida, Juan A. Gómez, José M. Badía |
PDP | 3 |
| 2006 | Parallel Solution of Large-Scale and Sparse Generalized Algebraic Riccati Equations
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 1 |
| 2006 | Parallelization of GSL: The Web Service InterfaceabstractWe present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements. José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
PDP | 2 |
| 2005 | Dynamic Hierarchical Compact Clustering Algorithm
Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
CIARP | 2 |
| 2005 | Parallel Order Reduction via Balanced Truncation for Optimal Cooling of Steel Profiles
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Jens Saak |
Euro-Par | 1 |
| 2005 | Solving the block-Toeplitz least-squares problem in parallelabstractIn this paper we present two versions of a parallel algorithm to solve the block–Toeplitz least-squares problem on distributed-memory architectures. We derive a parallel algorithm based on the seminormal equations arising from the triangular decomposition of the product T TT . Our parallel algorithm exploits the displacement structure of the Toeplitz-likematrices using theGeneralized SchurAlgorithm to obtain the solution in O(mn) flops instead of O(mn2) flops of the algorithms for non-structured matrices. The strong regularity of the previous product of matrices and an appropriate computation of the hyperbolic rotations improve the stability of the algorithms. We have reduced the communication cost of previous versions, and have also reduced the memory access cost by appropriately arranging the elements of the matrices. Furthermore, the second version of the algorithm has a very low spatial cost, because it does not store the triangular factor of the decomposition. The experimental results show a good scalability of the parallel algorithm on two different clusters of personal computers. Copyright c © 2005 John Wiley & Sons, Ltd. Pedro Alonso 0002, José M. Badía, Antonio M. Vidal |
Concurr. Pract. Exp. | 2 |
| 2005 | An Efficient Parallel Algorithm to Solve Block-Toeplitz Systems
Pedro Alonso 0002, José M. Badía, Antonio M. Vidal |
J. Supercomput. | 2 |
| 2004 | Parallel Algorithm for Extended Star Clustering
Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
CIARP | 2 |
| 2003 | Extended Star Clustering Algorithm
Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
CIARP | 2 |
| 2003 | A Parallel Algorithm for Incremental Compact Clustering
Reynaldo Gil-García, José M. Badía, Aurora Pons-Porrata |
Euro-Par | 2 |
| 2002 | Solving Large Sparse Lyapunov Equations on Parallel Computers (Research Note)
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 1 |
| 2000 | Inverse Toeplitz eigenproblem on personal computer networksabstractIn this paper we present a parallel algorithm for solving the inverse Toeplitz Eigenvalue Problem. The algorithm has been implemented by using a cluster of personal computers, interconnected by a high-performance Myrinet network. We have utilized standard public domain parallel environments for implementing the calculation part as well as the communications, thus producing portable software. The results obtained allow us to confirm the scalability and efficiency of the algorithm. Moreover, we have checked that by using the theoretical cost model provided by the ScaLAPACK we can predict the behaviour of the experimental results. Copyright © 2000 John Wiley & Sons, Ltd. José M. Badía, Antonio M. Vidal |
Concurr. Pract. Exp. | 1 |