VLDB 2026 Research / reviewers in the wild / expert
Jose A. Belloch
dblp:89/10398 · also Jose Antonio Belloch
· DBLP profile ↗
25ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0002-2595-1828ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorComputer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dependability analysis and hardening of vision transformers against soft errorsabstractAbstract The deployment of Vision Transformers (ViTs) in safety-critical domains needs a clear understanding of their resilience to soft errors, since their specific layer-level vulnerabilities are currently insufficiently characterized. This work presents a dependability analysis of the ViT-Base architecture against injection-induced soft errors. Using a high-fidelity, software-level fault injection methodology with custom CUDA kernels, the study injects random bit-flips directly into the IEEE 754 binary32 floating-point representation of the intermediate data tensors resulting from the Transformer modules to quantify model accuracy degradation across increasing bit error rates. As a primary result, a vulnerability map across ViT layers is presented, confirming that the results of normalization and fully connected layers exhibit critical sensitivity to soft errors. To address these vulnerabilities, the work evaluates targeted hardening strategies. These include Fault-Aware Training (FAT), applied both globally and selectively to linear layers, as well as practical runtime mitigations such as range-based value clipping and filtering of non-numeric values. The findings demonstrate that these software-only approaches can significantly protect model accuracy. Lester Frias-Dominguez, José M. Badía, German Leon, Adrian Amor-Martin, Jose A. Belloch |
J. Supercomput. | 5 |
| 2026 | Real-time object tracking with on-device deep learning for adaptive beamforming in dynamic acoustic environmentsabstractAbstract Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning–based tracking with beamforming to achieve precise sound source localization and directional audio capture in dynamic environments. The approach combines single-camera depth estimation and stereo vision to enable accurate 3D localization of moving objects. A planar concentric circular microphone array constructed with MEMS microphones provides a compact, energy-efficient platform supporting 2D beam steering across azimuth and elevation. Real-time tracking outputs continuously adapt the array’s focus, synchronizing the acoustic response with the target’s position. By uniting learned spatial awareness with dynamic steering, the system maintains robust performance in the presence of multiple or moving sources. Experimental evaluation demonstrates significant gains in signal-to-interference ratio, making the design well-suited for teleconferencing, smart home devices, and assistive technologies. Jorge Ortigoso-Narro, Jose A. Belloch, Adrian Amor-Martin, Sandra Roger 0002, Maximo Cobos |
J. Supercomput. | 2 |
| 2025 | Optimizing Millimeter Wave MIMO Channel Estimation Through GPU-Based Edge Artificial IntelligenceabstractIn the context of upcoming sixth-generation (6G) wireless communication systems, the use of millimeter wave (mmWave) frequencies is a key technology for achieving high-throughput communications. Accurate parametric estimation of mmWave channels is critical for effective beamforming design and configuration, requiring sophisticated models to capture the directional characteristics of these channels. This work considers an innovative artificial intelligence (AI) approach for accurate estimation of angle-of-arrival (AoA) and angle-of-departure (AoD) parameters from frequency-domain channel observations. Our approach is based on the implementation of two convolutional neural networks (CNNs): a residual CNN (ResNet) and a U-Net CNN. Specifically, this work focuses on the efficient implementation of both schemes in an embedded system suitable for edge AI. We performed the experiments in a low-power NVIDIA Jetson Orin Nano platform and evaluated the effect of modifying the frequencies of its CPU and GPU on the performance of the inference process, both in terms of execution time and energy consumption. Experimental results showed that the U-Net model is more power consuming, but as it is faster, it consumes less energy per channel. Diego Lloria, Sandra Roger 0002, German Leon, José M. Badía, Carmen Botella-Mascarell, Jose A. Belloch |
J. Supercomput. | 6 |
| 2025 | Evaluating and accelerating vision transformers on GPU-based embedded edge AI systemsabstractAbstract Many current embedded systems comprise heterogeneous computing components including quite powerful GPUs, which enables their application across diverse sectors. This study demonstrates the efficient execution of a medium-sized self-supervised audio spectrogram transformer (SSAST) model on a low-power system-on-chip (SoC). Through comprehensive evaluation, including real time inference scenarios, we show that GPUs outperform multi-core CPUs in inference processes. Optimization techniques such as adjusting batch size, model compilation with TensorRT, and reducing data precision significantly enhance inference time, energy consumption, and memory usage. In particular, negligible accuracy degradation is observed, with post-training quantization to 8-bit integers showing less than 1% loss. This research underscores the feasibility of deploying transformer neural networks on low-power embedded devices, ensuring efficiency in time, energy, and memory, while maintaining the accuracy of the results. Ignacio Martin-Salinas, José M. Badía, Óscar Valls, German Leon, Rocío del Amor, Jose A. Belloch, Adrian Amor-Martin, Valery Naranjo |
J. Supercomput. | 6 |
| 2024 | Urban sound classification using neural networks on embedded FPGAsabstractAbstract Sound classification using neural networks has recently produced very accurate results. A large number of different applications use this type of sound classifiers such as controlling and monitoring the type of activity in a city or identifying different types of animals in natural environments. While traditional acoustic processing applications have been developed on high-performance computing platforms equipped with expensive multi-channel audio interfaces, the Internet of Things (IoT) paradigm requires the use of more flexible and energy-efficient systems. Although software-based platforms exist for implementing general-purpose neural networks, they are not optimized for sound classification, wasting energy and computational resources. In this work, we have used FPGAs to develop an ad hoc system where only the hardware needed for our application is synthesized, resulting in faster and more energy-efficient circuits. The results show that our developments are accelerated by a factor of 35 compared to a software-based implementation on a Raspberry Pi. Jose A. Belloch, Raul Coronado, Óscar Valls, Rocío del Amor, German Leon, Valery Naranjo, Manuel F. Dolz, Adrian Amor-Martin, Gema Piñero |
J. Supercomput. | 1 |
| 2024 | Comparative analysis of soft-error sensitivity in LU decomposition algorithms on diverse GPUsabstractAbstract Graphics processing units (GPUs) have become integral to embedded systems and supercomputing centres due to their large memory, cutting-edge technology and high performance per watt. However, their susceptibility to transient errors requires a comprehensive analysis of error sensitivity, as well as the development of error mitigation techniques and fault-tolerant algorithms. This study focuses on evaluating the soft-error sensitivity of two distinct versions of LU decomposition algorithms implemented on two very different GPUs—a low-power SoC embedded GPU and a high-performance massively parallel GPU. Through extensive fault injection campaigns on both GPUs, we examine the vulnerability of the algorithms, identify error causes, and determine critical code components requiring enhanced protection. The experiments reveal that most single bit flip fault injections in the instruction results lead to erroneous outcomes or unrecoverable errors. Notably, efficient GPU resource utilisation can increase the number of masked errors, thereby enhancing error resilience. Additionally, while different parts of the code exhibit similar error occurrence types and rates, the propagation of errors to elements within the result matrix differs significantly. German Leon, José M. Badía, Jose A. Belloch, Almudena Lindoso, Luis Entrena |
J. Supercomput. | 3 |
| 2023 | Strategies to parallelize a finite element mesh truncation technique on multi-core and many-core architecturesabstractAbstract Achieving maximum parallel performance on multi-core CPUs and many-core GPUs is a challenging task depending on multiple factors. These include, for example, the number and granularity of the computations or the use of the memories of the devices. In this paper, we assess those factors by evaluating and comparing different parallelizations of the same problem on a multiprocessor containing a CPU with 40 cores and four P100 GPUs with Pascal architecture. We use, as study case, the convolutional operation behind a non-standard finite element mesh truncation technique in the context of open region electromagnetic wave propagation problems. A total of six parallel algorithms implemented using OpenMP and CUDA have been used to carry out the comparison by leveraging the same levels of parallelism on both types of platforms. Three of the algorithms are presented for the first time in this paper, including a multi-GPU method, and two others are improved versions of algorithms previously developed by some of the authors. This paper presents a thorough experimental evaluation of the parallel algorithms on a radar cross-sectional prediction problem. Results show that performance obtained on the GPU clearly overcomes those obtained in the CPU, much more so if we use multiple GPUs to distribute both data and computations. Accelerations close to 30 have been obtained on the CPU, while with the multi-GPU version accelerations larger than 250 have been achieved. José M. Badía, Adrian Amor-Martin, Jose A. Belloch, L. E. García-Castillo |
J. Supercomput. | 3 |
| 2023 | Hybrid CPU-GPU implementation of the transformed spatial domain channel estimation algorithm for mmWave MIMO systemsabstractAbstract Hybrid platforms combining multicore central processing units (CPU) with many-core hardware accelerators such as graphic processing units (GPU) can be smartly exploited to provide efficient parallel implementations of wireless communication algorithms for Fifth Generation (5G) and beyond systems. Massive multiple-input multiple-output (MIMO) systems are a key element of the 5G standard, involving several tens or hundreds of antenna elements for communication. Such a high number of antennas has a direct impact on the computational complexity of some MIMO signal processing algorithms. In this work, we focus on the channel estimation stage. In particular, we develop a parallel implementation of a recently proposed MIMO channel estimation algorithm. Its performance in terms of execution time is evaluated both in a multicore CPU and in a GPU. The results show that some computation blocks of the algorithm are more suitable for multicore implementation, whereas other parts are more efficiently implemented in the GPU, indicating that a hybrid CPU–GPU implementation would achieve the best performance in practical applications based on the tested platform. Diego Lloria, Pablo M. Aviles, Jose A. Belloch, Sandra Roger 0002, Carmen Botella-Mascarell, Almudena Lindoso |
J. Supercomput. | 3 |
| 2022 | Acceleration of the TSDCE MIMO Channel Estimation Algorithm on a Multi-core PlatformabstractThe use of Multi-Processor System-on-Chip (MPSoC) is becoming widespread in a huge number of signal processing systems, including wireless communications and vehicular technology applications. In those scenarios, when Multiple-Input Multiple-Output (MIMO) communication schemes are considered, the system usually has to deal with a high number of communication links that involve sensors and antennas from different vehicles and users. The use of MIMO systems with a high number of antennas increases the complexity of many signal processing algorithms which could benefit from computationally efficient implementations. The Xilinx Zynq UltraScale+ EG Heterogeneous MPSoC is a well-positioned platform to manage computationally-demanding communication systems. This platform holds a dual-core ARM Cortex-R5, a quad-core ARM Cortex-A53, a graphics processing unit (GPU) and a high-end Field Programmable Gate Array (FPGA). In particular, this work aims to evaluate the computational performance of the Transformed Spatial Domain Channel Estimation (TSDCE), a novel millimeter-wave MIMO channel estimation algorithm, on the proposed embedded platform. This work focuses firstly on developing an efficient sequential implementation that runs on the ARM Cortex-A53, so that we can afterwards leverage the use of the multi-core system to accelerate the sequential performance. Pablo M. Aviles, Diego Lloria, Jose A. Belloch, Sandra Roger 0002, Almudena Lindoso, Maximo Cobos |
EATIS | 3 |
| 2022 | Performance analysis of a millimeter wave MIMO channel estimation method in an embedded multi-core processorabstractAbstract The emerging Multi-Processor System-on-Chip (MPSoC) technology, which combines heterogeneous computing with the high performance of field programmable gate arrays (FPGA), is a promising platform for a large number of applications, including wireless communications and vehicular technology. In this specific application context, when multiple-input multiple-output (MIMO) scenarios are considered, the system usually has to manage a large number of communication links among sensors and antennas involving different vehicles and users. Millimeter wave (mmWave) communications are one of the key technology enablers toward achieving high data rates in beyond 5G systems (B5G). Communication at these frequency bands usually involves the use of large antenna arrays, often requiring high computational resources. One of the candidate platforms able to manage a huge number of communications is the Xilinx Zynq UltraScale+ EG Heterogeneous MPSoC, which is composed of a dual-core Cortex-R5, a quad-core ARM Cortex-A53, a graphics processing unit (GPU) and a high-end FPGA. This work analyzes the computational performance that requires a recent mmWave MIMO channel estimation algorithm in a platform of this kind. As a first approach, we will focus our work on the performance that can be achieved via the quad-core ARM Cortex-A53. To this end, we will use the libraries for numerical algebra (BLAS and LAPACK). The results show that our reference implementation is able to manage a large MIMO communication system with 256 antennas without exhausting platform resources. Pablo M. Aviles, Diego Lloria, Jose A. Belloch, Sandra Roger 0002, Almudena Lindoso, Maximo Cobos |
J. Supercomput. | 3 |
| 2022 | Multicore implementation of a multichannel parallel graphic equalizerabstractAbstract Numerous signal processing applications are emerging on mobile computing systems. These applications are subject to responsiveness constraints for user interactivity and, at the same time, must be optimized for energy efficiency. Many current embedded devices are composed of low-power multicore processors that offer a good trade-off between computational capacity and low power consumption. In this context, equalizers are widely used in multiple mobile-based applications such as “Music streaming” to adjust the levels of bass and treble in sound reproduction. In this study, we evaluate a graphic equalizer from audio, computational capacity, and energy efficiency perspectives, as well as the execution of multiple real-time equalizers running on an embedded quad-core processor of a mobile device. To this end, we experiment with the working frequencies as well as the parallelism that can be extracted from a quad-core ARM Cortex-A57. Results show that using high CPU frequencies and three or four cores, our parallel algorithm is able to equalize more than five channels per watt in real time with an audio buffer of 4096 samples, which implies a latency of 92.8 ms at the standard sample rate of 44.1 kHz. Jose A. Belloch, José M. Badía, German Leon, Balázs Bank, Vesa Välimäki |
J. Supercomput. | 1 |
| 2021 | On the performance of a GPU-based SoC in a distributed spatial audio system
Jose A. Belloch, José M. Badía, Diego Francisco Larios Marín, Enrique Personal, Miguel Ferrer 0001, Laura Fuster, Mihaita Lupoiu, Alberto González 0001, Carlos León 0001, Antonio M. Vidal, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2021 | Evaluating the computational performance of the Xilinx Ultrascale+ EG Heterogeneous MPSoC
Jose A. Belloch, German Leon, José M. Badía, Almudena Lindoso, Enrique San Millán |
J. Supercomput. | 1 |
| 2019 | Practical Considerations for Acoustic Source Localization in the IoT Era: Platforms, Energy Efficiency, and PerformanceabstractThe rapid development of the Internet of Things (IoT) has posed important changes in the way emerging acoustic signal processing applications are conceived. While traditional acoustic processing applications have been developed taking into account high-throughput computing platforms equipped with expensive multichannel audio interfaces, the IoT paradigm is demanding the use of more flexible and energy-efficient systems. In this context, algorithms for source localization and ranging in wireless acoustic sensor networks can be considered an enabling technology for many IoT-based environments, including security, industrial, and health-care applications. This paper is aimed at evaluating important aspects dealing with the practical deployment of IoT systems for acoustic source localization. Recent systems-on-chip composed of low-power multicore processors, combined with a small graphics accelerator (or GPU), yield a notable increment of the computational capacity needed in intensive signal processing algorithms while partially retaining the appealing low power consumption of embedded systems. Different algorithms and implementations over several state-of-the-art platforms are discussed, analyzing important aspects, such as the tradeoffs between performance, energy efficiency, and exploitation of parallelism by taking into account real-time constraints. Jose A. Belloch, José M. Badía, Francisco D. Igual, Maximo Cobos |
IEEE Internet Things J. | 1 |
| 2019 | Accelerating the SRP-PHAT algorithm on multi- and many-core platforms using OpenCL
José M. Badía, Jose A. Belloch, Maximo Cobos, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2019 | On the use of many-core machines for the acceleration of a mesh truncation technique for FEM
Jose A. Belloch, Adrian Amor-Martin, Daniel Garcia-Donoro, Francisco-Jose Martínez-Zaldívar, L. E. García-Castillo |
J. Supercomput. | 1 |
| 2017 | A Parallel Approach to HRTF Approximation and Interpolation Based on a Parametric Filter ModelabstractSpatial audio-rendering techniques using head-related transfer functions (HRTFs) are currently used in many different contexts such as immersive teleconferencing systems, gaming, or 3-D audio reproduction. Since all these applications usually involve real-time constraints, efficient processing structures for HRTF modeling and interpolation are necessary for providing real-time binaural audio solutions. This letter presents a parametric parallel model that allows us to perform HRTF filtering and interpolation efficiently from an input HRTF dataset. The resulting model, which is an adaptation from a recently proposed modeling technique, not only reduces the size of HRTF datasets significantly, but also allows for simplified interpolation and real-time computation over parallel processors. In order to discuss the suitability of this new model, an implementation over a graphic processing unit is presented. Germán Ramos, Maximo Cobos, Balázs Bank, Jose A. Belloch |
IEEE Signal Process. Lett. | 4 |
| 2017 | GPU-Based Dynamic Wave Field Synthesis Using Fractional Delay Filters and Room CompensationabstractWave field synthesis (WFS) is a multichannel audio reproduction method, of a considerable computational cost that renders an accurate spatial sound field using a large number of loudspeakers to emulate virtual sound sources. The moving of sound source locations can be improved by using fractional delay filters, and room reflections can be compensated by using an inverse filter bank that corrects the room effects at selected points within the listening area. However, both the fractional delay filters and the room compensation filters further increase the computational requirements of the WFS system. This paper analyzes the performance of a WFS system composed of 96 loudspeakers which integrates both strategies. In order to deal with the large computational complexity, we explore the use of a graphics processing unit (GPU) as a massive signal co-processor to increase the capabilities of the WFS system. The performance of the method as well as the benefits of the GPU acceleration are demonstrated by considering different sizes of room compensation filters and fractional delay filters of order 9. The results show that a 96-speaker WFS system that is efficiently implemented on a state-of-art GPU can synthesize the movements of 94 sound sources in real time and, at the same time, can manage 9216 room compensation filters having more than 4000 coefficients each. Jose A. Belloch, Alberto González 0001, Enrique S. Quintana-Ortí, Miguel Ferrer 0001, Vesa Välimäki |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Accelerating multi-channel filtering of audio signal on ARM processors
Jose A. Belloch, Fran J. Alventosa, Pedro Alonso 0002, Enrique S. Quintana-Ortí, Antonio M. Vidal |
J. Supercomput. | 1 |
| 2017 | Solving Weighted Least Squares (WLS) problems on ARM-based architectures
Jose A. Belloch, Balázs Bank, Francisco D. Igual, Enrique S. Quintana-Ortí, Antonio M. Vidal |
J. Supercomput. | 1 |
| 2016 | Efficient target-response interpolation for a graphic equalizerabstractA graphic equalizer is an adjustable filter in which the command gain of each frequency band is practically independent of the gains of other bands. Designing a graphic equalizer with a high precision requires evaluating a target response that interpolates the magnitude response at several frequency points between the command gains. Good accuracy has been previously achieved by using polynomial interpolation methods such as cubic Hermite or spline interpolation. However, these methods require large computational resources, which is a limitation in real-time applications. This paper proposes an efficient way of computing the target response without sacrificing the approximation accuracy. This new approach called Linear Interpolation with Constant Segments (LICS) reduces the computing time of the target response by 55% and has an intrinsic parallel structure. Performance of the LICS method is assessed on an ARM Cortex-A7 core, which is commonly used in embedded systems. Jose A. Belloch, Vesa Välimäki |
ICASSP | 1 |
| 2015 | On the performance of multi-GPU-based expert systems for acoustic localization involving massive microphone arrays
Jose A. Belloch, Alberto González 0001, Antonio M. Vidal, Maximo Cobos |
Expert Syst. Appl. | 1 |
| 2014 | Multi-channel IIR filtering of audio signals using a GPUabstractIn the audio signal processing field, multiple IIR filters are required in many applications. As an example, equalizing a Wave Field Synthesis system requires massive filter processing in real time. Graphics Processing Units (GPUs) are well known for their potential in highly parallel data processing. Up to now, the use of the GPUs for implementing IIR filters has not been clearly tackled in audio processing because of its feedback loop that prevents its total parallelization. However, using the Parallel form of IIR filters, this feedback is reduced, since every single sample is computed in a parallel way. This paper analyzes the performance of multiple IIR filters using GPUs and compares it with a powerful multi-core computer. The proposed GPU implementation can run up to 1256 concurrent IIR filters of order 256th in real time, which means 321,536 total filter order, with a latency time of 0.72 ms (sampling frequency of 44.1 kHz). This demonstrates that GPUs are well suited for computing massive IIR filtering. Jose A. Belloch, Balázs Bank, Lauri Savioja, Alberto González 0001, Vesa Välimäki |
ICASSP | 1 |
| 2011 | A real-time crosstalk canceller on a notebook GPUabstractCrosstalk cancellation is one of the main applications in multichannel acoustic signal processing. This field has experienced a major development in recent years because of the in crease in the number of sound sources used in playback applications available to users. Developing these applications re quires high computing capabilities because of its high number of operations. Graphics Processor Unit (GPU), a high parallel commodity programmable co-processors, offer the possibility of parallelizing these operations. This allows to obtain the results in a much shorter time and also to free up CPU resources which can be used for other tasks. One important aspect lies in the possibility to overlap the data transfer from CPU to GPU and vice versa with the computation, in order to carry out real-time applications. Thus, this work focuses on two main points: to describe an efficient implementation of a crosstalk cancellation on GPU and to incorporate it into a real-time application. Jose A. Belloch, Alberto González 0001, Francisco-Jose Martínez-Zaldívar, Antonio M. Vidal |
ICME | 1 |
| 2011 | Real-time massive convolution for audio applications on GPU - Massive convolution on GPU
Jose A. Belloch, Alberto González 0001, Francisco-Jose Martínez-Zaldívar, Antonio M. Vidal |
J. Supercomput. | 1 |