German Leon

dblp:25/3994 · also Germán León · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-7687-2126ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Energy and robustness trade-offs in adaptive neural mmWave channel estimation on edge devices
abstract
Abstract The evolution toward 6G will continue to leverage massive multiple-input multiple-output and millimeter-wave systems, which demand accurate angle-of-arrival (AoA) and angle-of-departure (AoD) estimation. While several deep learning models have demonstrated strong performance for this task, their accuracy, like that of most estimation methods, is often degraded by hardware non-idealities, which can be further exacerbated by time-varying operational factors such as component aging and adverse weather, among others. Building on a pre-trained U-Net architecture with demonstrated competitive performance for AoA/AoD estimation, we first propose an adaptation mechanism based on fine-tuning with impairment-augmented data. Specifically, we simulate hardware imperfections by introducing random phase errors in the antenna elements, ranging from mild fluctuations to severe signal distortions. The U-Net model with adaptation capabilities is then implemented on an NVIDIA Jetson Orin Nano device, a compact edge platform with heterogeneous computing resources. To this end, we design a co-execution strategy that performs AoA/AoD estimation (inference) on the CPU while simultaneously fine-tuning the model on the GPU, thus enabling continuous model adaptation to changing environmental or hardware conditions while preserving real-time inference performance. Experimental results show that impairment-aware fine-tuning effectively counters hardware degradation, particularly under significant phase impairments. In such scenarios, the fine-tuned model consistently preserves or even improves estimation accuracy, reducing the Root Mean Square Error (RMSE) by approximately 3.6% and increasing the Probability of Detection ( $$P_D$$ P D ) by up to 1 percentage point compared to the base model. Furthermore, a detailed energy-performance analysis demonstrates that while maximum frequency settings reduce training time by over 11 $$\times $$ × , they also increase power consumption by more than 5 $$\times $$ × , with optimal energy efficiency achieved at mid-range CPU and high GPU frequencies. This work establishes the feasibility of concurrent training and inference on resource-constrained heterogeneous hardware, paving the way for resilient and autonomous edge intelligence in future 6G systems.
Eric Meneses Albalá, Saúl Villaescusa, José M. Badía, German Leon, Carmen Botella-Mascarell, Sandra Roger 0002
J. Supercomput.4
2026 Accelerated deep learning denoising for edge AI in single-pixel imaging
abstract
Abstract Single-pixel imaging (SPI) uses pixel arrays with coded illumination and a single detector to reconstruct an image of a scene. However, it is time-consuming due to the many projected patterns and non-trivial reconstruction. We pursue a practical route to real-time SPI by pairing fast linear reconstructions with a compact U-Net denoiser on an embedded GPU. In a fully simulated pipeline, we form undersampled reconstructions from CelebA-derived $$64\times 64$$ 64 × 64 grayscale faces using Hadamard patterns ordered by Cake Cutting at $$M/N\in \{4,8,16,24,32\%\}$$ M / N ∈ { 4 , 8 , 16 , 24 , 32 % } . The denoiser is trained with mean squared error (MSE) on normalized images and deployed on a Jetson Orin NX 16 GB (FP32). We measure the acquisition time, reconstruction time, and inference time of the U-Net model per frame. Quality is reported with PSNR and SSIM, and performance with serial latency and pipelined throughput. Results indicate that GPU inference markedly cuts denoising time, shifting the bottleneck toward optical acquisition and/or the linear step as sampling grows. The study offers a reproducible recipe–data generation, models, and timing methodology–for assessing SPI denoising on edge hardware, and outlines levers to raise throughput: higher-rate pattern projection, optimized reconstruction kernels, and right-sized U-Net variants that preserve PSNR/SSIM while lowering latency.
Carlos Chabert-Ull, Heberley Tobón-Maya, Samuel I. Zapata-Valencia, Enrique Tajahuerce, German Leon
J. Supercomput.5
2026 Dependability analysis and hardening of vision transformers against soft errors
abstract
Abstract The deployment of Vision Transformers (ViTs) in safety-critical domains needs a clear understanding of their resilience to soft errors, since their specific layer-level vulnerabilities are currently insufficiently characterized. This work presents a dependability analysis of the ViT-Base architecture against injection-induced soft errors. Using a high-fidelity, software-level fault injection methodology with custom CUDA kernels, the study injects random bit-flips directly into the IEEE 754 binary32 floating-point representation of the intermediate data tensors resulting from the Transformer modules to quantify model accuracy degradation across increasing bit error rates. As a primary result, a vulnerability map across ViT layers is presented, confirming that the results of normalization and fully connected layers exhibit critical sensitivity to soft errors. To address these vulnerabilities, the work evaluates targeted hardening strategies. These include Fault-Aware Training (FAT), applied both globally and selectively to linear layers, as well as practical runtime mitigations such as range-based value clipping and filtering of non-numeric values. The findings demonstrate that these software-only approaches can significantly protect model accuracy.
Lester Frias-Dominguez, José M. Badía, German Leon, Adrian Amor-Martin, Jose A. Belloch
J. Supercomput.3
2025 Optimizing Millimeter Wave MIMO Channel Estimation Through GPU-Based Edge Artificial Intelligence
abstract
In the context of upcoming sixth-generation (6G) wireless communication systems, the use of millimeter wave (mmWave) frequencies is a key technology for achieving high-throughput communications. Accurate parametric estimation of mmWave channels is critical for effective beamforming design and configuration, requiring sophisticated models to capture the directional characteristics of these channels. This work considers an innovative artificial intelligence (AI) approach for accurate estimation of angle-of-arrival (AoA) and angle-of-departure (AoD) parameters from frequency-domain channel observations. Our approach is based on the implementation of two convolutional neural networks (CNNs): a residual CNN (ResNet) and a U-Net CNN. Specifically, this work focuses on the efficient implementation of both schemes in an embedded system suitable for edge AI. We performed the experiments in a low-power NVIDIA Jetson Orin Nano platform and evaluated the effect of modifying the frequencies of its CPU and GPU on the performance of the inference process, both in terms of execution time and energy consumption. Experimental results showed that the U-Net model is more power consuming, but as it is faster, it consumes less energy per channel.
Diego Lloria, Sandra Roger 0002, German Leon, José M. Badía, Carmen Botella-Mascarell, Jose A. Belloch
J. Supercomput.3
2025 Evaluating and accelerating vision transformers on GPU-based embedded edge AI systems
abstract
Abstract Many current embedded systems comprise heterogeneous computing components including quite powerful GPUs, which enables their application across diverse sectors. This study demonstrates the efficient execution of a medium-sized self-supervised audio spectrogram transformer (SSAST) model on a low-power system-on-chip (SoC). Through comprehensive evaluation, including real time inference scenarios, we show that GPUs outperform multi-core CPUs in inference processes. Optimization techniques such as adjusting batch size, model compilation with TensorRT, and reducing data precision significantly enhance inference time, energy consumption, and memory usage. In particular, negligible accuracy degradation is observed, with post-training quantization to 8-bit integers showing less than 1% loss. This research underscores the feasibility of deploying transformer neural networks on low-power embedded devices, ensuring efficiency in time, energy, and memory, while maintaining the accuracy of the results.
Ignacio Martin-Salinas, José M. Badía, Óscar Valls, German Leon, Rocío del Amor, Jose A. Belloch, Adrian Amor-Martin, Valery Naranjo
J. Supercomput.4
2024 Urban sound classification using neural networks on embedded FPGAs
abstract
Abstract Sound classification using neural networks has recently produced very accurate results. A large number of different applications use this type of sound classifiers such as controlling and monitoring the type of activity in a city or identifying different types of animals in natural environments. While traditional acoustic processing applications have been developed on high-performance computing platforms equipped with expensive multi-channel audio interfaces, the Internet of Things (IoT) paradigm requires the use of more flexible and energy-efficient systems. Although software-based platforms exist for implementing general-purpose neural networks, they are not optimized for sound classification, wasting energy and computational resources. In this work, we have used FPGAs to develop an ad hoc system where only the hardware needed for our application is synthesized, resulting in faster and more energy-efficient circuits. The results show that our developments are accelerated by a factor of 35 compared to a software-based implementation on a Raspberry Pi.
Jose A. Belloch, Raul Coronado, Óscar Valls, Rocío del Amor, German Leon, Valery Naranjo, Manuel F. Dolz, Adrian Amor-Martin, Gema Piñero
J. Supercomput.5
2024 Comparative analysis of soft-error sensitivity in LU decomposition algorithms on diverse GPUs
abstract
Abstract Graphics processing units (GPUs) have become integral to embedded systems and supercomputing centres due to their large memory, cutting-edge technology and high performance per watt. However, their susceptibility to transient errors requires a comprehensive analysis of error sensitivity, as well as the development of error mitigation techniques and fault-tolerant algorithms. This study focuses on evaluating the soft-error sensitivity of two distinct versions of LU decomposition algorithms implemented on two very different GPUs—a low-power SoC embedded GPU and a high-performance massively parallel GPU. Through extensive fault injection campaigns on both GPUs, we examine the vulnerability of the algorithms, identify error causes, and determine critical code components requiring enhanced protection. The experiments reveal that most single bit flip fault injections in the instruction results lead to erroneous outcomes or unrecoverable errors. Notably, efficient GPU resource utilisation can increase the number of masked errors, thereby enhancing error resilience. Additionally, while different parts of the code exhibit similar error occurrence types and rates, the propagation of errors to elements within the result matrix differs significantly.
German Leon, José M. Badía, Jose A. Belloch, Almudena Lindoso, Luis Entrena
J. Supercomput.1
2022 Multicore implementation of a multichannel parallel graphic equalizer
abstract
Abstract Numerous signal processing applications are emerging on mobile computing systems. These applications are subject to responsiveness constraints for user interactivity and, at the same time, must be optimized for energy efficiency. Many current embedded devices are composed of low-power multicore processors that offer a good trade-off between computational capacity and low power consumption. In this context, equalizers are widely used in multiple mobile-based applications such as “Music streaming” to adjust the levels of bass and treble in sound reproduction. In this study, we evaluate a graphic equalizer from audio, computational capacity, and energy efficiency perspectives, as well as the execution of multiple real-time equalizers running on an embedded quad-core processor of a mobile device. To this end, we experiment with the working frequencies as well as the parallelism that can be extracted from a quad-core ARM Cortex-A57. Results show that using high CPU frequencies and three or four cores, our parallel algorithm is able to equalize more than five channels per watt in real time with an audio buffer of 4096 samples, which implies a latency of 92.8 ms at the standard sample rate of 44.1 kHz.
Jose A. Belloch, José M. Badía, German Leon, Balázs Bank, Vesa Välimäki
J. Supercomput.3
2021 Evaluating the computational performance of the Xilinx Ultrascale+ EG Heterogeneous MPSoC
Jose A. Belloch, German Leon, José M. Badía, Almudena Lindoso, Enrique San Millán
J. Supercomput.2
2019 Noise estimation for hyperspectral subspace identification on FPGAs
German Leon, Carlos González 0002, Rafael Mayo 0002, Daniel Mozos, Enrique S. Quintana-Ortí
J. Supercomput.1
2015 Unveiling the performance-energy trade-off in iterative linear system solvers for multithreaded processors
abstract
Summary In this paper, we analyze the interactions occurring in the triangle performance‐power‐energy for the execution of a pivotal numerical algorithm, the iterative conjugate gradient (CG) method, on a diverse collection of parallel multithreaded architectures. This analysis is especially timely in a decade where the power wall has arisen as a major obstacle to build faster processors. Moreover, the CG method has recently been proposed as a complement to the LINPACK benchmark, as this iterative method is argued to be more archetypical of the performance of today's scientific and engineering applications. To gain insights about the benefits of hands‐on optimizations we include runtime and energy efficiency results for both out‐of‐the‐box usage relying exclusively on compiler optimizations, and implementations manually optimized for target architectures, that range from general‐purpose and digital signal multicore processors to manycore graphics processing units, all representative of current multithreaded systems. Copyright © 2014 John Wiley & Sons, Ltd.
José Ignacio Aliaga, Hartwig Anzt, María Isabel Castillo, Juan Carlos Fernández 0002, German Leon, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.5
2015 Exploring the performance-power-energy balance of low-power multicore and manycore architectures for anomaly detection in remote sensing
German Leon, José M. Molero, Ester M. Garzón, Inmaculada García, Antonio Plaza, Enrique S. Quintana-Ortí
J. Supercomput.1
2010 A reconfigurable platform for evaluating the performance of QoS networks
José M. Claver, P. Agustí, Miguel Arevalillo-Herráez, German Leon, Manel Canseco
J. Syst. Archit.4
2008 The UJI industrial robotics telelaboratory: Real-time vision and networking
abstract
In this video we present a work in progress in the UJI (i.e. the acronym for University Jaume I) robotics telelaboratory. This telelaboratory uses a remote control system based on networked robots and FPGAs technology. The main devices included in this cell are: a SCARA manipulator (AdeptOne), a robot arm with six degrees of freedom (Motoman), an industrial belt, several sensors and cameras, an FPGA that takes care of the computer vision algorithms (i.e. including grasping determination), and a distributed architecture that allows any user to control remotely via Internet a specific manufacturing task. The different components of this system are connected by a 100BaseT Ethernet network and follow the SNRP architecture (i.e. Simple Network Robot Protocol), which permits the integration of network robots and sensors within an e-learning platform in a simple and reliable manner. The whole telelaboratory is connected to the Internet through a router that permits the user to control the networked devices according to security constraints. This distributed architecture allows any user to control remotely via Internet a specific manufacturing task.
Jorge Sales, Reinel Beltran, Pedro J. Sanz, Raúl Marín Prades, Raul Wirz, German Leon, José M. Claver, Jaime Alemany
IROS6
2007 A Reprogrammable and Scalable Multimedia Traffic Generator/Monitor on FPGA
abstract
Nowadays, high performance System/Local Area Networks (SAN/LAN) are filled by heterogeneous traffic, consisting of information flows with different bandwidth and latency requirements. The bottleneck produced by the throughput of network processing elements (servers, routers,...) and the increasing bandwidth of links, makes it necessary to propose new designs for these network components. The MMR (and its simplified version, the SMMR), a router that supports QoS, is a very well-known proposal in this area. In this article we propose the architecture and implementation of a hardware reprogrammable traffic multimedia Generator/Monitor (G/M) to study these sorts of routers under different traffic conditions and server models.
José M. Claver, P. Agustí, German Leon, Manel Canseco
FPL3
2007 High Level Power Optimization by Type Inference on the Generation of Application Specific Circuits on FPGAs
abstract
We describe the optimization of power consumption obtained by a high level environment developed for the automatic generation of application specific circuits on FPGA. The methodology used is based on the transformation of the whole algorithm in a graph of LUTs that implements all the required operations without the use of library components. The quality of the obtained circuitry is guaranteed by the use of "type inference". Our environment automatically optimizes the word-length and size of operators, and at the same time, reduces the internal data paths and the switching activity. Thus, in the extreme cases testes, the resulting generated circuits offer an important improvement in area usage of up to 95%, and power consumption is reduced by up to 98%.
José M. Claver, German Leon
FPL2
2006 A Hardware NIC Scheduler to Guarantee QoS on High Performance Servers
José M. Claver, Manel Canseco, P. Agustí, German Leon
ISPA4
1999 Parity Sensitive Comparators
abstract
Parity sensitive comparators are a new type of comparators designed to take advantage of the parity information present in most buses. Instead of simply comparing the signals carried by the buses, parity information is used to select the probably correct output in case of mismatch, thus avoiding an important percentage of errors to stop system functioning. These devices verify parity for each pair of compared groups and discard erroneous data when comparing, propagating the probably correct data and signaling the fact with a "recoverable error" output signal. Benefits of parity sensitive comparators are analyzed by means of a deep probabilistic study of the "parity bit per data byte" case that shows how they are able to recover more than the 97% of errors signaled by classical comparators. The output information of the comparators is also described, as well as its uses according to the level of reliability required.
Germán Fabregat, Jose V. Marti, German Leon
PRDC3