Rubén Salvador

dblp:201/1377 · also Rubén Salvador Perea · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-0021-5808ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 eGoRG: GPU-accelerated depth estimation for immersive video applications based on graph cuts
abstract
Immersive video is gaining relevance across various fields, but its integration into real applications remains limited due to the technical challenges of depth estimation. Generating accurate depth maps is essential for 3D rendering, yet high-quality algorithms can require hundreds of seconds to produce a single frame. While real-time depth estimation solutions exist — particularly monocular deep learning-based methods and active sensors such as time-of-flight or plenoptic cameras — their depth accuracy and multiview consistency are often insufficient for depth image-based rendering (DIBR) and immersive video applications. This highlights the persistent challenge of jointly achieving real-time performance and high-quality, correlated depth across views. This paper introduces eGoRG, a GPU-accelerated depth estimation algorithm based on MPEG DERS, which employs graph cuts to achieve high-quality results. eGoRG contributes a novel GPU-based graph cuts stage, integrating block-based push-relabel acceleration and a simplified alpha expansion method. These optimizations deliver quality comparable to leading graph-cut approaches while greatly improving speed. Evaluation on an MPEG multiview dataset and a static NeRF dataset demonstrates the algorithm’s effectiveness across different scenarios. • The proposal is a novel GPU-accelerated depth estimation algorithm based on graph cuts. • Algorithm-dependent strategies are introduced to maximize the quality–time trade-off. • Depth results are comparable to high-performing graph-cut approaches while being substantially faster. • The method is training-free and can process dynamic scenes. • The algorithm is a good trade-off between quality and processing time achieving near real-time results.
Jaime Sancho, Manuel Villa, Miguel Chavarrías, Rubén Salvador, Eduardo Juárez Martínez, César Sanz
J. Vis. Commun. Image Represent.4
2026 Multi-Screaming-Channel Attacks: Frequency Diversity for Enhanced Attacks
abstract
Side-channel attacks consist of retrieving internal data from a victim system by analyzing its leakage, which usually requires proximity to the victim in the range of a few millimetres. Screaming channels are EM side channels transmitted at a distance of a few meters. They appear on mixed-signal devices integrating an RF module on the same silicon die as the digital part. Consequently, the side channels are modulated by legitimate RF signal carriers and appear at the harmonics of the digital clock frequency. While initial works have only considered collecting leakage at these harmonics, our work has demonstrated that the leakage is also present at frequencies other than these harmonics. This result significantly increases the number of available frequencies to perform a screaming-channel attack, which can be convenient in an environment where multiple harmonics are polluted. This paper studies how this diversity of frequencies carrying leakage can be used to improve attack performance. We first study how to combine multiple frequencies. Second, we demonstrate that frequency combination can improve attack performance and evaluate this improvement according to the performance of the combined frequencies. Finally, we demonstrate the interest of frequency combination in attacks at 15 and, for the first time, at 30 meters in an RF-polluted environment. One last important observation is that this frequency combination divides by at least 2 (and up to 3.76) the number of traces needed to reach a given attack performance.
Jeremy Guillaume, Maxime Pelcat, Amor Nafkha, Rubén Salvador
IEEE Trans. Inf. Forensics Secur.4
2025 Side-Channel Extraction of Dataflow AI Accelerator Hardware Parameters
abstract
Dataflow neural network accelerators efficiently process AI tasks on FPGAs, with deployment simplified by ready-to-use frameworks and pre-trained models. However, this convenience makes them vulnerable to malicious actors seeking to reverse engineer valuable Intellectual Property (IP) through Side-Channel Attacks (SCA). This paper proposes a methodology to recover the hardware configuration of dataflow accelerators generated with the FINN framework. Through unsupervised dimensionality reduction, we reduce the computational overhead compared to the state-of-the-art, enabling lightweight classifiers to recover both folding and quantization parameters. We demonstrate an attack phase requiring only 337 ms to recover the hardware parameters with an accuracy of more than 95% and 421 ms to fully recover these parameters with an averaging of 4 traces for a FINN-based accelerator running a CNN, both using a random forest classifier on side-channel traces, even with the accelerator dataflow fully loaded. This approach offers a more realistic attack scenario than existing methods, and compared to SoA attacks based on tsfresh, our method requires 940x and 110x less time for preparation and attack phases, respectively, and gives better results even without averaging traces.
Guillaume Lomet, Rubén Salvador, Brice Colombier, Vincent Grosso, Olivier Sentieys, Cédric Killian
IOLTS2
2025 Detecting Hardware Trojans in Microprocessors via Hardware Error Correction Code-based Modules
abstract
Software-exploitable Hardware Trojans (HTs) en-able attackers to execute unauthorized software or gain illicit access to privileged operations. This manuscript introduces a hardware-based methodology for detecting runtime HT activations using Error Correction Codes (ECCs) on a RISC-V mi-croprocessor. Specifically, it focuses on HTs that inject malicious instructions, disrupting the normal execution flow by triggering unauthorized programs. To counter this threat, the manuscript introduces a Hardware Security Checker (HSC) leveraging Hamming Single Error Correction (HSEC) architectures for effective HT detection. Experimental results demonstrate that the proposed solution achieves a 100% detection rate for potential HT activations, with no false positives or undetected attacks. The implementation incurs minimal overhead, requiring only 72 #LUTs, 24 #FFs, and 0.5 #BRAM while maintaining the microprocessor's original operating frequency and introducing no additional time delay.
Alessandro Palumbo, Rubén Salvador
IOLTS2
2023 Attacking at Non-harmonic Frequencies in Screaming-Channel Attacks
Jeremy Guillaume, Maxime Pelcat, Amor Nafkha, Rubén Salvador
CARDIS4
2018 Automatic instrumentation of dataflow applications using PAPI
abstract
The widening of the complexity-productivity gap witnessed in the last years is becoming unaffordable from the application development point of view. New design methods try to automate most designers tasks in order to bridge this gap. In addition, new Models of Computation (MoC), as those dataflow-based, ease the expression of parallelism within applications and lead to higher productivity.
Daniel Madroñal, Antoine Morvan, Raquel Lazcano, Rubén Salvador, Karol Desnos, Eduardo Juárez Martínez, César Sanz
CF4
2017 Porting a PCA-based hyperspectral image dimensionality reduction algorithm for brain cancer detection on a manycore architecture
Raquel Lazcano, Daniel Madroñal, Rubén Salvador, Karol Desnos, Maxime Pelcat, Raúl Guerra, Himar Fabelo, Samuel Ortega, Sebastián López, Gustavo M. Callicó, Eduardo Juárez Martínez, César Sanz
J. Syst. Archit.3
2017 SVM-based real-time hyperspectral image classifier on a manycore architecture
Daniel Madroñal, Raquel Lazcano, Rubén Salvador, Himar Fabelo, Samuel Ortega, Gustavo M. Callicó, Eduardo Juárez Martínez, César Sanz
J. Syst. Archit.3
2013 Self-Reconfigurable Evolvable Hardware System for Adaptive Image Processing
abstract
This paper presents an evolvable hardware system, fully contained in an FPGA, which is capable of autonomously generating digital processing circuits, implemented on an array of processing elements (PEs). Candidate circuits are generated by an embedded evolutionary algorithm and implemented by means of dynamic partial reconfiguration, enabling evaluation in the final hardware. The PE array follows a systolic approach, and PEs do not contain extra logic such as path multiplexers or unused logic, so array performance is high. Hardware evaluation in the target device and the fast reconfiguration engine used yield smaller reconfiguration than evaluation times. This means that the complete evaluation cycle is faster than software-based approaches and previous evolvable digital systems. The selected application is digital image filtering and edge detection. The evolved filters yield better quality than classic linear and nonlinear filters using mean absolute error as standard comparison metric. Results do not only show better circuit adaptation to different noise types and intensities, but also a nondegrading filtering behavior. This means they may be run iteratively to enhance filtering quality. These properties are even kept for high noise levels (40 percent). The system as a whole is a step toward fully autonomous, adaptive systems.
Rubén Salvador, Andrés Otero, Javier Mora 0001, Eduardo de la Torre, Teresa Riesgo, Lukás Sekanina
IEEE Trans. Computers1
2012 Implementation techniques for evolvable HW systems: virtual VS. dynamic reconfiguration
abstract
Adaptive hardware requires some reconfiguration capabilities. FPGAs with native dynamic partial reconfiguration (DPR) support pose a dilemma for system designers: whether to use native DPR or to build a virtual reconfigurable circuit (VRC) on top of the FPGA which allows selecting alternative functions by a multiplexing scheme. This solution allows much faster reconfiguration, but with higher resource overhead. This paper discusses the advantages of both implementations for a 2D image processing matrix. Results show how higher operating frequency is obtained for the matrix using DPR. However, this is compensated in the VRC during evolution due to the comparatively negligible reconfiguration time. Regarding area, the DPR implementation consumes slightly more resources due to the reconfiguration engine, but adds further more capabilities to the system.
Rubén Salvador, Andrés Otero, Javier Mora 0001, Eduardo de la Torre, Teresa Riesgo, Lukás Sekanina
FPL1
2010 High Level Validation of an Optimization Algorithm for the Implementation of Adaptive Wavelet Transforms in FPGAs
abstract
The work reported in this paper describes the steps given towards an FPGA-based implementation of evolvable wavelet transforms for image compression in embedded systems. An Evolutionary Algorithm (EA) for the design and optimization of the transform coefficients is tailored for a suitable System on Chip implementation. Several cut downs on the computing requirements have been done to the original algorithm, adapting it for the FPGA implementation. What this paper addresses more specifically is the validation of the algorithm using fixed point arithmetic for the whole optimization process. The results show how high quality transforms are evolved from scratch with limited precision arithmetic. Also, preliminary results of the implementation in an FPGA device are included.
Rubén Salvador, Félix Moreno, Teresa Riesgo, Lukás Sekanina
DSD1