Juan F. Sanjuan

dblp:284/3889 · also J. F. Sanjuan-Estrada, Juan Francisco Sanjuan, Juan Francisco Sanjuan-Estrada · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-2874-0903ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 A parallel framework for data input pipelines and online data augmentation in deep learning
abstract
Abstract Efficient data ingestion and online data augmentation remain challenges in deep learning workflows, particularly when dealing with datasets containing non-standard formats or massive multidimensional arrays that natively optimised functions cannot fully manage. This work presents a parallel framework that integrates and global shared memory through a ring buffer architecture, enabling high-throughput data loading and flexible on-the-fly augmentation. The framework decouples data production from consumption, allowing multiple CPU workers to load and preprocess batches in parallel while completely bypassing the Python GIL and memory bottlenecks. Crucially, the framework supports both CPU-side and GPU-side augmentation strategies, adapting to whether complex conditional transformations or framework-native operations are required. The proposed approach was validated on two representative tasks: (i) sign language recognition from human pose CSV sequences, and (ii) hyperspectral image classification using massive arrays. Relative to standard sequential baselines, the proposed framework achieved up to $$27\times $$ 27 × acceleration in isolated data ingestion and up to $$28\times $$ 28 × in end-to-end training. Importantly, even against natively optimised parallel TensorFlow and PyTorch pipelines, it still delivered up to $$8\times $$ 8 × faster data loading and up to $$7\times $$ 7 × faster full training in memory-intensive scenarios. Overall, the proposed framework provides a scalable, multi-GPU compatible solution for deep learning pipelines, showing robust performance across both I/O-bound and memory-constrained scenarios in TensorFlow and PyTorch while alleviating memory fragmentation and allocation constraints.
Antonio De Toro-Castro, Marcos Lupión, Vicente González Ruiz, Juan F. Sanjuan, Pilar Martínez Ortigosa
J. Supercomput.4
2023 THPoseLite, a Lightweight Neural Network for Detecting Pose in Thermal Images
abstract
Nowadays, smart environments (SEs) enable the monitoring of people with physical disabilities by incorporating activity recognition. Thermal cameras are being incorporated as they preserve privacy. Some deep learning (DL) solutions use the pose of the users because it removes external noise. Although there are robust DL solutions in the visible spectrum (VS), they fail in the thermal domain. Thus, we propose thermal human pose lite (THPoseLite), a convolutional neural network (CNN) based on MobileNetV2 that extracts pose from thermal images (TIs). In a novel way, an auto-labeling approach has been developed. It includes a background removal using an optical flow estimator. It also integrates Blazepose [a pose estimator for VS images (VSIs)] to obtain the poses in the preprocessed TIs. Results show that the preprocessing increases the percentage of detected poses by Blazepose from 19.55% to 76.85%. This allows the recording of human pose estimation (HPE) data sets in the VS without requiring VS cameras or manually annotating data sets. Furthermore, THPoseLite has been embedded in an Internet of Things (IoT) device incorporating an edge tensor processing unit (TPU) accelerator, which can process TIs recorded at 9 frames per second (FPS) in real time (12.28 FPS). It requires fewer than 6W of energy to run. It has been achieved using model quantization, decreasing the accuracy in estimating the poses by only 1%. The mean-squared error of MobileNetV2 in test images is 35.48, obtaining accurate poses in 21% of the images that Blazepose is not able to detect any pose.
Marcos Lupión, Vicente González Ruiz, Javier Medina 0001, Juan F. Sanjuan, Pilar Martínez Ortigosa
IEEE Internet Things J.4
2023 Accelerating neural network architecture search using multi-GPU high-performance computing
Marcos Lupión, Nicolas C. Cruz, Juan F. Sanjuan, Ben Paechter, Pilar Martínez Ortigosa
J. Supercomput.3
2022 On the limits of Conditional Generative Adversarial Neural Networks to reconstruct the identification of inhabitants from IoT low-resolution thermal sensors
abstract
One of the main objectives of smart homes is to facilitate daily life by increasing user comfort, with the potential to play a key role in revolutionizing healthcare for the elderly, the disabled and people with functional limitations. To achieve this end, smart homes will have to be able to distinguish the identity of users, their location and the activities they are performing, while also being implemented in a non-invasive way that protects the privacy of these users. Computer vision is one of the main technologies included in smart homes. However, there are drawbacks to traditional cameras, given their dependence on light and privacy-related concerns. Thermal cameras provide a solution, as they operate regardless of light conditions (e.g. at night) while respecting users’ privacy. In this work, image reconstruction and identification of inhabitants from facial images collected by low-resolution thermal sensors has been carried out by using Conditional Generative Adversarial Neural Networks (CGANs). The system has been implemented through an IoT device with raspberry Pi and dual-vision thermal and visible-spectrum sensors installed in a real smart home to automatically collect paired visible-spectrum and thermal images. Thus, different configurations of CGANs have been implemented and analyzed to achieve the following outcomes: (1) inhabitant identification (normal and masked face) with enhanced user privacy and (2) transfer from thermal to color images in the visible spectrum. Results show that the proposed CGAN achieves a recognition rate of 95% and 94% for uncovered and masked faces. This enables user identification without registering accurate facial expressions in the color image reconstruction, protecting user privacy. Furthermore, the developed system outperformed similar approaches using low-resolution datasets and has demonstrated that the accuracy of image reconstruction depends on the resolution of the input visible-spectrum images. In addition, a contribution to high-performance computing has been made by designing a CGAN that runs efficiently on multiple GPUs, achieving increased performance and response speed of the network, as well as applicability to larger problems.
Marcos Lupión, Aurora Polo Rodríguez, Javier Medina 0001, Juan F. Sanjuan, Pilar Martínez Ortigosa
Expert Syst. Appl.4
2022 Using a Multi-GPU node to accelerate the training of Pix2Pix neural networks
abstract
Abstract Generative adversarial networks are gaining importance in problems such as image conversion, cross-domain translation and fast styling. However, the training of these networks remains unclear because it often results in unexpected behavior caused by non-convergence, model collapse or overly long training, causing the training task to have to be supervised by the user and vary with each dataset. To increase the speed of training in Pix2Pix (image-to-image translation) networks, this work incorporates multi-GPU training using mixed precision, along with optimizations in the GPU image input process. In addition, in order to make the training unsupervised and to terminate it when the best transformations are performed, an early stopping method using the peak signal noise ratio (PSNR) metric is proposed.
Marcos Lupión, Juan F. Sanjuan, Pilar Martínez Ortigosa
J. Supercomput.2
2012 Performance Driven Cooperation between Kernel and Auto-tuning Multi-threaded Interval B&B Applications
Juan F. Sanjuan, Leocadio G. Casado, Inmaculada García, Eligius M. T. Hendrix
ICCSA (1)1
2011 Adaptive Parallel Interval Global Optimization Algorithms Based on their Performance for Non-dedicated Multicore Architectures
abstract
Branch and Bound (B&B) algorithms are highlyparallelizable but they are irregular and dynamic load balancing techniques have been used to avoid idle processors. In previous work, authors use a dynamic number of threads at run time, which depends on the measured performance of the application for just one interval B&B algorithm running on the system. In this way, load balancing is achieved by thread generation decisions. In this work, we extend the study of these models to non-dedicated systems. In order to have a controlled test bed and comparable results, several instances of the interval global optimization algorithm are executed in the system, with the same model and problem to solve. Therefore, a non-dedicated system is simulated because the execution of one application affects the execution of the other instances. This paper discusses different methods and models to decide when a thread should be created. Experiments show which of the proposed methods performs best in terms of maximum running time per application, using the fewest running threads. Following this parallel programming methodology, which is well suited for other B&B codes, applications can adapt their parallelism level to their performance and load of the system(at run time). This work represents a step forward towards increasing the performance of parallel algorithm running inn on-dedicated and heterogeneous systems. The adaptive model discussed in this work is able to reduce the overall execution time for a set of instances of the same application running simultaneously. It also exempts the user from specifying the number of threads each application should use.
Juan F. Sanjuan, Leocadio G. Casado, Inmaculada García
PDP1
2011 Adaptive parallel interval branch and bound algorithms based on their performance for multicore architectures
Juan F. Sanjuan, Leocadio G. Casado, Inmaculada García
J. Supercomput.1