Nicolás Guil

dblp:g/NicolasGuilMata · also Nicolás Guil Mata · DBLP profile ↗
← Back
68ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-3431-6516ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 4 since 2021Systems, architecture and hardware · 21 · 1 first-author · 5 since 2021Security and privacy · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 End-to-End Multitask CNN-Based Model for Palm Vein Biometrics
José I. Santamaría, Franco Lara, Francisco M. Castro, Ricardo J. Barrientos, Nicolás Guil, Ruber Hernández-García
CIARP (2)5
2025 Leveraging Implicit 3D Geometry for Biometric and Anthropometric Estimation from Gait
abstract
Estimating biometric and anthropometric attributes from gait sequences presents a promising alternative to traditional body measurement techniques, particularly in unconstrained or low-resource environments. However, inferring metric attributes from 2D silhouette-based gait representations remains challenging due to the lack of volumetric cues. In this work, we propose a novel training strategy that leverages the geometric priors encoded by PIFuHD—a high-resolution implicit function-based model for 3D human reconstruction—to inject structural supervision into a gait encoder. Our method introduces a feature reconstruction branch that distills volumetric knowledge from precomputed PIFuHD embeddings during training, enabling accurate prediction of anthropometric attributes at inference time from 2D silhouette sequences alone. We evaluate our approach on the Health&Gait dataset, achieving substantial improvements over baselines in predicting multiple attributes, including height, weight, BMI, and body circumferences. These results demonstrate that shape-aligned supervision from implicit 3D models can effectively bridge the gap between geometric reasoning and efficient biometric estimation from visual gait data.
Nicolás Cubero, Jorge Zafra-Palma, Francisco M. Castro, Nicolás Guil, Manuel J. Marín-Jiménez
IJCB4
2025 Empirical study of human pose representations for gait recognition
abstract
Gait recognition has gained attention for its ability to identify individuals from afar. Current state-of-the-art approaches predominantly utilize visual information, such as silhouettes, or a combination of visual data and basic body pose information, including skeleton joint coordinates. However, the role of human pose in gait recognition is still underexplored, often leading to poorer results compared to visual approaches. In this work, we propose a novel hierarchical limb-based representation that enhances the depiction of body pose and can be applied to various pose descriptors. Our representation consists of three hierarchical levels: full body, body limbs (arms and legs), and middle limbs (forearms, lower arms, thighs, and shins). This structure enriches the gait description of the overall pose by incorporating the specific movements of each limb. Particularly, we investigate the application of our hierarchical arrangement using two different rich pose descriptors: heatmaps derived from 2D body skeletons and a dense representation obtained from pixel-wise estimation of body pose ( i.e DensePose). Furthermore, we introduce the PoseGaitGL family of models to better leverage the features derived from our pose representations. By employing our hierarchical pose representations, the proposed model achieves state-of-the-art results in pose-based gait recognition. Thus, the hierarchical heatmap-based and hierarchical DensePose representations attain Rank-1 accuracy of 82.2% and 92.0%, respectively, on the cross-view setup of CASIA-B, and 99.3% and 99.8%, respectively, on TUM-GAID, establishing a new benchmark for pose-based methods. Source code is available at https://github.com/Nico-Cubero/PoseGaitGL .
Nicolás Cubero, Francisco M. Castro, Julián Ramos Cózar, Nicolás Guil, Manuel J. Marín-Jiménez
Expert Syst. Appl.4
2025 Real-time unsupervised video object detection on the edge
abstract
Object detection in video is an essential computer vision task. Consequently, many efforts have been devoted to developing precise and fast deep-learning models for this task. These models are commonly deployed on discrete and powerful GPU devices to meet both frame rate performance and detection accuracy requirements. Furthermore, model training is usually performed in a strongly supervised way so that samples must be previously labelled by humans using a slow and costly process. In this paper, we develop a real-time implementation for unsupervised object detection in video employing a low-power device. We improve typical approaches for object detection using information supplied by optical flow to detect moving objects. Besides, we use an unsupervised clustering algorithm to group similar detections that avoid manual object labelling. Finally, we propose a methodology to optimize the deployment of our resulting framework on an embedded heterogeneous platform. Thus, we illustrate how all the computational resources of a Jetson AGX Xavier (CPU, GPU, and DLAs) can be used to fulfil frame rate, accuracy, and energy consumption requirements. Three different data representations (FP32, FP16 and INT8) are studied for the pipeline networks in order to evaluate the impact of all of them in our pipeline. Obtained results show that our proposed optimizations can improve up to 23 . 6 × energy consumption and 32 . 2 × execution time with respect to the non-optimized pipeline without penalizing the original mAP (59.44). This computational complexity reduction is achieved through knowledge distillation, using FP16 data precision, and deploying concurrent tasks in different computing units.
Paula Ruiz-Barroso, Francisco M. Castro, Nicolás Guil
Future Gener. Comput. Syst.3
2025 Lightweight Structure-Aware Attention for Visual Understanding
Heeseung Kwon, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Karteek Alahari
Int. J. Comput. Vis.4
2024 AttenGait: Gait recognition with attention and rich modalities
Francisco M. Castro, Rubén Delgado-Escaño, Ruber Hernández-García, Manuel J. Marín-Jiménez, Nicolás Guil
Pattern Recognit.5
2023 Irregular alignment of arbitrarily long DNA sequences on GPU
abstract
Abstract The use of Graphics Processing Units to accelerate computational applications is increasingly being adopted due to its affordability, flexibility and performance. However, achieving top performance comes at the price of restricted data-parallelism models. In the case of sequence alignment, most GPU-based approaches focus on accelerating the Smith-Waterman dynamic programming algorithm due to its regularity. Nevertheless, because of its quadratic complexity, it becomes impractical when comparing long sequences, and therefore heuristic methods are required to reduce the search space. We present GPUGECKO, a CUDA implementation for the sequential, seed-and-extend sequence-comparison algorithm, GECKO. Our proposal includes optimized kernels based on collective operations capable of producing arbitrarily long alignments while dealing with heterogeneous and unpredictable load. Contrary to other state-of-the-art methods, GPUGECKO employs a batching mechanism that prevents memory exhaustion by not requiring to fit all alignments at once into the device memory, therefore enabling to run massive comparisons exhaustively with improved sensitivity while also providing up to 6x average speedup w.r.t. the CUDA acceleration of BLASTN.
Esteban Pérez-Wohlfeil, Oswaldo Trelles, Nicolás Guil
J. Supercomput.3
2022 A Hybrid Piece-Wise Slowdown Model for Concurrent Kernel Execution on GPU
Bernabé López-Albelda, Francisco M. Castro, José María González-Linares, Nicolás Guil
Euro-Par4
2022 CAVLCU: an efficient GPU-based implementation of CAVLC
abstract
Abstract CAVLC (Context-Adaptive Variable Length Coding) is a high-performance entropy method for video and image compression. It is the most commonly used entropy method in the video standard H.264. In recent years, several hardware accelerators for CAVLC have been designed. In contrast, high-performance software implementations of CAVLC (e.g., GPU-based) are scarce. A high-performance GPU-based implementation of CAVLC is desirable in several scenarios. On the one hand, it can be exploited as the entropy component in GPU-based H.264 encoders, which are a very suitable solution when GPU built-in H.264 hardware encoders lack certain necessary functionality, such as data encryption and information hiding. On the other hand, a GPU-based implementation of CAVLC can be reused in a wide variety of GPU-based compression systems for encoding images and videos in formats other than H.264, such as medical images. This is not possible with hardware implementations of CAVLC, as they are non-separable components of hardware H.264 encoders. In this paper, we present CAVLCU, an efficient implementation of CAVLC on GPU, which is based on four key ideas. First, we use only one kernel to avoid the long latency global memory accesses required to transmit intermediate results among different kernels, and the costly launches and terminations of additional kernels. Second, we apply an efficient synchronization mechanism for thread-blocks (In this paper, to prevent confusion, a block of pixels of a frame will be referred to as simply block and a GPU thread block as thread-block.) that process adjacent frame regions (in horizontal and vertical dimensions) to share results in global memory space. Third, we exploit fully the available global memory bandwidth by using vectorized loads to move directly the quantized transform coefficients to registers. Fourth, we use register tiling to implement the zigzag sorting, thus obtaining high instruction-level parallelism. An exhaustive experimental evaluation showed that our approach is between 2.5 $$\times$$ × and 5.4 $$\times$$ × faster than the only state-of-the-art GPU-based implementation of CAVLC.
Antonio Fuentes-Alventosa, Juan Gómez-Luna, José María González-Linares, Nicolás Guil, Rafael Medina Carnicer
J. Supercomput.4
2022 FlexSched: Efficient scheduling techniques for concurrent kernel execution on GPUs
Bernabé López-Albelda, Francisco M. Castro, José María González-Linares, Nicolás Guil
J. Supercomput.4
2021 ReSGait: The Real-Scene Gait Dataset
abstract
Many studies have shown that gait recognition can be used to identify humans at a long distance, with promising results on current datasets. However, those datasets are collected under controlled situations and predefined conditions, which limits the extrapolation of the results to unconstrained situations in which the subjects walk freely in scenes. To cover this gap, we release a novel real-scene gait dataset (ReSGait), which is the first dataset collected in unconstrained scenarios with freely moving subjects and not controlled environmental parameters. Overall, our dataset is composed of 172 subjects and 870 video sequences, recorded over 15 months. Video sequences are labeled with gender, clothing, carrying conditions, taken walking route, and whether mobile phones were used or not. Therefore, the main characteristics of our dataset that differentiate it from other datasets are as follows: (i) uncontrolled real-life scenes and (ii) long recording time. Finally, we empirically assess the difficulty of the proposed dataset by evaluating state-of-the-art gait approaches for silhouette and pose modalities. The results reveal an accuracy of less than 35%, showing the inherent level of difficulty of our dataset compared to other current datasets, in which accuracies are higher than 90%. Thus, our proposed dataset establishes a new level of difficulty in the gait recognition problem, much closer to real life.
Zihao Mu, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Yan-Ran Li 0001, Shiqi Yu 0001
IJCB4
2021 Multimodal Gait Recognition Under Missing Modalities
abstract
Multimodal systems for gait recognition have gained a lot of attention. However, there is a clear gap in the study of missing modalities, which represents real-life scenarios where sensors fail or data get corrupted. Here, we investigate how to handle missing modalities for gait recognition. We propose a single and flexible framework that uses a variable number of input modalities. For each modality, it consists of a branch and a binary unit indicating whether the modality is available; these are gated and merged together. Finally, it generates a single and compact ‘multimodal’ gait signature that encodes biometric information of the input. Our framework outperforms the state of the art on TUM-GAID and extensive experiments reveal its effectiveness for handling missing modalities even in the multiview setup of CASIA-B. The code is available online: https://github.com/avagait/gaitmiss.
Rubén Delgado-Escaño, Francisco M. Castro, Nicolás Guil, Vicky Kalogeiton, Manuel J. Marín-Jiménez
ICIP3
2021 Anomalous object detection by active search with PTZ cameras
Ezequiel López-Rubio, Miguel A. Molina-Cabello, Francisco M. Castro, Rafael Marcos Luque Baena, Manuel J. Marín-Jiménez, Nicolás Guil
Expert Syst. Appl.6
2021 UGaitNet: Multimodal Gait Recognition With Missing Input Modalities
abstract
Gait recognition systems typically rely solely on silhouettes for extracting gait signatures. Nevertheless, these approaches struggle with changes in body shape and dynamic backgrounds; a problem that can be alleviated by learning from multiple modalities. However, in many real-life systems some modalities can be missing, and therefore most existing multimodal frameworks fail to cope with missing modalities. To tackle this problem, in this work, we propose UGaitNet, a unifying framework for gait recognition, robust to missing modalities. UGaitNet handles and mingles various types and combinations of input modalities, i.e. pixel gray value, optical flow, depth maps, and silhouettes, while being camera agnostic. We evaluate UGaitNet on two public datasets for gait recognition: CASIA-B and TUM-GAID, and show that it obtains compact and state-of-the-art gait descriptors when leveraging multiple or missing modalities. Finally, we show that UGaitNet with optical flow and grayscale inputs achieves almost perfect (98.9%) recognition accuracy on CASIA-B (same-view “normal”) and 100% on TUM-GAID (“ellapsed time”). Code will be available.
Manuel J. Marín-Jiménez, Francisco M. Castro, Rubén Delgado-Escaño, Vicky Kalogeiton, Nicolás Guil
IEEE Trans. Inf. Forensics Secur.5
2020 iLGaCo: Incremental Learning of Gait Covariate Factors
abstract
Gait is a popular biometric pattern used for identifying people based on their way of walking. Traditionally, gait recognition approaches based on deep learning are trained using the whole training dataset. In fact, if new data (classes, view-points, walking conditions, etc.) need to be included, it is necessary to re-train again the model with old and new data samples. In this paper, we propose iLGaCo, the first incremental learning approach of covariate factors for gait recognition, where the deep model can be updated with new information without re-training it from scratch by using the whole dataset. Instead, our approach performs a shorter training process with the new data and a small subset of previous samples. This way, our model learns new information while retaining previous knowledge. We evaluate iLGaCo on CASIA-B dataset in two incremental ways: adding new view-points and adding new walking conditions. In both cases, our results are close to the classical `training-from-scratch' approach, obtaining a marginal drop in accuracy ranging from 0.2% to 1.2%, what shows the efficacy of our approach. In addition, the comparison of iLGaCo with other incremental learning methods, such as LwF and iCarl, shows a significant improvement in accuracy, between 6% and 15% depending on the experiment.
Zihao Mu, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Yan-Ran Li 0001, Shiqi Yu 0001
IJCB4
2020 Heuristics for concurrent task scheduling on GPUs
abstract
Summary Concurrent execution of tasks in GPUs can reduce the computation time of a workload by overlapping data transfer and execution commands. However, it is difficult to implement an efficient runtime scheduler that minimizes the workload makespan as many execution orderings should be evaluated. In this paper, we employ scheduling theory to build a model that takes into account the device capabilities, workload characteristics, constraints, and objective functions. In our model, GPU tasks scheduling is reformulated as a flow shop scheduling problem, which allow us to apply and compare well‐known heuristics already developed in the operations research field. In addition, we develop a new heuristic, specifically focused on executing GPU commands, that achieves better scheduling results than previous ones. It leverages on a precise GPU command execution model for both computation and data transfers to carry out more advantageous scheduling decisions. A comprehensive evaluation, showing the suitability and robustness of this new approach, is conducted in three different NVIDIA architectures (Kepler, Maxwell, and Pascal). Results confirm the proposed heuristic achieves the best results in more than 90% of the experiments. Furthermore, a comparison has been made with MPS (Multi‐Process Service), the NVIDIA API that deals with the execution of concurrent tasks, which shows that our solution obtains speed‐ups ranging from 1.15 to 1.20.
Bernabé López-Albelda, A. J. Lázaro-Muñoz, José María González-Linares, Nicolás Guil
Concurr. Comput. Pract. Exp.4
2020 Multimodal feature fusion for CNN-based gait recognition: an empirical comparison
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Nicolas Pérez de la Blanca
Neural Comput. Appl.3
2019 Energy-based tuning of convolutional neural networks on multi-GPUs
abstract
Summary Deep Learning (DL) applications are gaining momentum in the realm of Artificial Intelligence, particularly after GPUs have demonstrated remarkable skills for accelerating their challenging computational requirements. Within this context, Convolutional Neural Network (CNN) models constitute a representative example of success on a wide set of complex applications, particularly on datasets where the target can be represented through a hierarchy of local features of increasing semantic complexity. In most of the real scenarios, the roadmap to improve results relies on CNN settings involving brute force computation, and researchers have lately proven Nvidia GPUs to be one of the best hardware counterparts for acceleration. Our work complements those findings with an energy study on critical parameters for the deployment of CNNs on flagship image and video applications, ie, object recognition and people identification by gait, respectively. We evaluate energy consumption on four different networks based on the two most popular ones (ResNet/AlexNet), ie, ResNet (167 layers), a 2D CNN (15 layers), a CaffeNet (25 layers), and a ResNetIm (94 layers) using batch sizes of 64, 128, and 256, and then correlate those with speed‐up and accuracy to determine optimal settings. Experimental results on a multi‐GPU server endowed with twin Maxwell and twin Pascal Titan X GPUs demonstrate that energy correlates with performance and that Pascal may have up to 40% gains versus Maxwell. Larger batch sizes extend performance gains and energy savings, but we have to keep an eye on accuracy, which sometimes shows a preference for small batches. We expect this work to provide a preliminary guidance for a wide set of CNN and DL applications in modern HPC times, where the GFLOPS/w ratio constitutes the primary goal.
Francisco M. Castro, Nicolás Guil, Manuel J. Marín-Jiménez, Jesús Pérez Serrano, Manuel Ujaldon
Concurr. Comput. Pract. Exp.2
2018 End-to-End Incremental Learning
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Cordelia Schmid, Karteek Alahari
ECCV (12)3
2018 Improving Bag-of-Visual-Words model using visual n-grams for human action classification
Ruber Hernández-García, Julián Ramos Cózar, Nicolás Guil, Edel B. García Reyes, Hichem Sahli
Expert Syst. Appl.3
2017 Deep multi-task learning for gait-based biometrics
abstract
The task of identifying people by the way they walk is known as `gait recognition'. Although gait is mainly used for identification, additional tasks as gender recognition or age estimation may be addressed based on gait as well. In such cases, traditional approaches consider those tasks as independent ones, defining separated task-specific features and models for them. This paper shows that by training jointly more than one gait-based tasks, the identification task converges faster than when it is trained independently, and the recognition performance of multi-task models is equal or superior to more complex single-task ones. Our model is a multi-task CNN that receives as input a fixed-length sequence of optical flow channels and outputs several biometric features (identity, gender and age).
Manuel J. Marín-Jiménez, Francisco M. Castro, Nicolás Guil, Fernando De la Torre, Rafael Medina Carnicer
ICIP3
2017 Fisher Motion Descriptor for Multiview Gait Recognition
abstract
The goal of this paper is to identify individuals by analyzing their gait. Instead of using binary silhouettes as input data (as done in many previous works) we propose and evaluate the use of motion descriptors based on densely sampled short-term trajectories. We take advantage of state-of-the-art people detectors to define custom spatial configurations of the descriptors around the target person, obtaining a rich representation of the gait motion. The local motion features (described by the Divergence-Curl-Shear descriptor [M. Jain, H. Jegou and P. Bouthemy, Better exploiting motion for better action recognition, in Proc. IEEE Conf. Computer Vision Pattern Recognition (CVPR) (2013), pp. 2555–2562.]) extracted on the different spatial areas of the person are combined into a single high-level gait descriptor by using the Fisher Vector encoding [F. Perronnin, J. Sánchez and T. Mensink, Improving the Fisher kernel for large-scale image classification, in Proc. European Conf. Computer Vision (ECCV) (2010), pp. 143–156]. The proposed approach, coined Pyramidal Fisher Motion, is experimentally validated on ‘CASIA’ dataset [S. Yu, D. Tan and T. Tan, A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition, in Proc. Int. Conf. Pattern Recognition, Vol. 4 (2006), pp. 441–444]. (parts B and C), ‘TUM GAID’ dataset, [M. Hofmann, J. Geiger, S. Bachmann, B. Schuller and G. Rigoll, The TUM Gait from Audio, Image and Depth (GAID) database: Multimodal recognition of subjects and traits, J. Vis. Commun. Image Represent. 25(1) (2014) 195–206]. ‘CMU MoBo’ dataset [R. Gross and J. Shi, The CMU Motion of Body (MoBo) database, Technical Report CMU-RI-TR-01-18, Robotics Institute (2001)]. and the recent ‘AVA Multiview Gait’ dataset [D. López-Fernández, F. Madrid-Cuevas, A. Carmona-Poyato, M. Marín-Jiménez and R. Muñoz-Salinas, The AVA multi-view dataset for gait recognition, in Activity Monitoring by Multiple Distributed Sensing, Lecture Notes in Computer Science (Springer, 2014), pp. 26–39]. The results show that this new approach achieves state-of-the-art results in the problem of gait recognition, allowing to recognize walking people from diverse viewpoints on single and multiple camera setups, wearing different clothes, carrying bags, walking at diverse speeds and not limited to straight walking paths.
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Rafael Muñoz-Salinas
Int. J. Pattern Recognit. Artif. Intell.3
2017 A tasks reordering model to reduce transfers overhead on GPUs
A. J. Lázaro-Muñoz, José María González-Linares, Juan Gómez-Luna, Nicolás Guil
J. Parallel Distributed Comput.4
2016 Multimodal features fusion for gait, gender and shoes recognition
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil
Mach. Vis. Appl.3
2016 Configurable XOR Hash Functions for Banked Scratchpad Memories in GPUs
abstract
Scratchpad memories in GPU architectures are employed as software-controlled caches to increase the effective GPU memory bandwidth. Through the use of well-known optimization techniques, such as privatization and tiling, they are properly exploited. Typically, they are banked memories which are addressed with a$\text{mod}(2^N)$bank indexing scheme. Although their bandwidth is fully exploited for linear memory accesses, their performance is burdened when non-unit strides appear in memory access patterns because they provoke bank conflicts. This paper explores the use of configurablebit-vectorandbitwiseXOR-based hash functions to evenly distribute memory addresses of the access patterns over the memory banks, reducing the number of bank conflicts. An exhaustive, but lightweight, search is used to configure bit-vector hash functions. Bitwise hash functions are configured with heuristics. Hardware and software implementations are carried out. For the hardware approach, the experimental results show 24 percent performance speed-up for 22 benchmarks on GPGPU-Sim, a Fermi-like simulator. Bank conflicts are reduced by 96 percent with bit-vector hash functions, and 97 percent with bitwise hash functions using our proposed Minimum Imbalance Heuristic. The software approach, using bit-vector hash functions, demonstrates 23 percent speed-up and 96 percent bank conflict reduction on a Fermi GPU, and 33 percent speed-up and 99 percent bank conflict reduction on a Kepler GPU.
Gert-Jan van den Braak, Juan Gómez-Luna, José María González-Linares, Henk Corporaal, Nicolás Guil
IEEE Trans. Computers5
2016 In-Place Matrix Transposition on GPUs
abstract
Matrix transposition is an important algorithmic building block for many numeric algorithms such as FFT. With more and more algebra libraries offloading to GPUs, a high performance in-place transposition becomes necessary. Intuitively, in-place transposition should be a good fit for GPU architectures due to limited available on-board memory capacity and high throughput. However, direct application of CPU in-place transposition algorithms lacks the amount of parallelism and locality required by GPU to achieve good performance. In this paper we present our in-place matrix transposition approach for GPUs that is performed using elementary tile-wise transpositions. We propose low-level optimizations for the elementary transpositions, and find the best performing configurations for them. Then, we compare all sequences of transpositions that achieve full transposition, and detect which is the most favorable for each matrix. We present an heuristic to guide the selection of tile sizes, and compare them to brute-force search. We diagnose the drawback of our approach, and propose a solution using minimal padding. With fast padding and unpadding kernels, the overall throughput is significantly increased. Finally, we compare our method to another recent implementation.
Juan Gómez-Luna, I-Jui Sung, Li-Wen Chang, José María González-Linares, Nicolás Guil, Wen-Mei W. Hwu
IEEE Trans. Parallel Distributed Syst.5
2015 Empirical Study of Audio-Visual Features Fusion for Gait Recognition
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil
CAIP (1)3
2015 In-Place Data Sliding Algorithms for Many-Core Architectures
abstract
In-place data manipulation is very desirable in many-core architectures with limited on-board memory. This paper deals with the in-place implementation of a class of primitives that perform data movements in one direction. We call these primitives Data Sliding (DS) algorithms. Notable among them are relational algebra primitives (such as select and unique), padding to insert empty elements in a data structure, and stream compaction to reduce memory requirements. Their in-place implementation in a bulk synchronous parallel model, such as GPUs, is specially challenging due to the difficulties in synchronizing threads executing on different compute units. Using a novel adjacent work-group synchronization technique, we propose two algorithmic schemes for regular and irregular DS algorithms. With a set of 5 benchmarks, we validate our approaches and compare them to the state-of-the-art implementations of these benchmarks. Our regular DS algorithms demonstrate up to 9.11x and 73.25x on NVIDIA and AMD GPUs, respectively, the throughput of their competitors. Our irregular DS algorithms outperform NVIDIA Thrust library by up to 3.24x on the three most recent generations of NVIDIA GPUs.
Juan Gómez-Luna, Li-Wen Chang, I-Jui Sung, Wen-Mei W. Hwu, Nicolás Guil
ICPP5
2015 Demystifying the 16 × 16 thread-block for stencils on the GPU
abstract
Summary Stencil computation is of paramount importance in many fields, in image processing, structural biology and biomedicine, among others. There exists a permanent demand of maximizing the performance of stencils on state‐of‐the‐art architectures, such graphics processing units (GPUs). One of the important issues when optimizing these kernels for the GPU is the selection of the best thread‐block that maximizes the overall performance. Usually, programmers look for the optimal thread‐block configuration in a reduced space of square thread‐block configurations or simply use the best configurations reported in previous works, which is usually 16 × 16. This paper provides a better understanding of the impact of thread‐block configurations on the performance of stencils on the GPU. In particular, we model locality and parallelism and consider that the optimal configurations are within the space that provides: (1) a small number of global memory communications; (2) a good shared memory utilization with small numbers of conflicts; (3) a good streaming multi‐processors utilization; and (4) a high efficiency of the threads within a thread‐block. The model determines the set of optimal thread‐block configurations without the need of executing the code. We validate the proposed model using six stencils with different halo widths and show that it reduces the optimization space to around 25% of the total valid space. The configurations in this space achieve at least a throughput of 75% of the best configuration and guarantee the inclusion of the best configurations. Copyright © 2015 John Wiley & Sons, Ltd.
Siham Tabik, Maurice Peemen, Nicolás Guil, Henk Corporaal
Concurr. Comput. Pract. Exp.3
2015 Calculation of dense trajectory descriptors on a heterogeneous embedded architecture
Julián Ramos Cózar, Manuel J. Marín-Jiménez, José María González-Linares, Nicolás Guil, Juan Gómez-Luna
J. Syst. Archit.4
2015 On how to improve tracklet-based gait recognition systems
Manuel J. Marín-Jiménez, Francisco M. Castro, Ángel Carmona-Poyato, Nicolás Guil
Pattern Recognit. Lett.4
2014 Human Action Classification Using N-Grams Visual Vocabulary
Ruber Hernández-García, Edel B. García Reyes, Julián Ramos Cózar, Nicolás Guil
CIARP4
2014 In-place transposition of rectangular matrices on accelerators
abstract
Matrix transposition is an important algorithmic building block for many numeric algorithms such as FFT. It has also been used to convert the storage layout of arrays. With more and more algebra libraries offloaded to GPUs, a high performance in-place transposition becomes necessary. Intuitively, in-place transposition should be a good fit for GPU architectures due to limited available on-board memory capacity and high throughput. However, direct application of CPU in-place transposition algorithms lacks the amount of parallelism and locality required by GPUs to achieve good performance. In this paper we present the first known in-place matrix transposition approach for the GPUs. Our implementation is based on a novel 3-stage transposition algorithm where each stage is performed using an elementary tiled-wise transposition. Additionally, when transposition is done as part of the memory transfer between GPU and host, our staged approach allows hiding transposition overhead by overlap with PCIe transfer. We show that the 3-stage algorithm allows larger tiles and achieves 3X speedup over a traditional 4-stage algorithm, with both algorithms based on our high-performance elementary transpositions on the GPU. We also show our proposed low-level optimizations improve the sustained throughput to more than 20 GB/s. Finally, we propose an asynchronous execution scheme that allows CPU threads to delegate in-place matrix transposition to GPU, achieving a throughput of more than 3.4 GB/s (including data transfers costs), and improving current multithreaded implementations of in-place transposition on CPU.
I-Jui Sung, Juan Gómez-Luna, José María González-Linares, Nicolás Guil, Wen-Mei W. Hwu
PPoPP4
2013 Simulation and architecture improvements of atomic operations on GPU scratchpad memory
abstract
GPUs are increasingly used as compute accelerators. With a large number of cores executing an even larger number of threads, significant speed-ups can be attained for parallel workloads. Applications that rely on atomic operations, such as histogram and Hough transform, suffer from serialization of threads in case they update the same memory location. Previous work shows that reducing this serialization with software techniques can increase performance by an order of magnitude. We observe, however, that some serialization remains and still slows down these applications. Therefore, this paper proposes to use a hash function in both the addressing of the banks and the locks of the scratchpad memory. To measure the effects of these changes, we first implement a detailed model of atomic operations on scratchpad memory in GPGPU-Sim, and verify its correctness. Second, we test our proposed hardware changes. They result in a speed-up up to 4.9× and 1.8× on implementations utilizing the aforementioned software techniques for histogram and Hough transform applications respectively, with minimum hardware costs.
Gert-Jan van den Braak, Juan Gómez-Luna, Henk Corporaal, José María González-Linares, Nicolás Guil
ICCD5
2013 An optimized approach to histogram computation on GPU
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil
Mach. Vis. Appl.4
2013 Performance Modeling of Atomic Additions on GPU Scratchpad Memory
abstract
GPU application implementations using scatter approaches will fall into write contention due to atomic updates of output elements, if these result from more than one input element. Colliding threads will be serialized, seriously harming performance. Dealing with these issues requires a proper understanding of the behavior of the scratchpad or shared memory under conflicting accesses caused by concurrent threads. Thus, this paper presents an exhaustive microbenchmark-based analysis of atomic additions in shared memory that quantifies the impact of access conflicts on latency and throughput. This analysis has led us to discover the lock mechanism that enables atomic updates to shared memory and to propose a performance model to estimate the latency penalties due to collisions by position or bank conflicts. Then, we have derived experiments from this model that show us the way to optimize applications using atomic operations. Position and bank conflicts can be diminished by replication and padding, respectively. The benefits of such techniques are illustrated with the optimization of two widely used voting processes: the centroid updating step in k-means clustering, and histogram calculation.
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil
IEEE Trans. Parallel Distributed Syst.4
2012 Reducing Vocabulary Size in Human Action Classification
abstract
Human action classification is an important task in computer vision. Bag-of-Words using spatio-temporal features and some classification algorithm is one of the most successful methods in this context. In this work we have studied the effect of reducing the vocabulary size using a video word ranking method. We have used the KTH dataset to obtain a vocabulary with more descriptive words and, at the same time, more compact and efficient. Results for different vocabulary sizes show an improvement of the recognition rate whilst reducing the number of words due to the fact that non-descriptive words are removed.
Julián Ramos Cózar, Ruber Hernández, Yanio Heredia, José María González-Linares, Nicolás Guil
KES5
2012 Pixel-based background initialization using spatio-temporal restrictions
abstract
A precise background detection is required for video surveillance applications in order to detect foreground objects. However, in the presence of cluttered scenes, standard techniques for background segmentation can fail. In this work we present a new technique for foreground detection that is able to detect the correct background in complex scenes. It works grouping neighbour pixels that fulfill some kind of spatial and temporal criteria. Spatial criterion is based on the appearance similarity of neighbour pixels while temporal criterion looks for the best temporal correlation in the whole video sequence. Initially, a set of seed points of the image are selected and both criteria are applied in an alternate way until all the pixels of the image have been visited and their background value has been calculated.
Juan Villalba Espinosa, José María González-Linares, Julián Ramos Cózar, Nicolás Guil
KES4
2012 Performance models for asynchronous data transfers on consumer Graphics Processing Units
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil
J. Parallel Distributed Comput.4
2011 Detection of logos in low quality videos
abstract
This paper presents a novel framework for logo detection in low quality videos. Our method assumes the logo template is unknown in advance and exploits the property that logotype pixels appearance through several consecutive frames has a lower variance than the others. Segmentation is difficult to accomplish if logo continuity is broken. In this work we propose the use of both edge and appearance continuity to carry out the segmentation. By checking edge continuity, the video is split into sequences with stable content. Later, sequences with similar static content are merged in order to build a longer sequence. Next a Gaussian mixture is used to model the variance of the pixels values in the merged sequences. Finally, a threshold that allows identification of the logo pixels is calculated. The new method is compared with a state-of-the-art method, obtaining better results in both accuracy and false logo rejection.
Julián Ramos Cózar, Pablo Nieto, José María González-Linares, Nicolás Guil, Yanio Hernández Heredia
ISDA4
2010 Learning a generic 3D face model from 2D image databases using incremental Structure-from-Motion
Jose Gonzalez-Mora, Fernando De la Torre, Nicolás Guil, Emilio L. Zapata
Image Vis. Comput.3
2009 Efficient image alignment using linear appearance models
abstract
Visual tracking is a key component in many computer vision applications. Linear subspace techniques (e.g. eigen-tracking) are one of the most popular approaches to align templates with appearance variations (e.g. illumination, iconic changes). A number of well known tracking algorithms have been proposed in the last years to accurately fit these models to images. Computational efficiency is an important limitation in object tracking algorithms and different efficient techniques, such as the “projected-out” optimization, have been proposed. They reduce the computational cost using an efficient formulation in which many of the involved operations can be precomputed. On the other hand, alternative “simultaneous” algorithms jointly optimize pose and appearance parameters, providing better performance but increasing the computational cost. In this paper, we propose an algorithm for efficient linear appearance model fitting based on the inverse compositional simultaneous optimization of pose and appearance. We introduce a novel formulation which reduces the required computational time while maintaining similar convergence properties of previous “simultaneous” approaches. Experimental results illustrate the capabilities of this algorithm in face tracking.
Jose Gonzalez-Mora, Nicolás Guil, Emilio L. Zapata, Fernando De la Torre
CVPR2
2009 Parallelization of a Video Segmentation Algorithm on CUDA-Enabled Graphics Processing Units
Juan Gómez-Luna, José María González-Linares, José Ignacio Benavides Benítez, Nicolás Guil
Euro-Par4
2008 Topic 11: Distributed and High-Performance Multimedia
Frank J. Seinstra, Nicolás Guil, Zoltan Juhasz
Euro-Par2
2008 A TV-logo classification and learning system
abstract
Logotypes superimposed to broadcasted videos supply important information for semantic video annotation, such as the content creator. In this work a novel logo classification and learning system for TV broadcast videos is presented. Logos are segmented from the video stream but scale change, position shift, clutter and noise makes difficult to classify and to recognize them. Several robust features that use edges and shape information have been selected, and a Bayesian network classifier is used to classify the logos. New logos are recognized as such for the first time they appear and passed to a semi-supervised learning system. The learning process clusters the set of new logos to group different instances of the same new logo. A logo model is obtained for each cluster that must be validated by a human to incorporate them into the classification system. Comprehensive tests with a set of 724 TV logos show the high performance of our classification and learning system.
Pablo Nieto, Julián Ramos Cózar, José María González-Linares, Nicolás Guil
ICIP4
2008 Recognition of circular patterns on GPUs: Performance analysis and contributions
Antonio Ruiz 0001, Nicolás Guil, Manuel Ujaldon
J. Parallel Distributed Comput.2
2008 On the computation of the Circle Hough Transform by a GPU rasterizer
Manuel Ujaldon, Antonio Ruiz 0001, Nicolás Guil
Pattern Recognit. Lett.3
2007 Bilinear Active Appearance Models
abstract
Appearance Models have been applied to model the space of human faces over the last two decades. In particular, Active Appearance Models (AAMs) have been successfully used for face tracking, synthesis and recognition, and they are one of the state-of-the-art approaches due to its efficiency and representational power. Although widely employed, AAMs suffer from a few drawbacks, such as the inability to isolate pose, identity and expression changes. This paper proposes Bilinear Active Appearance Models (BAAMs), an extension of AAMs, that effectively decouple changes due to pose and expression/identity. We derive a gradient-descent algorithm to efficiently fit BAAMs to new images. Experimental results show how BAAMs improve generalization and convergence with respect to the linear model. In addition, we illustrate decoupling benefits of BAAMs in face recognition across pose. We show how the pose normalization provided by BAAMs increase the recognition performance of commercial systems.
Jose Gonzalez-Mora, Fernando De la Torre, Rajesh Murthi, Nicolás Guil, Emilio L. Zapata
ICCV4
2007 Logotype detection to support semantic-based video annotation
Julián Ramos Cózar, Nicolás Guil, José María González-Linares, Emilio L. Zapata, Ebroul Izquierdo
Signal Process. Image Commun.2
2006 Tracking of Linear Appearance Models Using Second Order Minimization
Jose Gonzalez-Mora, Nicolás Guil, Emilio L. Zapata
ACIVS2
2006 Video Cataloging Based on Robust Logotype Detection
abstract
In this paper a technique for video cataloging based on logo detection is shown. No a priori knowledge about shape or spatial-temporal location of logos is assumed. The method implements a new algorithm for online logo detection based on temporal and spatial segmentation of broadcasted videos. Temporal segmentation identifies constant luminance regions within video frames while spatial segmentation helps to refine previous segmented regions. In a final step, identified logos are searched in a database and classified into candidate or learnt logotypes. Learnt logos can be directly tracked through the video. Candidate logotypes are assigned to a cluster of similar logos. After a promotion process, all the candidate logos belonging to the same cluster are used to create a new learnt logotype.
Julián Ramos Cózar, Nicolás Guil, José María González-Linares, Emilio L. Zapata
ICIP2
2005 Analysis and Description of the Semantic Content of Cell Biological Videos
Andrés Rodríguez Moreno, Nicolás Guil, David M. Shotton, Oswaldo Trelles
Multim. Tools Appl.2
2004 Combining luminance and edge based metrics for robust temporal video segmentation
abstract
This work presents a new video indexing technique able to detect shot cuts as well as gradual transitions in MPEG compressed video sequences. Moreover, this technique also allows to perform global motion estimation in order to successfully detect swing, zoom and displacement effects in video streams. In fact, motion information is used during the gradual shot transition detection algorithm in order to reduce false positives caused by video motion. Two different metrics are used. The first one is based on luminance histograms, while the second is based on the Generalized Hough transform. The use of this technique greatly improves the results obtained by other shot cut detection algorithms. Also, it is a key technique to the success of both; the gradual shot transition detection and the global motion estimation algorithms.
Edmundo Saez, José Ignacio Benavides Benítez, Nicolás Guil
ICIP3
2004 Reliable real time scene change detection in MPEG compressed video
abstract
This work presents a new scene change detection method in MPEG compressed video. This method is based on two different video characteristics, edges and luminance, and takes the best of them to define two distance functions. On one hand, a contour based distance function that requires no registration technique is introduced. On the other hand, a new distance function based on luminance histograms correlation is defined. In order to efficiently combine both distance functions, a new estimator based on the divergence of distributions is proposed. The method has been tested using several videos from the MPEG-7 content set.
Edmundo Saez, José Ignacio Benavides Benítez, Nicolás Guil
ICME3
2004 Automatic analysis of the content of cell biological videos and database organization of their metadata descriptors
abstract
We present a video content analysis and metadata organizational system for research videos arising from biological microscopy of living cells. Automated procedures are described to determine the position, size, shape and orientation of cells in each video frame. From the temporal changes in the values of these simple metadata parameters, high-level descriptors are derived that describe the semantic content of the video. This content information (specific intrinsic metadata) is of high information value, since it describes the behavior of cells and the timing of events within the video, including changes in environmental conditions experienced by the cells. When such metadata are properly organized in a searchable database, a content-based video query and retrieval system may be developed to locate particular objects, events or behaviors. Moreover, the availability of such semantic contents in the formal and generic format we propose will allow the application of data mining techniques and the amassing of more elaborate knowledge, e.g., species classification depending on behavior, patterns in response to environment changes, etc. The suitability and functionality of the proposed metadata model is demonstrated by the automated analysis of five different types of biological experiments, recording epithelial wound healing, bacterial multiplication, the rotations of tethered bacteria, and the swimming of motile bacteria and of human sperm.
Andrés Rodríguez Moreno, Nicolás Guil, David M. Shotton, Oswaldo Trelles
IEEE Trans. Multim.2
2003 Global motion estimation algorithm for video segmentation
Edmundo Saez, José M. Palomares, José Ignacio Benavides Benítez, Nicolás Guil
VCIP4
2003 An efficient 2D deformable objects detection and location algorithm
José María González-Linares, Nicolás Guil, Emilio L. Zapata
Pattern Recognit.2
2002 Planar object detection under scaled orthographic projection
Julián Ramos Cózar, Nicolás Guil, Emilio L. Zapata
Pattern Recognit. Lett.2
2001 A deformable model for image segmentation in noisy medical images
abstract
Deformable-model-based segmentation techniques can overcome some limitations of the traditional image processing techniques. Currently developed deformable models can cope with gaps and another irregularities in object boundaries. However, they present problems in noisy images. Our approach is able to segment objects in noisy images by defining a new energy function associated with image noise and avoiding the tendency of contour points to bunch up. The model is validated for vessel segmentation on mammograms.
Francisco L. Valverde, Nicolás Guil, José Muñoz, Takashi Aoyama, Kunio Doi
ICIP (3)2
2001 An evaluation criterion for edge detection techniques in noisy images
abstract
Segmentation in noisy images is an important and difficult problem in pattern recognition. Edge detection is a crucial step in this process. Current subjective and objective methods for evaluation and comparison of segmentation techniques are inadequate or not applicable to edge detection techniques. A general framework for segmentation evaluation in noisy images is introduced after a brief review of previous work. Several measures based on similarity between true and result segmented images are defined. These measures are, then, combined in a unique criterion as a proposed global measure of performance. The results indicate that this global measure can be helpful in the evaluation and comparison of segmentation techniques applied to noisy images.
Francisco L. Valverde, Nicolás Guil, José Muñoz, Robert M. Nishikawa, Kunio Doi
ICIP (1)2
2001 Detection of arbitrary planar shapes with 3D pose
Julián Ramos Cózar, Nicolás Guil, Emilio L. Zapata
Image Vis. Comput.2
2000 Deformable Shapes Detection by Stochastic Optimization
abstract
A new approach to the detection of shapes under global deformations is presented. The algorithm is based in the combination of a generalized Hough transform (GHT) and an universal evolutionary global optimizer (UEGO). This method exploits the invariant characteristics to rotation, scale and displacement of the GHT to detect shapes deformed by a global deformation model, and without an initial positioning of the template. The GHT is used as an objective function for the UEGO, an optimizer that is able to find multiple optima with a low computational cost.
José María González-Linares, Nicolás Guil, Emilio L. Zapata, Pilar Martínez Ortigosa, Inmaculada García
ICIP2
2000 Object Tracking and Event Recognition in Biological Microscopy Videos
abstract
We present a video analysis and content-based video query and retrieval system for research videos arising from biological microscopy that: (a) uses image processing procedures to identify objects and events automatically within the digitized videos, thus leading to the generation of content information (specific intrinsic metadata) of high information value by automated analysis of their visual content; and (b) permits subsequent queries to be performed on the specific intrinsic metadata thus generated, allowing important factual and analytical information to be obtained and selected video sequences matching the query criteria to be retrieved. Such a system requires the novel use of image and video processing techniques, and new approaches to the organization, accessing and querying of video metadata databases.
David M. Shotton, Andrés Rodríguez Moreno, Nicolás Guil, Oswaldo Trelles
ICPR3
1999 Bidimensional shape detection using an invariant approach
Nicolás Guil, José María González-Linares, Emilio L. Zapata
Pattern Recognit.1
1997 Fast Hough Transform on Multiprocessors: A Branch and Bound Approach
Nicolás Guil, Emilio L. Zapata
J. Parallel Distributed Comput.1
1997 Lower order circle and ellipse Hough transform
Nicolás Guil, Emilio L. Zapata
Pattern Recognit.1
1996 Parallelization of irregular algorithms for shape detection
abstract
There are a series of very efficient sequential algorithms that generate irregular trees during the process of detecting shapes in images. These algorithms are based on the fast Hough transform and are used for solving the most complex stages of detection when the production of the parameters is uncoupled. However, the parallelization of these algorithms is complex, and the problem of load distribution is crucial. We present three parallel algorithms for solving this problem. One of the solutions employs static load balancing. The other two use dynamic balancing with two different control policies: distributed and centralized. These algorithms may also be used for solving other problems, such as the branch and bound, that generate irregular trees.
Nicolás Guil, Emilio L. Zapata
ICIP (2)1
1995 A fast Hough transform for segment detection
abstract
The authors describe a new algorithm for the fast Hough transform (FHT) that satisfactorily solves the problems other fast algorithms propose in the literature-erroneous solutions, point redundance, scaling, and detection of straight lines of different sizes-and needs less storage space. By using the information generated by the algorithm for the detection of straight lines, they manage to detect the segments of the image without appreciable computational overhead. They also discuss the performance and the parallelization of the algorithm and show its efficiency with some examples.
Nicolás Guil, Julio Villalba, Emilio L. Zapata
IEEE Trans. Image Process.1