VLDB 2026 Research / reviewers in the wild / expert
Manuel J. Marín-Jiménez
dblp:44/4520 · also Manuel Jesús Marín-Jiménez
· DBLP profile ↗
60ranked-venue papers
14as first author
22since 2021 · last 2025
0000-0001-9294-6714ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 8 since 2021Security and privacy · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What Does Gait Reveal About Health? Investigating Human Motion as an Indicator
Rafael Aguilar-Ortega, Shiqi Yu 0001, Nuria Marín-Jiménez, Manuel J. Marín-Jiménez |
CAIP (2) | 4 |
| 2025 | Leveraging Implicit 3D Geometry for Biometric and Anthropometric Estimation from GaitabstractEstimating biometric and anthropometric attributes from gait sequences presents a promising alternative to traditional body measurement techniques, particularly in unconstrained or low-resource environments. However, inferring metric attributes from 2D silhouette-based gait representations remains challenging due to the lack of volumetric cues. In this work, we propose a novel training strategy that leverages the geometric priors encoded by PIFuHD—a high-resolution implicit function-based model for 3D human reconstruction—to inject structural supervision into a gait encoder. Our method introduces a feature reconstruction branch that distills volumetric knowledge from precomputed PIFuHD embeddings during training, enabling accurate prediction of anthropometric attributes at inference time from 2D silhouette sequences alone. We evaluate our approach on the Health&Gait dataset, achieving substantial improvements over baselines in predicting multiple attributes, including height, weight, BMI, and body circumferences. These results demonstrate that shape-aligned supervision from implicit 3D models can effectively bridge the gap between geometric reasoning and efficient biometric estimation from visual gait data. Nicolás Cubero, Jorge Zafra-Palma, Francisco M. Castro, Nicolás Guil, Manuel J. Marín-Jiménez |
IJCB | 5 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 35 |
| 2025 | Edge artificial intelligence and super-resolution for enhanced weapon detection in video surveillanceabstractThe prevalence of crimes involving handguns and knives underscores the importance of early weapon detection. This, along with the spread of video surveillance systems, boosted the development of automatic approaches for weapon detection from surveillance cameras. Despite the advancements from classical computer vision to Deep Learning (DL) techniques, accurately detecting weapons in real-time remains challenging due to their small size. Current DL methods, which attempt to mitigate this issue using complex detection architectures, are resource-intensive, resulting in high costs and energy usage, and hindering their deployment on efficient edge devices. This creates challenges in resource-limited environments, making these methods impractical for edge and real-time applications. To address these shortcomings, our work proposes YOLOSR, which integrates a You Only Look Once (YOLO) v8-small model with an Enhanced Deep Super Resolution (EDSR)-based network using a shared backbone. During training, the auxiliary Super Resolution (SR) helps in learning better features, which could benefit the weapon detection task. During inference, the SR branch is removed, keeping the detector’s computational complexity unchanged. The YOLOSR’s accuracy and efficiency were validated on our WeaponSense dataset and on a NVIDIA Jetson Nano, against other weapon detectors. The results exhibited that YOLOSR, compared to the state-of-the-art YOLOv8-small model, maintained the same computational complexity with 28.8 billion floating point operations and on-device latency of 101 ms per image, while increasing the Average Precision by 10.2 percentage points. Thus, the YOLOSR emerges as an effective solution for real-time weapon detection in resource-constrained environments, achieving an optimal trade-off between efficiency and accuracy. • YOLOSR merges Super-Resolution (SR) with YOLOv8 to enhance small weapon detection. • SR boosts training and is discarded at inference, with no extra computational cost. • YOLOSR achieved a 10.2% AP gain compared to YOLOv8, keeping the same real-time speed. • The approach runs at 101ms/image on Jetson Nano, enabling real-time edge inference. • The approach was validated on WeaponSense, a realistic video-surveillance dataset. Daniele Berardini, Lucia Migliorelli, Alessandro Galdelli, Manuel J. Marín-Jiménez |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Empirical study of human pose representations for gait recognitionabstractGait recognition has gained attention for its ability to identify individuals from afar. Current state-of-the-art approaches predominantly utilize visual information, such as silhouettes, or a combination of visual data and basic body pose information, including skeleton joint coordinates. However, the role of human pose in gait recognition is still underexplored, often leading to poorer results compared to visual approaches. In this work, we propose a novel hierarchical limb-based representation that enhances the depiction of body pose and can be applied to various pose descriptors. Our representation consists of three hierarchical levels: full body, body limbs (arms and legs), and middle limbs (forearms, lower arms, thighs, and shins). This structure enriches the gait description of the overall pose by incorporating the specific movements of each limb. Particularly, we investigate the application of our hierarchical arrangement using two different rich pose descriptors: heatmaps derived from 2D body skeletons and a dense representation obtained from pixel-wise estimation of body pose ( i.e DensePose). Furthermore, we introduce the PoseGaitGL family of models to better leverage the features derived from our pose representations. By employing our hierarchical pose representations, the proposed model achieves state-of-the-art results in pose-based gait recognition. Thus, the hierarchical heatmap-based and hierarchical DensePose representations attain Rank-1 accuracy of 82.2% and 92.0%, respectively, on the cross-view setup of CASIA-B, and 99.3% and 99.8%, respectively, on TUM-GAID, establishing a new benchmark for pose-based methods. Source code is available at https://github.com/Nico-Cubero/PoseGaitGL . Nicolás Cubero, Francisco M. Castro, Julián Ramos Cózar, Nicolás Guil, Manuel J. Marín-Jiménez |
Expert Syst. Appl. | 5 |
| 2025 | Lightweight Structure-Aware Attention for Visual Understanding
Heeseung Kwon, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Karteek Alahari |
Int. J. Comput. Vis. | 3 |
| 2024 | Cross-Modality Gait Recognition: Bridging LiDAR and Camera Modalities for Human IdentificationabstractCurrent gait recognition research mainly focuses on identifying pedestrians captured by the same type of sensor, neglecting the fact that individuals may be captured by different sensors in order to adapt to various environments. A more practical approach should involve cross-modality matching across different sensors. Hence, this paper focuses on investigating the problem of cross-modality gait recognition, with the objective of accurately identifying pedestrians across diverse vision sensors. We present CrossGait inspired by the feature alignment strategy, capable of cross retrieving diverse data modalities. Specifically, we investigate the cross-modality recognition task by initially extracting features within each modality and subsequently aligning these features across modalities. To further enhance the cross-modality performance, we propose a Prototypical Modality-shared Attention Module that learns modality-shared features from two modality-specific features. Additionally, we design a Cross-modality Feature Adapter that transforms the learned modality-specific features into a unified feature space. Extensive experiments conducted on the SUSTech1K dataset demonstrate the effectiveness of CrossGait: (1) it exhibits promising cross-modality ability in retrieving pedestrians across various modalities from different sensors in diverse scenes, and (2) CrossGait not only learns modality-shared features for cross-modality gait recognition but also maintains modality-specific features for single-modality recognition. Chuanfu Shen, Manuel J. Marín-Jiménez, George Q. Huang, Shiqi Yu 0001 |
IJCB | 3 |
| 2024 | ReSLAM: Reusable SLAM with heterogeneous cameras
Francisco José Romero-Ramírez, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez, Ángel Carmona-Poyato, Rafael Medina Carnicer |
Neurocomputing | 3 |
| 2024 | DeepArUco++: Improved detection of square fiducial markers in challenging lighting conditions
Rafael Berral-Soler, Rafael Muñoz-Salinas, Rafael Medina Carnicer, Manuel J. Marín-Jiménez |
Image Vis. Comput. | 4 |
| 2024 | Guest Editorial: Special issue on IbPRIA 2023
Antonio Javier Gallego 0001, Manuel J. Marín-Jiménez, Raquel Justo, Hélder Oliveira, Antonio Pertusa |
Pattern Anal. Appl. | 2 |
| 2024 | Proxemics-net++: classification of human interactions in still imagesabstractAbstract Human interaction recognition (HIR) is a significant challenge in computer vision that focuses on identifying human interactions in images and videos. HIR presents a great complexity due to factors such as pose diversity, varying scene conditions, or the presence of multiple individuals. Recent research has explored different approaches to address it, with an increasing emphasis on human pose estimation. In this work, we propose Proxemics-Net++, an extension of the Proxemics-Net model, capable of addressing the problem of recognizing human interactions in images through two different tasks: the identification of the types of “touch codes” or proxemics and the identification of the type of social relationship between pairs. To achieve this, we use RGB and body pose information together with the state-of-the-art deep learning architecture, ConvNeXt, as the backbone. We performed an ablative analysis to understand how the combination of RGB and body pose information affects these two tasks. Experimental results show that body pose information contributes significantly to proxemic recognition (first task) as it allows to improve the existing state of the art, while its contribution in the classification of social relations (second task) is limited due to the ambiguity of labelling in this problem, resulting in RGB information being more influential in this task. Isabel Jiménez-Velasco, Jorge Zafra-Palma, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez |
Pattern Anal. Appl. | 4 |
| 2024 | AttenGait: Gait recognition with attention and rich modalities
Francisco M. Castro, Rubén Delgado-Escaño, Ruber Hernández-García, Manuel J. Marín-Jiménez, Nicolás Guil |
Pattern Recognit. | 4 |
| 2024 | Editorial: Special session on IbPRIA 2023
Antonio Javier Gallego 0001, Manuel J. Marín-Jiménez, Raquel Justo, Hélder Oliveira, Antonio Pertusa |
Pattern Recognit. Lett. | 2 |
| 2023 | A Comparison of Neural Network-Based Super-Resolution Models on 3D Rendered Images
Rafael Berral-Soler, Francisco José Madrid-Cuevas, Sebastián Ventura, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez |
CAIP (1) | 5 |
| 2023 | Empirical Study of Attention-Based Models for Automatic Classification of Gastrointestinal Endoscopy Images
Ricardo Espantaleón-Pérez, Isabel Jiménez-Velasco, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez |
CAIP (2) | 4 |
| 2022 | Preface to the Special Issue on Human Pose, Motion, Activities and Shape in 3D
Manuel J. Marín-Jiménez, Javier Romero 0002, Hao Li 0015, Grégory Rogez |
Int. J. Comput. Vis. | 1 |
| 2022 | LAEO-Net++: Revisiting People Looking at Each Other in VideosabstractCapturing the 'mutual gaze' of people is essential for understanding and interpreting the social interactions between them. To this end, this paper addresses the problem of detecting people Looking At Each Other (LAEO) in video sequences. For this purpose, we propose LAEO-Net++, a new deep CNN for determining LAEO in videos. In contrast to previous works, LAEO-Net++ takes spatio-temporal tracks as input and reasons about the whole track. It consists of three branches, one for each character's tracked head and one for their relative position. Moreover, we introduce two new LAEO datasets: UCO-LAEO and AVA-LAEO. A thorough experimental evaluation demonstrates the ability of LAEO-Net++ to successfully determine if two people are LAEO and the temporal window where it happens. Our model achieves state-of-the-art results on the existing TVHID-LAEO video dataset, significantly outperforming previous approaches. Finally, we apply LAEO-Net++ to a social network, where we automatically infer the social relationship between pairs of people based on the frequency and duration that they LAEO, and show that LAEO can be a useful tool for guided search of human interactions in videos. Manuel J. Marín-Jiménez, Vicky Kalogeiton, Pablo Medina-Suarez, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | ReSGait: The Real-Scene Gait DatasetabstractMany studies have shown that gait recognition can be used to identify humans at a long distance, with promising results on current datasets. However, those datasets are collected under controlled situations and predefined conditions, which limits the extrapolation of the results to unconstrained situations in which the subjects walk freely in scenes. To cover this gap, we release a novel real-scene gait dataset (ReSGait), which is the first dataset collected in unconstrained scenarios with freely moving subjects and not controlled environmental parameters. Overall, our dataset is composed of 172 subjects and 870 video sequences, recorded over 15 months. Video sequences are labeled with gender, clothing, carrying conditions, taken walking route, and whether mobile phones were used or not. Therefore, the main characteristics of our dataset that differentiate it from other datasets are as follows: (i) uncontrolled real-life scenes and (ii) long recording time. Finally, we empirically assess the difficulty of the proposed dataset by evaluating state-of-the-art gait approaches for silhouette and pose modalities. The results reveal an accuracy of less than 35%, showing the inherent level of difficulty of our dataset compared to other current datasets, in which accuracies are higher than 90%. Thus, our proposed dataset establishes a new level of difficulty in the gait recognition problem, much closer to real life. Zihao Mu, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Yan-Ran Li 0001, Shiqi Yu 0001 |
IJCB | 3 |
| 2021 | Multimodal Gait Recognition Under Missing ModalitiesabstractMultimodal systems for gait recognition have gained a lot of attention. However, there is a clear gap in the study of missing modalities, which represents real-life scenarios where sensors fail or data get corrupted. Here, we investigate how to handle missing modalities for gait recognition. We propose a single and flexible framework that uses a variable number of input modalities. For each modality, it consists of a branch and a binary unit indicating whether the modality is available; these are gated and merged together. Finally, it generates a single and compact ‘multimodal’ gait signature that encodes biometric information of the input. Our framework outperforms the state of the art on TUM-GAID and extensive experiments reveal its effectiveness for handling missing modalities even in the multiview setup of CASIA-B. The code is available online: https://github.com/avagait/gaitmiss. Rubén Delgado-Escaño, Francisco M. Castro, Nicolás Guil, Vicky Kalogeiton, Manuel J. Marín-Jiménez |
ICIP | 5 |
| 2021 | Anomalous object detection by active search with PTZ cameras
Ezequiel López-Rubio, Miguel A. Molina-Cabello, Francisco M. Castro, Rafael Marcos Luque Baena, Manuel J. Marín-Jiménez, Nicolás Guil |
Expert Syst. Appl. | 5 |
| 2021 | RealHePoNet: a robust single-stage ConvNet for head pose estimation in the wild
Rafael Berral-Soler, Francisco José Madrid-Cuevas, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez |
Neural Comput. Appl. | 4 |
| 2021 | UGaitNet: Multimodal Gait Recognition With Missing Input ModalitiesabstractGait recognition systems typically rely solely on silhouettes for extracting gait signatures. Nevertheless, these approaches struggle with changes in body shape and dynamic backgrounds; a problem that can be alleviated by learning from multiple modalities. However, in many real-life systems some modalities can be missing, and therefore most existing multimodal frameworks fail to cope with missing modalities. To tackle this problem, in this work, we propose UGaitNet, a unifying framework for gait recognition, robust to missing modalities. UGaitNet handles and mingles various types and combinations of input modalities, i.e. pixel gray value, optical flow, depth maps, and silhouettes, while being camera agnostic. We evaluate UGaitNet on two public datasets for gait recognition: CASIA-B and TUM-GAID, and show that it obtains compact and state-of-the-art gait descriptors when leveraging multiple or missing modalities. Finally, we show that UGaitNet with optical flow and grayscale inputs achieves almost perfect (98.9%) recognition accuracy on CASIA-B (same-view “normal”) and 100% on TUM-GAID (“ellapsed time”). Code will be available. Manuel J. Marín-Jiménez, Francisco M. Castro, Rubén Delgado-Escaño, Vicky Kalogeiton, Nicolás Guil |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | iLGaCo: Incremental Learning of Gait Covariate FactorsabstractGait is a popular biometric pattern used for identifying people based on their way of walking. Traditionally, gait recognition approaches based on deep learning are trained using the whole training dataset. In fact, if new data (classes, view-points, walking conditions, etc.) need to be included, it is necessary to re-train again the model with old and new data samples. In this paper, we propose iLGaCo, the first incremental learning approach of covariate factors for gait recognition, where the deep model can be updated with new information without re-training it from scratch by using the whole dataset. Instead, our approach performs a shorter training process with the new data and a small subset of previous samples. This way, our model learns new information while retaining previous knowledge. We evaluate iLGaCo on CASIA-B dataset in two incremental ways: adding new view-points and adding new walking conditions. In both cases, our results are close to the classical `training-from-scratch' approach, obtaining a marginal drop in accuracy ranging from 0.2% to 1.2%, what shows the efficacy of our approach. In addition, the comparison of iLGaCo with other incremental learning methods, such as LwF and iCarl, shows a significant improvement in accuracy, between 6% and 15% depending on the experiment. Zihao Mu, Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Yan-Ran Li 0001, Shiqi Yu 0001 |
IJCB | 3 |
| 2020 | Multimodal feature fusion for CNN-based gait recognition: an empirical comparison
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Nicolas Pérez de la Blanca |
Neural Comput. Appl. | 2 |
| 2020 | Editorial for special section at Pattern Recognition Letters - IbPRIA 2019
Manuel J. Marín-Jiménez, Aythami Morales, Julian Fierrez, Antonio Pertusa, Hugo Proença 0001, J. Salvador Sánchez 0001 |
Pattern Recognit. Lett. | 1 |
| 2019 | LAEO-Net: Revisiting People Looking at Each Other in VideosabstractCapturing the ‘mutual gaze’ of people is essential for understanding and interpreting the social interactions between them. To this end, this paper addresses the problem of detecting people Looking At Each Other (LAEO) in video sequences. For this purpose, we propose LAEO-Net, a new deep CNN for determining LAEO in videos. In contrast to previous works, LAEO-Net takes spatio-temporal tracks as input and reasons about the whole track. It consists of three branches, one for each character’s tracked head and one for their relative position. Moreover, we introduce two new LAEO datasets: UCO-LAEO and AVA-LAEO. A thorough experimental evaluation demonstrates the ability of LAEO-Net to successfully determine if two people are LAEO and the temporal window where it happens. Our model achieves state-of-the-art results on the existing TVHID-LAEO video dataset, significantly outperforming previous approaches. Manuel J. Marín-Jiménez, Vicky Kalogeiton, Pablo Medina-Suarez, Andrew Zisserman |
CVPR | 1 |
| 2019 | Energy-based tuning of convolutional neural networks on multi-GPUsabstractSummary Deep Learning (DL) applications are gaining momentum in the realm of Artificial Intelligence, particularly after GPUs have demonstrated remarkable skills for accelerating their challenging computational requirements. Within this context, Convolutional Neural Network (CNN) models constitute a representative example of success on a wide set of complex applications, particularly on datasets where the target can be represented through a hierarchy of local features of increasing semantic complexity. In most of the real scenarios, the roadmap to improve results relies on CNN settings involving brute force computation, and researchers have lately proven Nvidia GPUs to be one of the best hardware counterparts for acceleration. Our work complements those findings with an energy study on critical parameters for the deployment of CNNs on flagship image and video applications, ie, object recognition and people identification by gait, respectively. We evaluate energy consumption on four different networks based on the two most popular ones (ResNet/AlexNet), ie, ResNet (167 layers), a 2D CNN (15 layers), a CaffeNet (25 layers), and a ResNetIm (94 layers) using batch sizes of 64, 128, and 256, and then correlate those with speed‐up and accuracy to determine optimal settings. Experimental results on a multi‐GPU server endowed with twin Maxwell and twin Pascal Titan X GPUs demonstrate that energy correlates with performance and that Pascal may have up to 40% gains versus Maxwell. Larger batch sizes extend performance gains and energy savings, but we have to keep an eye on accuracy, which sometimes shows a preference for small batches. We expect this work to provide a preliminary guidance for a wide set of CNN and DL applications in modern HPC times, where the GFLOPS/w ratio constitutes the primary goal. Francisco M. Castro, Nicolás Guil, Manuel J. Marín-Jiménez, Jesús Pérez Serrano, Manuel Ujaldon |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | SPM-SLAM: Simultaneous localization and mapping with squared planar markers
Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez, Rafael Medina Carnicer |
Pattern Recognit. | 2 |
| 2018 | End-to-End Incremental Learning
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Cordelia Schmid, Karteek Alahari |
ECCV (12) | 2 |
| 2018 | Robust identification of fiducial markers in challenging conditions
Víctor Manuel Mondéjar-Guerra, Sergio Garrido-Jurado, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez, Rafael Medina Carnicer |
Expert Syst. Appl. | 4 |
| 2018 | 3D human pose estimation from depth maps using a deep combination of poses
Manuel J. Marín-Jiménez, Francisco José Romero-Ramírez, Rafael Muñoz-Salinas, Rafael Medina Carnicer |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Mapping and localization from planar markersabstractSquared planar markers are a popular tool for fast, accurate and robust camera localization, but its use is frequently limited to a single marker, or at most, to a small set of them for which their relative pose is known beforehand. Mapping and localization from a large set of planar markers is yet a scarcely treated problem in favour of keypoint-based approaches. However, while keypoint detectors are not robust to rapid motion, large changes in viewpoint, or significant changes in appearance, fiducial markers can be robustly detected under a wider range of conditions. This paper proposes a novel method to simultaneously solve the problems of mapping and localization from a set of squared planar markers. First, a quiver of pairwise relative marker poses is created, from which an initial pose graph is obtained. The pose graph may contain small pairwise pose errors, that when propagated, leads to large errors. Thus, we distribute the rotational and translational error along the basis cycles of the graph so as to obtain a corrected pose graph. Finally, we perform a global pose optimization by minimizing the reprojection errors of the planar markers in all observed frames. The experiments conducted show that our method performs better than Structure from Motion and visual SLAM techniques. Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez, Enrique Yeguas-Bolívar, Rafael Medina Carnicer |
Pattern Recognit. | 2 |
| 2017 | Deep multi-task learning for gait-based biometricsabstractThe task of identifying people by the way they walk is known as `gait recognition'. Although gait is mainly used for identification, additional tasks as gender recognition or age estimation may be addressed based on gait as well. In such cases, traditional approaches consider those tasks as independent ones, defining separated task-specific features and models for them. This paper shows that by training jointly more than one gait-based tasks, the identification task converges faster than when it is trained independently, and the recognition performance of multi-task models is equal or superior to more complex single-task ones. Our model is a multi-task CNN that receives as input a fixed-length sequence of optical flow channels and outputs several biometric features (identity, gender and age). Manuel J. Marín-Jiménez, Francisco M. Castro, Nicolás Guil, Fernando De la Torre, Rafael Medina Carnicer |
ICIP | 1 |
| 2017 | Mixing body-parts model for 2D human pose estimation in stereo videosabstractThis study targets 2D articulated human pose estimation (i.e. localisation of body limbs) in stereo videos. Although in recent years depth‐based devices (e.g. Microsoft Kinect) have gained popularity, as they perform very well in controlled indoor environments (e.g. living rooms, operating theatres or gyms), they suffer clear problems in outdoor scenarios and, therefore, human pose estimation is still an interesting unsolved problem. The authors propose here a novel approach that is able to localise upper‐body keypoints (i.e. shoulders, elbows, and wrists) in temporal sequences of stereo image pairs. The authors’ method starts by locating and segmenting people in the image pairs by using disparity and appearance information. Then, a set of candidate body poses is computed for each view independently. Finally, temporal and stereo consistency is applied to estimate a final 2D pose. The authors’ validate their model on three challenging datasets: ‘stereo human pose estimation dataset’, ‘poses in the wild’ and ‘INRIA 3DMovie’. The experimental results show that the authors’ model not only establishes new state‐of‐the‐art results on stereo sequences, but also brings improvements in monocular sequences. Manuel I. López-Quintero, Manuel J. Marín-Jiménez, Rafael Muñoz-Salinas, Rafael Medina Carnicer |
IET Comput. Vis. | 2 |
| 2017 | Fisher Motion Descriptor for Multiview Gait RecognitionabstractThe goal of this paper is to identify individuals by analyzing their gait. Instead of using binary silhouettes as input data (as done in many previous works) we propose and evaluate the use of motion descriptors based on densely sampled short-term trajectories. We take advantage of state-of-the-art people detectors to define custom spatial configurations of the descriptors around the target person, obtaining a rich representation of the gait motion. The local motion features (described by the Divergence-Curl-Shear descriptor [M. Jain, H. Jegou and P. Bouthemy, Better exploiting motion for better action recognition, in Proc. IEEE Conf. Computer Vision Pattern Recognition (CVPR) (2013), pp. 2555–2562.]) extracted on the different spatial areas of the person are combined into a single high-level gait descriptor by using the Fisher Vector encoding [F. Perronnin, J. Sánchez and T. Mensink, Improving the Fisher kernel for large-scale image classification, in Proc. European Conf. Computer Vision (ECCV) (2010), pp. 143–156]. The proposed approach, coined Pyramidal Fisher Motion, is experimentally validated on ‘CASIA’ dataset [S. Yu, D. Tan and T. Tan, A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition, in Proc. Int. Conf. Pattern Recognition, Vol. 4 (2006), pp. 441–444]. (parts B and C), ‘TUM GAID’ dataset, [M. Hofmann, J. Geiger, S. Bachmann, B. Schuller and G. Rigoll, The TUM Gait from Audio, Image and Depth (GAID) database: Multimodal recognition of subjects and traits, J. Vis. Commun. Image Represent. 25(1) (2014) 195–206]. ‘CMU MoBo’ dataset [R. Gross and J. Shi, The CMU Motion of Body (MoBo) database, Technical Report CMU-RI-TR-01-18, Robotics Institute (2001)]. and the recent ‘AVA Multiview Gait’ dataset [D. López-Fernández, F. Madrid-Cuevas, A. Carmona-Poyato, M. Marín-Jiménez and R. Muñoz-Salinas, The AVA multi-view dataset for gait recognition, in Activity Monitoring by Multiple Distributed Sensing, Lecture Notes in Computer Science (Springer, 2014), pp. 26–39]. The results show that this new approach achieves state-of-the-art results in the problem of gait recognition, allowing to recognize walking people from diverse viewpoints on single and multiple camera setups, wearing different clothes, carrying bags, walking at diverse speeds and not limited to straight walking paths. Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Rafael Muñoz-Salinas |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2017 | New method for obtaining optimal polygonal approximations to solve the min- $$\varepsilon$$ ε problem
Ángel Carmona-Poyato, Eusebio J. Aguilera-Aguilera, Francisco José Madrid-Cuevas, Manuel J. Marín-Jiménez, Nicolás Luis Fernández García |
Neural Comput. Appl. | 4 |
| 2016 | Fast computation of optimal polygonal approximations of digital planar closed curves
Eusebio J. Aguilera-Aguilera, Ángel Carmona-Poyato, Francisco José Madrid-Cuevas, Manuel J. Marín-Jiménez |
Graph. Model. | 4 |
| 2016 | Viewpoint-independent gait recognition through morphological descriptions of 3D human reconstructions
David López-Fernández, Francisco José Madrid-Cuevas, Ángel Carmona-Poyato, Manuel J. Marín-Jiménez, Rafael Muñoz-Salinas, Rafael Medina Carnicer |
Image Vis. Comput. | 4 |
| 2016 | Simultaneous reconstruction and calibration for multi-view structured light scanning
Sergio Garrido-Jurado, Rafael Muñoz-Salinas, Francisco José Madrid-Cuevas, Manuel J. Marín-Jiménez |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Multimodal features fusion for gait, gender and shoes recognition
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil |
Mach. Vis. Appl. | 2 |
| 2016 | Stereo Pictorial Structure for 2D articulated human pose estimation
Manuel I. López-Quintero, Manuel J. Marín-Jiménez, Rafael Muñoz-Salinas, Francisco José Madrid-Cuevas, Rafael Medina Carnicer |
Mach. Vis. Appl. | 2 |
| 2015 | Empirical Study of Audio-Visual Features Fusion for Gait Recognition
Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil |
CAIP (1) | 2 |
| 2015 | Keypoint descriptor fusion with Dempster-Shafer theory
Víctor Manuel Mondéjar-Guerra, Rafael Muñoz-Salinas, Manuel J. Marín-Jiménez, Ángel Carmona-Poyato, Rafael Medina Carnicer |
Int. J. Approx. Reason. | 3 |
| 2015 | Calculation of dense trajectory descriptors on a heterogeneous embedded architecture
Julián Ramos Cózar, Manuel J. Marín-Jiménez, José María González-Linares, Nicolás Guil, Juan Gómez-Luna |
J. Syst. Archit. | 2 |
| 2015 | On how to improve tracklet-based gait recognition systems
Manuel J. Marín-Jiménez, Francisco M. Castro, Ángel Carmona-Poyato, Nicolás Guil |
Pattern Recognit. Lett. | 1 |
| 2014 | Pyramidal Fisher Motion for Multiview Gait RecognitionabstractThe goal of this paper is to identify individuals by analyzing their gait. Instead of using binary silhouettes as input data (as done in many previous works) we propose and evaluate the use of motion descriptors based on densely sampled short-term trajectories. We take advantage of state-of-the-art people detectors to define custom spatial configurations of the descriptors around the target person. Thus, obtaining a pyramidal representation of the gait motion. The local motion features (described by the Divergence-Curl-Shear descriptor [1]) extracted on the different spatial areas of the person are combined into a single high-level gait descriptor by using the Fisher Vector encoding [2]. The proposed approach, coined Pyramidal Fisher Motion, is experimentally validated on the recent 'AVA Multiview Gait' dataset [3]. The results show that this new approach achieves promising results in the problem of gait recognition. Francisco M. Castro, Manuel J. Marín-Jiménez, Rafael Medina Carnicer |
ICPR | 2 |
| 2014 | Conflict-based pruning of a solution space within a constructive geometric constraint solver
Enrique Yeguas-Bolívar, Manuel J. Marín-Jiménez, Rafael Muñoz-Salinas, Rafael Medina Carnicer |
Appl. Intell. | 2 |
| 2014 | Detecting People Looking at Each Other in Videos
Manuel J. Marín-Jiménez, Andrew Zisserman, Marcin Eichner, Vittorio Ferrari |
Int. J. Comput. Vis. | 1 |
| 2014 | Human interaction categorization by using audio-visual cues
Manuel J. Marín-Jiménez, Rafael Muñoz-Salinas, Enrique Yeguas-Bolívar, Nicolas Pérez de la Blanca |
Mach. Vis. Appl. | 1 |
| 2014 | Human action recognition from simple feature pooling
Manuel J. Marín-Jiménez, Nicolas Pérez de la Blanca, Maria Ángeles Mendoza |
Pattern Anal. Appl. | 1 |
| 2014 | Automatic generation and detection of highly reliable fiducial markers under occlusion
Sergio Garrido-Jurado, Rafael Muñoz-Salinas, Francisco José Madrid-Cuevas, Manuel J. Marín-Jiménez |
Pattern Recognit. | 4 |
| 2013 | Exploring STIP-based models for recognizing human interactions in TV videos
Manuel J. Marín-Jiménez, Enrique Yeguas-Bolívar, Nicolas Pérez de la Blanca |
Pattern Recognit. Lett. | 1 |
| 2012 | 2D Articulated Human Pose Estimation and Retrieval in (Almost) Unconstrained Still Images
Marcin Eichner, Manuel J. Marín-Jiménez, Andrew Zisserman, Vittorio Ferrari |
Int. J. Comput. Vis. | 2 |
| 2011 | "Here's looking at you, kid". Detecting people looking at each other in videosabstractThe objective of this work is to determine if people are interacting in TV video by detecting whether they are looking at each other or not. We determine both the temporal period of the interaction and also spatially localize the relevant people. We make the following three contributions: (i) head pose estimation in unconstrained scenarios (TV video) using Gaussian Process regression; (ii) propose and evaluate several methods for assessing whether and when pairs of people are looking at each other in a video shot; and (iii) introduce new ground truth annotation for this task, extending the TV Human Interactions Dataset [22]. The peformance of the methods is evaluated on this dataset, which consists of 300 video clips extracted from TV shows. despite the variety and difficulty of this video material, our best method obtains an average precision of 86: 2%. Manuel J. Marín-Jiménez, Andrew Zisserman, Vittorio Ferrari |
BMVC | 1 |
| 2011 | Exploring Alternative Spatial and Temporal Dense Representations for Action Recognition
Pau Agustí, V. Javier Traver, Manuel J. Marín-Jiménez, Filiberto Pla |
CAIP (2) | 3 |
| 2010 | RBM-based Silhouette Encoding for Human Action ModellingabstractIn this paper we evaluate the use of Restricted Bolzmann Machines (RBM) in the context of learning and recognizing human actions. The features used as basis are binary silhouettes of persons. We test the proposed approach on two datasets of human actions where binary silhouettes are available: ViHASi (synthetic data) and Weizmann (real data). In addition, on Weizmann dataset, we combine features based on optical flow with the associated binary silhouettes. The results show that thanks to the use of RBM-based models, very informative and shorter feature vectors can be obtained for the classification tasks, improving the classification performance. Manuel J. Marín-Jiménez, Nicolas Pérez de la Blanca, Maria Ángeles Mendoza |
ICPR | 1 |
| 2010 | Tracking people in video sequences using multiple models
Manuel J. Lucena, José Manuel Fuertes, Nicolas Pérez de la Blanca, Manuel J. Marín-Jiménez |
Multim. Tools Appl. | 4 |
| 2009 | Fitting Product of HMM to Human Motions
Maria Ángeles Mendoza, Nicolas Pérez de la Blanca, Manuel J. Marín-Jiménez |
CAIP | 3 |
| 2009 | Pose search: Retrieving people using their poseabstractWe describe a method for retrieving shots containing a particular 2D human pose from unconstrained movie and TV videos. The method involves first localizing the spatial layout of the head, torso and limbs in individual frames using pictorial structures, and associating these through a shot by tracking. A feature vector describing the pose is then constructed from the pictorial structure. Shots can be retrieved either by querying on a single frame with the desired pose, or through a pose classifier trained from a set of pose examples. Our main contribution is an effective system for retrieving people based on their pose, and in particular we propose and investigate several pose descriptors which are person, clothing, background and lighting independent. As a second contribution, we improve the performance over existing methods for localizing upper body layout on unconstrained video. We compare the spatial layout pose retrieval to a baseline method where poses are retrieved using a HOG descriptor. Performance is assessed on five episodes of the TV series 'Buffy the Vampire Slayer', and pose retrieval is demonstrated also on three Hollywood movies.. Vittorio Ferrari, Manuel J. Marín-Jiménez, Andrew Zisserman |
CVPR | 2 |
| 2008 | Progressive search space reduction for human pose estimationabstractThe objective of this paper is to estimate 2D human pose as a spatial configuration of body parts in TV and movie video shots. Such video material is uncontrolled and extremely challenging. We propose an approach that progressively reduces the search space for body parts, to greatly improve the chances that pose estimation will succeed. This involves two contributions: (i) a generic detector using a weak model of pose to substantially reduce the full pose search space; and (ii) employing 'grabcut' initialized on detected regions proposed by the weak model, to further prune the search space. Moreover, we also propose (Hi) an integrated spatio- temporal model covering multiple frames to refine pose estimates from individual frames, with inference using belief propagation. The method is fully automatic and self-initializing, and explains the spatio-temporal volume covered by a person moving in a shot, by soft-labeling every pixel as belonging to a particular body part or to the background. We demonstrate upper-body pose estimation by an extensive evaluation over 70000 frames from four episodes of the TV series Buffy the vampire slayer, and present an application to full- body action recognition on the Weizmann dataset. Vittorio Ferrari, Manuel J. Marín-Jiménez, Andrew Zisserman |
CVPR | 2 |