Álvaro García-Martín

dblp:39/1542 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-1705-3972ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 A Data-Centric Approach to Pedestrian Attribute Recognition: Synthetic Augmentation via Prompt-Driven Diffusion Models
abstract
Pedestrian Attribute Recognition (PAR) is a challenging task as models are required to generalize across numerous attributes in real-world data. Traditional approaches focus on complex methods, yet recognition performance is often constrained by training dataset limitations, particularly the under-representation of certain attributes. In this paper, we propose a data-centric approach to improve PAR by synthetic data augmentation guided by textual descriptions. First, we define a protocol to identify weakly recognized attributes across multiple datasets. Second, we propose a prompt-driven pipeline that leverages diffusion models to generate synthetic pedestrian images while preserving the consistency of PAR datasets. Finally, we derive a strategy to seamlessly incorporate synthetic samples into training data, which considers prompt-based annotation rules and modifies the loss function. Results on popular PAR datasets demonstrate that our approach not only boosts recognition of underrepresented attributes but also improves overall model performance beyond the targeted attributes. Notably, this approach strengthens zero-shot generalization without requiring architectural changes of the model, presenting an efficient and scalable solution to improve the recognition of attributes of pedestrians in the real world.
Sawaiz A. Chaudhry, Juan C. SanMiguel, Álvaro García-Martín, Pablo Ayuso-Albizu, Pablo Carballeira
AVSS4
2025 An image-processing toolkit for remote photoplethysmography
abstract
Abstract Objective. Image-processing-based remote photoplethysmography algorithms are usually composed of steps where different methods are used, and often, researchers perform these steps using methods that are not necessarily the best for their application. With our toolkit, we want to provide easy and fast access to different state-of-the-art methods for the most common image-processing steps in remote photoplethysmography algorithms. Methods. Our toolkit was programmed in Python and was developed with sequential workflow in mind, making it versatile and easy to use in interactive environments. It also includes tools so the users can modify or extend it if they want to, and will be updated as new methods for the different steps are published. Results. Our use case examples and validation show an effective approach and how the toolkit can be used for exhaustive evaluation and ablation studies in a simple way. We also show how choosing different methods can affect the final heart rate estimation accuracy at the cost of computation time. Conclusion. With this toolkit we are providing researchers with a versatile, easy-to-use tool, with access to different methods for the most common steps in remote photoplethysmography algorithms. Significance. Our toolkit is a relevant tool for researchers in the remote photoplethysmography field due to their versatility, ease of use, and adaptability. (It will be available on https://github.com/Montyro/rppgtk github upon acceptance).
Javier Montalvo, Álvaro García-Martín, José María Martínez Sanchez
Multim. Tools Appl.2
2024 Synthmanticlidar: A Synthetic Dataset For Semantic Segmentation On Lidar Imaging
abstract
Semantic segmentation on LiDAR imaging is increasingly gaining attention, as it can provide useful knowledge for perception systems and potential for autonomous driving. However, collecting and labeling real LiDAR data is an expensive and time-consuming task. While datasets such as SemanticKITTI [1] have been manually collected and labeled, the introduction of simulation tools such as CARLA [2], has enabled the creation of synthetic datasets on demand. In this work, we present a modified CARLA simulator designed with LiDAR semantic segmentation in mind, with new classes, more consistent object labeling with their counter-parts from real datasets such as SemanticKITTI, and the possibility to adjust the object class distribution. Using this tool, we have generated SynthmanticLiDAR, a synthetic dataset for semantic segmentation on LiDAR imaging, designed to be similar to SemanticKITTI, and we evaluate its contribution to the training process of different semantic segmentation algorithms by using a naive transfer learning approach. Our results show that incorporating SynthmanticLiDAR into the training process improves the overall performance of tested algorithms, proving the usefulness of our dataset, and therefore, our adapted CARLA simulator. The dataset and simulator are available in https:// github.com/vpulab/SynthmanticLiDAR.
Javier Montalvo, Pablo Carballeira, Álvaro García-Martín
ICIP3
2024 Long-Term Geo-Positioned Re-Identification Dataset of Urban Elements
abstract
This paper introduces UrbAM-ReID, a new long-term geo-positioned urban ReID dataset. It is composed by four subdatasets recording the same trajectory at the UAM Campus, each one recorded in different seasons and including an inverse direction recording. While most of the current datasets in the state-of-the-art focus on person re-identification, with vehicles as the second most explored object, our work specifically addresses urban objects re-identification, currently, waste containers, rubbish bins, and crosswalks. The dataset provides different attributes of the annotated objects, like their classes, their foreground or background status and the geo-position. Several evaluation configurations can be defined to simulate realistic scenarios that may arise in actual situations within the management of urban elements, considering the utilization of just visual data, or incorporating additional attributes, providing different complexity levels. Finally, the dataset is used for defining a benchmark where two state-of-the-art systems are evaluated. The dataset and supplementary material is available in https://github.com/vpulab/UrbAMReID
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez
ICIP2
2024 Improved transferability of self-supervised learning models through batch normalization finetuning
Kirill Sirotkin, Marcos Escudero-Viñolo, Pablo Carballeira, Álvaro García-Martín
Appl. Intell.4
2023 Self-supervised Monocular Depth Estimation on Unseen Synthetic Cameras
Cecilia Diana-Albelda, Juan Ignacio Bravo Pérez-Villar, Javier Montalvo, Álvaro García-Martín, Jesús Bescós
CIARP4
2023 Vehicle Re-Identification Based on Unsupervised Domain Adaptation by Incremental Generation of Pseudo-Labels
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez
CIARP2
2023 On exploring weakly supervised domain adaptation strategies for semantic segmentation using synthetic data
abstract
Abstract Pixel-wise image segmentation is key for many Computer Vision applications. The training of deep neural networks for this task has expensive pixel-level annotation requirements, thus, motivating a growing interest on synthetic data to provide unlimited data and its annotations. In this paper, we focus on the generation and application of synthetic data as representative training corpuses for semantic segmentation of urban scenes. First, we propose a synthetic data generation protocol, which identifies key features affecting performance and provides datasets with variable complexity. Second, we adapt two popular weakly supervised domain adaptation approaches (combined training, fine-tuning) to employ synthetic and real data. Moreover, we analyze several backbone models, real/synthetic datasets and their proportions when combined. Third, we propose a new curriculum learning strategy to employ several synthetic and real datasets. Our major findings suggest the high performance impact of pace and order of synthetic and real data presentation, achieving state of the art results for well-known models. The results by training with the proposed dataset outperform popular alternatives, thus demonstrating the effectiveness of the proposed protocol. Our code and dataset are available at http://www-vpu.eps.uam.es/publications/WSDA_semantic/
Roberto Alcover-Couso, Juan C. SanMiguel, Marcos Escudero-Viñolo, Álvaro García-Martín
Multim. Tools Appl.4
2023 Exploiting semantic segmentation to boost reinforcement learning in video game environments
abstract
Abstract In this work we explore enhancing performance of reinforcement learning algorithms in video game environments by feeding it better, more relevant data. For this purpose, we use semantic segmentation to transform the images that would be used as input for the reinforcement learning algorithm from their original domain to a simplified semantic domain with just silhouettes and class labels instead of textures and colors, and then we train the reinforcement learning algorithm with these simplified images. We have conducted different experiments to study multiple aspects: feasibility of our proposal, and potential benefits to model generalization and transfer learning. Experiments have been performed with the Super Mario Bros video game as the testing environment. Our results show multiple advantages for this method. First, it proves that using semantic segmentation enables reaching higher performance than the baseline reinforcement learning algorithm without modifying the actual algorithm, and in fewer episodes; second, it shows noticeable performance improvements when training on multiple levels at the same time; and finally, it allows to apply transfer learning for models trained on visually different environments. We conclude that using semantic segmentation can certainly help reinforcement learning algorithms that work with visual data, by refining it. Our results also suggest that other computer vision techniques may also be beneficial for data prepossessing. Models and code will be available on github upon acceptance.
Javier Montalvo, Álvaro García-Martín, Jesús Bescós
Multim. Tools Appl.2
2023 Enhancing vehicle re-identification via synthetic training datasets and re-ranking based on video-clips information
abstract
Abstract Vehicle re-identification (ReID) aims to find a specific vehicle identity across multiple non-overlapping cameras. The main challenge of this task is the large intra-class and small inter-class variability of vehicles appearance, sometimes related with large viewpoint variations, illumination changes or different camera resolutions. To tackle these problems, we proposed a vehicle ReID system based on ensembling deep learning features and adding different post-processing techniques. In this paper, we improve that proposal by: incorporating large-scale synthetic datasets in the training step; performing an exhaustive ablation study showing and analyzing the influence of synthetic content in ReID datasets, in particular CityFlow-ReID and VeRi-776; and extending post-processing by including different approaches to the use of gallery video-clips of the target vehicles in the re-ranking step. Additionally, we present an evaluation framework in order to evaluate CityFlow-ReID: as this dataset has not public ground truth annotations, AI City Challenge provided an on-line evaluation service which is no more available; our evaluation framework allows researchers to keep on evaluating the performance of their systems in the CityFlow-ReID dataset.
Paula Moral, Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós
Multim. Tools Appl.2
2020 Semantic-aware scene recognition
Alejandro López-Cifuentes, Marcos Escudero-Viñolo, Jesús Bescós, Álvaro García-Martín
Pattern Recognit.4
2019 Incorporating wheelchair users in people detection
Rafael Martin Nieto, Álvaro García-Martín, José María Martínez Sanchez
Multim. Tools Appl.2
2019 Automatic Vacant Parking Places Management System Using Multicamera Vehicle Detection
abstract
This paper presents a multicamera system for vehicles detection and their corresponding mapping into the parking spots of a parking lot. Approaches from the state-of-the-art system, which work properly in controlled scenarios, have been validated using small amount of sequences and without more challenging realistic conditions (illumination changes and different weather conditions). On the other hand, most of them are not complete systems, but provide only parts of them, usually detectors. The proposed system has been designed for realistic scenarios considering different cases of occlusion, illumination changes, and different climatic conditions; a real scenario (the International Pittsburgh Airport parking lot) has been targeted with the condition that existing parking security cameras can be used, avoiding the deployment of new cameras or other sensors infrastructures. For design and validation, a new multicamera data set has been recorded. The system is based on existing object detectors (the results of two of them are shown) and different proposed postprocessing stages. The results clearly show that the proposed system works correctly in challenging scenarios including almost total occlusions, illumination changes, and different weather conditions.
Rafael Martin Nieto, Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez
IEEE Trans. Intell. Transp. Syst.2
2017 Adaptive people detection based on cross-correlation maximization
abstract
Applying people detectors to unseen data is challenging since patterns distributions may significantly differ from the ones of the training dataset. In this paper, we propose a framework to adapt people detectors during runtime classification. Such adaptation takes advantage of multiple detectors to identify their best configurations (i.e. detection thresholds) without requiring manually labeled ground truth. We maximize the mutual information of detectors by pair-wise correlating their outputs to obtain a set of hypotheses for the detection thresholds. These hypotheses are later combined by weighted voting to obtain a final decision for the detection threshold of each detector. The proposed approach does not require re-training detectors and uses standard people detector outputs, i.e., bounding boxes, therefore it can employ various types of detectors. The experimental results demonstrate that the proposed approach outperforms state-of-the-art detectors whose optimal configuration is learned from training data.
Álvaro García-Martín, Juan C. SanMiguel
ICIP1
2015 Post-processing approaches for improving people detection performance
Álvaro García-Martín, José María Martínez Sanchez
Comput. Vis. Image Underst.1
2015 People detection in surveillance: classification and evaluation
abstract
Nowadays, people detection in video surveillance environments is a task that has been generating great interest. There are many approaches trying to solve the problem either in controlled scenarios or in very specific surveillance applications. The main objective of this study is to give a comprehensive and extensive evaluation of the state of the art of people detection regardless of the final surveillance application. For this reason, first, the different processing tasks involved in the automatic people detection in video sequences have been defined, then a proper classification of the state of the art of people detection has been made according to the two most critical tasks, object detection and person model, that are needed in every detection approach. Finally, experiments have been performed on an extensive dataset with different approaches that completely cover the proposed classification and support the conclusions drawn from the state of the art.
Álvaro García-Martín, José María Martínez Sanchez
IET Comput. Vis.1
2012 People-background segmentation with unequal error cost
abstract
We address the problem of segmenting a video in two classes of different semantic value, namely background and people, with the goal of guaranteeing that no people (or body parts) are classified as background. Body parts classified as background are given a higher classification error cost (segmentation with bias on background), as opposed to traditional approaches focused on people detection. To generate the people-background segmentation mask, the proposed approach first combines detection confidence maps of body parts and then extends them in order to derive a background mask, which is finally post-processed using morphological operators. Experiments validate the performance of our algorithm in different complex indoor and outdoor scenes with both static and moving cameras.
Álvaro García-Martín, Andrea Cavallaro, José María Martínez Sanchez
ICIP1
2012 On collaborative people detection and tracking in complex scenarios
Álvaro García-Martín, José María Martínez Sanchez
Image Vis. Comput.1
2012 A corpus for benchmarking of people detection algorithms
Álvaro García-Martín, José María Martínez Sanchez, Jesús Bescós
Pattern Recognit. Lett.1
2011 People detection based on appearance and motion models
abstract
The main contribution of this paper is a new people detection algorithm based on motion information. The algorithm builds a people motion model based on the Implicit Shape Model (ISM) Framework and the MoSIFT descriptor. We also propose a detection system that integrates appearance, motion and tracking information. Experimental results over sequences extracted from the TRECVID dataset show that our new people motion detector produces results comparable to the state of the art and that the proposed multimodal fusion system improves the obtained results combining the three information sources.
Álvaro García-Martín, Alex Hauptmann 0001, José María Martínez Sanchez
AVSS1
2010 Robust Real Time Moving People Detection in Surveillance Scenarios
abstract
In this paper an improved real time algorithm for detecting pedestrians in surveillance video is proposed. The algorithm is based on people appearance and defines a person model as the union of four models of body parts. Firstly, motion segmentation is performed to detect moving pixels. Then, moving regions are extracted and tracked. Finally, the detected moving objects are classified as human or nonhuman objects. In order to test and validate the algorithm, we have developed a dataset containing annotated surveillance sequences of different complexity levels focused on the pedestrians detection. Experimental results over this dataset show that our approach performs considerably well at real time and even better than other real and non-real time approaches from the state of art.
Álvaro García-Martín, José María Martínez Sanchez
AVSS1
2009 An Ontology for Event Detection and its Application in Surveillance Video
abstract
In this paper, we propose an ontology for representing the prior knowledge related to video event analysis. It is composed of two types of knowledge related to the application domain and the analysis system. Domain knowledge involves all the high level semantic concepts in the context of each examined domain (objects, events, context...) whilst system knowledge involves the capabilities of the analysis system (algorithms, reactions to events...). The proposed ontology has been structured in two parts: the basic ontology (composed of the basic concepts and their specializations) and the domain-specific extensions. Additionally, a video analysis framework based on the proposed ontology is defined for the analysis of different application domains showing the potential use of the proposed ontology. In order to show the real applicability of the proposed ontology, it is specialized for the underground video-surveillance domain showing some results that demonstrate the usability and effectiveness of the proposed ontology.
Juan C. SanMiguel, José María Martínez Sanchez, Álvaro García-Martín
AVSS3
2008 Video Object Segmentation Based on Feedback Schemes Guided by a Low-Level Scene Ontology
Álvaro García-Martín, Jesús Bescós
ACIVS1