Manuel Gil-Martín

dblp:219/9397 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-4285-6224ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Lightweight Model for Accurate Multi-View Hand Pose Recognition
abstract
This paper introduces a lightweight architecture for multi-view hand pose recognition on multimodal fusion of images and landmarks. The proposed model employs a compact Convolutional Neural Network (CNN) to extract visual features from dual-view grayscale images, while a Multi-Layer Perceptron (MLP) processes the corresponding Leap Motion Controller 2 hand landmarks. The two modalities are fused to create an efficient yet discriminative representation. Compared to the Vision Transformer (ViT)+MLP baseline, which achieves an F1 score of 79.33 ± 0.09 % with 8.95 × 107 parameters, our CNN+MLP model reaches a higher recognition accuracy of 85.36 ± 0.08 % while requiring only 2.13 × 105 parameters, which corresponds to an important reduction of the model size. Moreover, a landmarks-only variant using the MLP achieves 85.22 ± 0.08 % accuracy with just 6.46 × 104 parameters. These results, obtained on the Multi-view Leap2 Hand Pose Dataset under a Leave-One-Subject-Out Cross-Validation protocol, demonstrate that accurate multi-view hand pose recognition can be achieved with dramatically fewer parameters, enabling efficient deployment in resource-constrained environments.
Artur Kadyrzhanov, Sergio Esteban Romero, Manuel Gil-Martín, Marco Raoul Marini
ICAART (3)3
2026 Optimizing Sequential Models through Temporal Landmark Selection and Normalization for Sign Language Recognition
Sergio Esteban Romero, Iván Martín-Fernández, Cristina Luna Jiménez, Manuel Gil-Martín, Fernando Fernández Martínez, Elisabeth André
ICAART (3)4
2026 A Case Study on Large Visual-Language Model Attention Explainability After Adaptation Using Persuasion Strategies in Advertisements
Iván Martín-Fernández, Mihai Gabriel Constantin, Bogdan Ionescu, Sergio Esteban Romero, Fernando Fernández Martínez, Manuel Gil-Martín
MMM (1)6
2026 A comprehensive study on contrastive pre-training and fine tuning of vision and text transformers for video memorability prediction
abstract
Abstract Video memorability prediction has emerged as a key challenge for improving information retrieval, content design, and user engagement. Prior work has shown that semantic cues play a crucial role in determining memorability, with recent studies leveraging Contrastive Language-Image Pre-training (CLIP) encoders to incorporate semantic information. However, the specific improvements attributable to CLIP models remain unclear, as few studies systematically compare their performance against equivalent unimodal encoders or explore fine-tuning strategies. This work addresses that gap through a comprehensive, controlled evaluation of CLIP-based and unimodal encoders for video memorability prediction. We propose FCLIP, a domain-adapted extension of CLIP that undergoes additional contrastive pre-training on memorability-specific image-text pairs. Our experiments assess both feature extraction and supervised fine-tuning, ensuring fair comparisons across models with matched architecture and parameter count. Results show that FCLIP image encoders achieve a Spearman Rank Correlation Coefficient (SRCC) of 0.672 on the Memento10k dataset, significantly outperforming unimodal Vision Transformers. FCLIP text encoders similarly outperform unimodal baselines, reaching an SRCC of 0.632. These findings demonstrate that contrastive learning and domain adaptation substantially improve memorability prediction, highlighting the importance of semantic and multimodal pre-training in developing advanced content analysis systems.
Iván Martín-Fernández, Sergio Esteban Romero, Manuel Gil-Martín, Fernando Fernández Martínez
Multim. Tools Appl.3
2026 Principled Evaluation of Multi-Label Persuasion in Advertisements with Large Vision-Language Models
abstract
The automatic detection of persuasive strategies in advertisements presents a uniquely multimodal challenge at the intersection of vision, language, and social cognition. While recent advances in Large Vision-Language Models (LVLMs) offer promising capabilities for such tasks, current approaches often rely on restrictive evaluation schemes that do not reflect the inherently multi-label nature of persuasive messaging. In this work, we reframe persuasion strategy detection as a genuine multi-label classification problem and propose a principled evaluation framework to enhance interpretability and robustness. We apply this approach to both image and video datasets, examine their characteristics, and introduce novel input-agnostic baselines that achieve macro F1 scores of 0.082 and 0.289 on the respective test sets. As part of this analysis, we study label co-occurrence patterns and dataset ambiguities, providing insights that inform both model interpretation and future dataset design. To assess the native capabilities of LVLMs, we benchmark three open source models—PaliGemma, PaliGemma2, and Qwen2.5-VL—on the image persuasion dataset. Under zero-shot conditions, we demonstrate that querying each strategy individually with a logit-based decision threshold outperforms guided text generation. The best-performing zero-shot model, Qwen2.5-VL, achieves a macro F1 score of 0.227 and a sample F1 score of 0.234 on the image test set, and 0.381 and 0.400 respectively on the video test set. We further explore and compare two lightweight fine-tuning strategies that update only small subsets of model parameters while keeping the remaining weights frozen: fine-tuning of the image-to-text tokens linear projection and Low Rank Adaptation (LoRA) of the language model. Linear projector fine-tuning yields a top macro F1 score of 0.396 on the image test set, marking a substantial improvement over zero-shot performance. To evaluate cross-modal generalization, we apply fine-tuned image models to the video dataset. Our experiments reveal that, while projection-based fine-tuning enables partial knowledge transfer from image to video (macro F1 = 0.416 on a testing subset), LoRA adaptation severely disrupts cross-modal performance (macro F1 = 0.189). Finally, we perform a per-strategy performance analysis, looking into annotator- and data-centric factors that may influence LVLM performance. These findings highlight the viability of open-weight LVLMs for fine-grained persuasion analysis and suggest efficient pathways for domain-specific adaptation under realistic resource constraints.
Iván Martín-Fernández, Mihai Gabriel Constantin, Bogdan Ionescu, Manuel Gil-Martín, Fernando Fernández Martínez
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Exploring the Effect of Size, Architecture and Fine-Tuning Hyperparameters on Large Visual-Language Model Adaptation for Video Memorability Prediction
abstract
Predicting video memorability involves modeling audiovisual content to estimate its likelihood of being remembered-relevant in areas like recommendation, marketing, and education. Leveraging the semantic power of Large Vision-Language Models (LVLMs), we adapt Qwen2.5-VL models using Low-Rank Adaptation (LoRA) to this task. We evaluate the influence of model size (3B vs. 7B) and LoRA hyperparameters ($r$, α) on performance. Our best result, an SRCC of 0.7658 on the Memento10k development dataset using 5-Fold Cross Validation, surpasses previous approaches and confirms the importance of careful tuning. Notably, the model with the best zero-shot performance (3B) also achieved the highest fine-tuned score, showing that zero-shot evaluation can predict adaptation success. Results challenge assumptions about scaling in LVLMs and highlight the potential of smaller, well-adapted models.
David Luna-García, Iván Martín-Fernández, Sergio Esteban Romero, Manuel Gil-Martín, Fernando Fernández Martínez
CBMI4
2025 Hand Gesture Recognition Using MediaPipe Landmarks and Deep Learning Networks
abstract
Advanced Human Computer Interaction techniques are commonly used in multiple application areas, from entertainment to rehabilitation. In this context, this paper proposes a framework to recognize hand gestures using a limited number of landmarks from the video images. This hand gesture recognition system comprises an image processing module that extracts and processes the coordinates of 21 hand points called landmarks, and a deep neural network module that models and classifies the hand gestures. These landmarks are extracted automatically through MediaPipe software. The experiments were carried out over the IPN Hand dataset in an independent-user scenario using a Subject-Wise Cross Validation. They cover the use of different landmark-based formats, normalizations, lengths of the gesture representations, and number of landmarks used as inputs. The system obtains significantly better accuracy when using the raw coordinates of the 21 landmarks through 125 timesteps and a light Recurre nt Neural Network architecture (80.56 ± 1.19 %) or the hand anthropometric measures (82.20 ± 1.15 %) compared to using the speed of the hand landmarks through the gesture (72.93 ± 1.34 %). The proposed framework studied the effect of different landmark-based normalizations over the raw coordinates, obtaining an accuracy of 83.67 ± 1.12 % when using as reference the wrist landmark from each frame, and an accuracy of 84.66 ± 1.09 % when using as reference the wrist landmark from the first video frame of the current gesture. In addition, the proposed solution provided high recognition performance even when only using the coordinates from 6 (82.15 ± 1.16 %) or 4 (81.46 ± 1.17 %) specific hand landmarks using as reference the wrist landmark from the first video frame of the current gesture.
Manuel Gil-Martín, Marco Raoul Marini, Iván Martín-Fernández, Sergio Esteban Romero, Luigi Cinque
ICAART (3)1
2025 Towards Multi-View Hand Pose Recognition Using a Fusion of Image Embeddings and Leap 2 Landmarks
abstract
This paper presents a novel approach for multi-view hand pose recognition through image embeddings and hand landmarks. The method integrates raw image data with structural hand landmarks derived from the Leap Motion Controller 2. A Vision Transformer (ViT) pretrained model was used to extract visual features from dual-view grayscale images, which were fused with the corresponding Leap 2 hand landmarks, creating a multimodal representation that encapsulates both visual and landmark data for each sample. These fused embeddings were then classified using a multi-layer perceptron to distinguish among 17 distinct hand poses from the Multi-view Leap2 Hand Pose Dataset, which includes data from 21 subjects. Using a Leave-OneSubject-Out Cross-Validation (LOSO-CV) strategy, we demonstrate that this fusion approach offers a robust recognition performance (F1 Score of 79.33 ± 0.09 %), particularly in scenarios where hand occlusions or challenging angles may limit the utility of single-modality data.
Sergio Esteban Romero, Romeo Lanzino, Marco Raoul Marini, Manuel Gil-Martín
ICAART (3)4
2024 A Comprehensive Analysis of Parkinson's Disease Detection Through Inertial Signal Processing
abstract
When developing deep learning systems for Parkinson's Disease (PD) detection using inertial sensors, a comprehensive analysis of some key factors, including data distribution, signal processing domain, number of sensors, and analysis window size, is imperative to refine tremor detection methodologies. Leveraging the PD-BioStampRC21 dataset with accelerometer recordings, our state-of-the-art deep learning architecture extracts a PD biomarker. Applying Fast Fourier Transform (FFT) magnitude coefficients as a preprocessing step improves PD detection in Leave-One-Subject-Out Cross-Validation (LOSO CV), achieving 66.90% accuracy with a single sensor and 6.4-second windows, compared to 60.33% using raw samples. Integrating information from all five sensors boosts performance to 75.10%. Window size analysis shows that 3.2- second windows of FFT coefficients from all sensors outperform shorter or longer windows, with a window-level accuracy of 80.49% and a user-level accuracy of 93.55% in a LOSO scenario.
Manuel Gil-Martín, Sergio Esteban Romero, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
ICAART (3)1
2024 Parkinson's Disease Detection Through Inertial Signals and Posture Insights
abstract
In the development of deep learning systems aimed at detecting Parkinson's Disease (PD) using inertial sensors, some aspects could be essential to refine tremor detection methodologies in realistic scenarios. This work analyses the effect of the subjects’ posture during tremor recordings and the required amount of data to assess a proper PD detection in a Leave-One-Subject-Out Cross-Validation (LOSO CV) scenario. We propose a deep learning architecture that learns a PD biomarker from accelerometer signals to classify subjects between healthy and PD patients. This study uses the PD-BioStampRC21 dataset, containing accelerometer recordings from healthy and PD participants equipped with five inertial sensors. An increment of performance was obtained when using sitting windows compared to using lying windows for Fast Fourier Transform (FFT) input signal domain. Moreover, using 5 minutes per subject could be sufficient to properly evaluate the PD status of a patient without losing performance, reaching a windowlevel accuracy of 77.71 ± 1.07 % and a user-level accuracy of 87.10 ± 11.80 %. Furthermore, a knowledge transfer could be performed when training the system with sitting instances and testing with lying examples, indicating that the sitting activity contains valuable information that allows an effective generalization to lying instances.
Manuel Gil-Martín, Sergio Esteban Romero, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
ICAART (3)1
2024 Evaluating emotional and subjective responses in synthetic art-related dialogues: A multi-stage framework with large language models
abstract
The appearance of Large Language Models (LLM) has implied a qualitative step forward in the performance of conversational agents, and even in the generation of creative texts. However, previous applications of these models in generating dialogues neglected the impact of ‘hallucinations’ in the context of generating synthetic dialogues, thus omitting this central aspect in their evaluations. For this reason, we propose an open-source and flexible framework called GenEvalGPT framework: a comprehensive multi-stage evaluation strategy utilizing diverse metrics. The objective is two-fold: first, the goal is to assess the extent to which synthetic dialogues between a chatbot and a human align with the specified commands, determining the successful creation of these dialogues based on the provided specifications; and second, to evaluate various aspects of emotional and subjective responses. Assuming that dialogues to be evaluated were synthetically produced from specific profiles, the first evaluation stage utilizes LLMs to reconstruct the original templates employed in dialogue creation. The success of this reconstruction is then assessed in a second stage using lexical and semantic objective metrics. On the other hand, crafting a chatbot’s behaviors demands careful consideration to encompass a diverse range of interactions it is meant to engage in. Synthetic dialogues play a pivotal role in this context, as they can be deliberately synthesized to emulate various behaviors. This is precisely the objective of the third stage: evaluating whether the generated dialogues adhere to the required aspects concerning emotional and subjective responses. To validate the capabilities of the proposed framework, we applied it to recognize whether the chatbot exhibited one of two distinct behaviors in the synthetically generated dialogues: being emotional and providing subjective responses, or remaining neutral. This evaluation will encompass traditional metrics and automatic metrics generated by the LLM. In our use case of art-related dialogues, our findings reveal that the capacity to recover templates or profiles is more effective for information or profile items that are objective and factual, in contrast to those related to mental states or subjective facts. For the emotional and subjective behavior assessment, rule-based metrics achieved a 79% of accuracy in detecting emotions or subjectivity (anthropic), and an 82% on the LLM automatic metrics. The combination of these metrics and stages could help to decide which of the generated dialogues should be maintained depending on the applied policy, which could vary from preserving between 57% to 93% of the initial dialogues.
Cristina Luna Jiménez, Manuel Gil-Martín, Luis Fernando D'Haro, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
Expert Syst. Appl.2
2023 Video Memorability Prediction From Jointly-learnt Semantic and Visual Features
abstract
The memorability of a video is defined as an intrinsic property of its visual features that dictates the fraction of people who recall having watched it on a second viewing within a memory game. Still, unravelling what are the key features to predict memorability remains an obscure matter. This challenge is addressed here by fine-tuning text and image encoders using a cross-modal strategy known as Contrastive Language-Image Pre-training (CLIP). The resulting video-level data representations learned include semantics and topic-descriptive information as observed from both modalities, hence enhancing the predictive power of our algorithms. Our proposal achieves in the text domain a significantly greater Spearman Rank Correlation Coefficient (SRCC) than a default pre-trained text encoder (0.575 ± 0.007 and 0.538 ± 0.007, respectively) over the Memento10K dataset. A similar trend, although less pronounced, can be noticed in the visual domain. We believe these findings signal the potential benefits that cross-modal predictive systems can extract from being fine-tuned to the specific issue of media memorability.
Iván Martín-Fernández, Ricardo Kleinlein, Cristina Luna Jiménez, Manuel Gil-Martín, Fernando Fernández Martínez
CBMI4
2021 Time Analysis in Human Activity Recognition
Manuel Gil-Martín, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Javier Ferreiros
Neural Process. Lett.1
2020 Improving physical activity recognition using a new deep learning architecture and post-processing techniques
Manuel Gil-Martín, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Javier Ferreiros
Eng. Appl. Artif. Intell.1
2020 Robust Biometrics from Motion Wearable Sensors Using a D-vector Approach
Manuel Gil-Martín, Rubén San-Segundo-Hernández, Ricardo de Córdoba, José Manuel Pardo
Neural Process. Lett.1
2018 Robust Human Activity Recognition using smartwatches and smartphones
Rubén San-Segundo-Hernández, Henrik Blunck, José Moreno-Pimentel, Allan Stisen, Manuel Gil-Martín
Eng. Appl. Artif. Intell.5