Fernando Fernández Martínez

dblp:13/11156 · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-3877-0089ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Optimizing Sequential Models through Temporal Landmark Selection and Normalization for Sign Language Recognition
Sergio Esteban Romero, Iván Martín-Fernández, Cristina Luna Jiménez, Manuel Gil-Martín, Fernando Fernández Martínez, Elisabeth André
ICAART (3)5
2026 A Case Study on Large Visual-Language Model Attention Explainability After Adaptation Using Persuasion Strategies in Advertisements
Iván Martín-Fernández, Mihai Gabriel Constantin, Bogdan Ionescu, Sergio Esteban Romero, Fernando Fernández Martínez, Manuel Gil-Martín
MMM (1)5
2026 A comprehensive study on contrastive pre-training and fine tuning of vision and text transformers for video memorability prediction
abstract
Abstract Video memorability prediction has emerged as a key challenge for improving information retrieval, content design, and user engagement. Prior work has shown that semantic cues play a crucial role in determining memorability, with recent studies leveraging Contrastive Language-Image Pre-training (CLIP) encoders to incorporate semantic information. However, the specific improvements attributable to CLIP models remain unclear, as few studies systematically compare their performance against equivalent unimodal encoders or explore fine-tuning strategies. This work addresses that gap through a comprehensive, controlled evaluation of CLIP-based and unimodal encoders for video memorability prediction. We propose FCLIP, a domain-adapted extension of CLIP that undergoes additional contrastive pre-training on memorability-specific image-text pairs. Our experiments assess both feature extraction and supervised fine-tuning, ensuring fair comparisons across models with matched architecture and parameter count. Results show that FCLIP image encoders achieve a Spearman Rank Correlation Coefficient (SRCC) of 0.672 on the Memento10k dataset, significantly outperforming unimodal Vision Transformers. FCLIP text encoders similarly outperform unimodal baselines, reaching an SRCC of 0.632. These findings demonstrate that contrastive learning and domain adaptation substantially improve memorability prediction, highlighting the importance of semantic and multimodal pre-training in developing advanced content analysis systems.
Iván Martín-Fernández, Sergio Esteban Romero, Manuel Gil-Martín, Fernando Fernández Martínez
Multim. Tools Appl.4
2026 Principled Evaluation of Multi-Label Persuasion in Advertisements with Large Vision-Language Models
abstract
The automatic detection of persuasive strategies in advertisements presents a uniquely multimodal challenge at the intersection of vision, language, and social cognition. While recent advances in Large Vision-Language Models (LVLMs) offer promising capabilities for such tasks, current approaches often rely on restrictive evaluation schemes that do not reflect the inherently multi-label nature of persuasive messaging. In this work, we reframe persuasion strategy detection as a genuine multi-label classification problem and propose a principled evaluation framework to enhance interpretability and robustness. We apply this approach to both image and video datasets, examine their characteristics, and introduce novel input-agnostic baselines that achieve macro F1 scores of 0.082 and 0.289 on the respective test sets. As part of this analysis, we study label co-occurrence patterns and dataset ambiguities, providing insights that inform both model interpretation and future dataset design. To assess the native capabilities of LVLMs, we benchmark three open source models—PaliGemma, PaliGemma2, and Qwen2.5-VL—on the image persuasion dataset. Under zero-shot conditions, we demonstrate that querying each strategy individually with a logit-based decision threshold outperforms guided text generation. The best-performing zero-shot model, Qwen2.5-VL, achieves a macro F1 score of 0.227 and a sample F1 score of 0.234 on the image test set, and 0.381 and 0.400 respectively on the video test set. We further explore and compare two lightweight fine-tuning strategies that update only small subsets of model parameters while keeping the remaining weights frozen: fine-tuning of the image-to-text tokens linear projection and Low Rank Adaptation (LoRA) of the language model. Linear projector fine-tuning yields a top macro F1 score of 0.396 on the image test set, marking a substantial improvement over zero-shot performance. To evaluate cross-modal generalization, we apply fine-tuned image models to the video dataset. Our experiments reveal that, while projection-based fine-tuning enables partial knowledge transfer from image to video (macro F1 = 0.416 on a testing subset), LoRA adaptation severely disrupts cross-modal performance (macro F1 = 0.189). Finally, we perform a per-strategy performance analysis, looking into annotator- and data-centric factors that may influence LVLM performance. These findings highlight the viability of open-weight LVLMs for fine-grained persuasion analysis and suggest efficient pathways for domain-specific adaptation under realistic resource constraints.
Iván Martín-Fernández, Mihai Gabriel Constantin, Bogdan Ionescu, Manuel Gil-Martín, Fernando Fernández Martínez
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Exploring the Effect of Size, Architecture and Fine-Tuning Hyperparameters on Large Visual-Language Model Adaptation for Video Memorability Prediction
abstract
Predicting video memorability involves modeling audiovisual content to estimate its likelihood of being remembered-relevant in areas like recommendation, marketing, and education. Leveraging the semantic power of Large Vision-Language Models (LVLMs), we adapt Qwen2.5-VL models using Low-Rank Adaptation (LoRA) to this task. We evaluate the influence of model size (3B vs. 7B) and LoRA hyperparameters ($r$, α) on performance. Our best result, an SRCC of 0.7658 on the Memento10k development dataset using 5-Fold Cross Validation, surpasses previous approaches and confirms the importance of careful tuning. Notably, the model with the best zero-shot performance (3B) also achieved the highest fine-tuned score, showing that zero-shot evaluation can predict adaptation success. Results challenge assumptions about scaling in LVLMs and highlight the potential of smaller, well-adapted models.
David Luna-García, Iván Martín-Fernández, Sergio Esteban Romero, Manuel Gil-Martín, Fernando Fernández Martínez
CBMI5
2024 A Comprehensive Analysis of Parkinson's Disease Detection Through Inertial Signal Processing
abstract
When developing deep learning systems for Parkinson's Disease (PD) detection using inertial sensors, a comprehensive analysis of some key factors, including data distribution, signal processing domain, number of sensors, and analysis window size, is imperative to refine tremor detection methodologies. Leveraging the PD-BioStampRC21 dataset with accelerometer recordings, our state-of-the-art deep learning architecture extracts a PD biomarker. Applying Fast Fourier Transform (FFT) magnitude coefficients as a preprocessing step improves PD detection in Leave-One-Subject-Out Cross-Validation (LOSO CV), achieving 66.90% accuracy with a single sensor and 6.4-second windows, compared to 60.33% using raw samples. Integrating information from all five sensors boosts performance to 75.10%. Window size analysis shows that 3.2- second windows of FFT coefficients from all sensors outperform shorter or longer windows, with a window-level accuracy of 80.49% and a user-level accuracy of 93.55% in a LOSO scenario.
Manuel Gil-Martín, Sergio Esteban Romero, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
ICAART (3)3
2024 Parkinson's Disease Detection Through Inertial Signals and Posture Insights
abstract
In the development of deep learning systems aimed at detecting Parkinson's Disease (PD) using inertial sensors, some aspects could be essential to refine tremor detection methodologies in realistic scenarios. This work analyses the effect of the subjects’ posture during tremor recordings and the required amount of data to assess a proper PD detection in a Leave-One-Subject-Out Cross-Validation (LOSO CV) scenario. We propose a deep learning architecture that learns a PD biomarker from accelerometer signals to classify subjects between healthy and PD patients. This study uses the PD-BioStampRC21 dataset, containing accelerometer recordings from healthy and PD participants equipped with five inertial sensors. An increment of performance was obtained when using sitting windows compared to using lying windows for Fast Fourier Transform (FFT) input signal domain. Moreover, using 5 minutes per subject could be sufficient to properly evaluate the PD status of a patient without losing performance, reaching a windowlevel accuracy of 77.71 ± 1.07 % and a user-level accuracy of 87.10 ± 11.80 %. Furthermore, a knowledge transfer could be performed when training the system with sitting instances and testing with lying examples, indicating that the sitting activity contains valuable information that allows an effective generalization to lying instances.
Manuel Gil-Martín, Sergio Esteban Romero, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
ICAART (3)3
2024 Evaluating emotional and subjective responses in synthetic art-related dialogues: A multi-stage framework with large language models
abstract
The appearance of Large Language Models (LLM) has implied a qualitative step forward in the performance of conversational agents, and even in the generation of creative texts. However, previous applications of these models in generating dialogues neglected the impact of ‘hallucinations’ in the context of generating synthetic dialogues, thus omitting this central aspect in their evaluations. For this reason, we propose an open-source and flexible framework called GenEvalGPT framework: a comprehensive multi-stage evaluation strategy utilizing diverse metrics. The objective is two-fold: first, the goal is to assess the extent to which synthetic dialogues between a chatbot and a human align with the specified commands, determining the successful creation of these dialogues based on the provided specifications; and second, to evaluate various aspects of emotional and subjective responses. Assuming that dialogues to be evaluated were synthetically produced from specific profiles, the first evaluation stage utilizes LLMs to reconstruct the original templates employed in dialogue creation. The success of this reconstruction is then assessed in a second stage using lexical and semantic objective metrics. On the other hand, crafting a chatbot’s behaviors demands careful consideration to encompass a diverse range of interactions it is meant to engage in. Synthetic dialogues play a pivotal role in this context, as they can be deliberately synthesized to emulate various behaviors. This is precisely the objective of the third stage: evaluating whether the generated dialogues adhere to the required aspects concerning emotional and subjective responses. To validate the capabilities of the proposed framework, we applied it to recognize whether the chatbot exhibited one of two distinct behaviors in the synthetically generated dialogues: being emotional and providing subjective responses, or remaining neutral. This evaluation will encompass traditional metrics and automatic metrics generated by the LLM. In our use case of art-related dialogues, our findings reveal that the capacity to recover templates or profiles is more effective for information or profile items that are objective and factual, in contrast to those related to mental states or subjective facts. For the emotional and subjective behavior assessment, rule-based metrics achieved a 79% of accuracy in detecting emotions or subjectivity (anthropic), and an 82% on the LLM automatic metrics. The combination of these metrics and stages could help to decide which of the generated dialogues should be maintained depending on the applied policy, which could vary from preserving between 57% to 93% of the initial dialogues.
Cristina Luna Jiménez, Manuel Gil-Martín, Luis Fernando D'Haro, Fernando Fernández Martínez, Rubén San-Segundo-Hernández
Expert Syst. Appl.4
2023 Video Memorability Prediction From Jointly-learnt Semantic and Visual Features
abstract
The memorability of a video is defined as an intrinsic property of its visual features that dictates the fraction of people who recall having watched it on a second viewing within a memory game. Still, unravelling what are the key features to predict memorability remains an obscure matter. This challenge is addressed here by fine-tuning text and image encoders using a cross-modal strategy known as Contrastive Language-Image Pre-training (CLIP). The resulting video-level data representations learned include semantics and topic-descriptive information as observed from both modalities, hence enhancing the predictive power of our algorithms. Our proposal achieves in the text domain a significantly greater Spearman Rank Correlation Coefficient (SRCC) than a default pre-trained text encoder (0.575 ± 0.007 and 0.538 ± 0.007, respectively) over the Memento10K dataset. A similar trend, although less pronounced, can be noticed in the visual domain. We believe these findings signal the potential benefits that cross-modal predictive systems can extract from being fine-tuned to the specific issue of media memorability.
Iván Martín-Fernández, Ricardo Kleinlein, Cristina Luna Jiménez, Manuel Gil-Martín, Fernando Fernández Martínez
CBMI5
2022 Sampling Based On Natural Image Statistics Improves Local Surrogate Explainers
Ricardo Kleinlein, Alexander Hepburn, Raúl Santos-Rodríguez, Fernando Fernández Martínez
BMVC4
2021 A dynamic term discovery strategy for automatic speech recognizers with evolving dictionaries
Alejandro Coucheiro-Limeres, Javier Ferreiros, Fernando Fernández Martínez, Ricardo de Córdoba
Expert Syst. Appl.3
2021 Time Analysis in Human Activity Recognition
Manuel Gil-Martín, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Javier Ferreiros
Neural Process. Lett.3
2020 Improving physical activity recognition using a new deep learning architecture and post-processing techniques
Manuel Gil-Martín, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Javier Ferreiros
Eng. Appl. Artif. Intell.3
2019 Attention-Based Word Vector Prediction with LSTMs and its Application to the OOV Problem in ASR
abstract
We propose three architectures for a word vector prediction system (WVPS) built with LSTMs that consider both past and future contexts of a word for predicting a vector in an embedded space where its surrounding area is semantically related to the considered word. We introduce an attention mechanism in one of the architectures so the system is able to assess the specific contribution of each context word to the prediction. All the architectures are trained under the same conditions and the same training material, following a curricular-learning fashion in the presentation of the data. For the inputs, we employ pretrained word embeddings. We evaluate the systems after the same number of training steps, over two different corpora composed of ground-truth speech transcriptions in Spanish language from TCSTAR and TV recordings used in the Search on Speech Challenge of IberSPEECH 2018. The results show that we are able to reach significant differences between the architectures, consistently across both corpora. The attention-based architecture achieves the best results, suggesting its adequacy for the task. Also, we illustrate the usefulness of the systems for resolving out-of-vocabulary (OOV) regions marked by an ASR system capable of detecting OOV occurrences
Alejandro Coucheiro-Limeres, Fernando Fernández Martínez, Rubén San-Segundo-Hernández, Javier Ferreiros
INTERSPEECH2
2019 Predicting Group-Level Skin Attention to Short Movies from Audio-Based LSTM-Mixture of Experts Models
abstract
Electrodermal activity (EDA) is a psychophysiological indicator that can be considered a somatic marker of the emotional and attentional reaction of subjects towards stimuli like audiovisual content. EDA measurements are not biased by the cognitive process of giving an opinion or a score to characterize the subjective perception, and group-level EDA recordings integrate the reaction of an audience, thus reducing the signal noise. This paper contributes to the field of audience's attention prediction to video content, extending previous novel work on the use of EDA as ground truth for prediction algorithms. Videos are segmented into shorter clips attending to the audience's increasing or decreasing attention, and we process videos' audio waveform to extract meaningful aural embeddings from a VG-Gish model pretrained on the Audioset database. While previous similar work on attention level prediction using only audio accomplished 69.83% accuracy, we propose a Mixture of Experts approach to train a binary classifier that outperforms the main existing state-of-the-art approaches predicting increasing and decreasing attention levels with 81.76% accuracy. These results confirm the usefulness of providing acoustic features with a semantic significance, and the convenience of considering experts over partitions of the dataset in order to predict group-level attention from audio.
Ricardo Kleinlein, Cristina Luna Jiménez, Juan Manuel Montero-Martínez, Zoraida Callejas Carrión, Fernando Fernández Martínez
INTERSPEECH5
2019 A multi-threshold approach and a realistic error measure for vanishing point detection in natural landscapes
Álvaro García-Faura, Fernando Fernández Martínez, Ricardo Kleinlein, Rubén San-Segundo-Hernández, Fernando Díaz-de-María
Eng. Appl. Artif. Intell.2
2019 Emotion and attention: Audiovisual models for group-level skin response recognition in short movies
abstract
The electrodermal activity (EDA) is a psychophysiological indicator which can be considered a somatic marker of the emotional and attentional reaction of subjects towards stimuli. EDA measurements are not biased by the cognitive process of giving an opinion or a score to characterize the subjective perception, and group-level EDA recordings integrate the reaction of the whole audience, thus reducing the signal noise. This paper contributes to the field of affective video content analysis, extending previous novel work on the use of EDA as ground truth for prediction algorithms. Here, we label short video clips according to the audience’s emotion (high vs. low) and attention (increasing vs. decreasing), derived from EDA records. Then, we propose a set of low-level audiovisual descriptors and train binary classifiers that predict the emotion and attention with 75% and 80% accuracy, respectively. These results, along with those of previous works, reinforce the usefulness of such low-level audiovisual descriptors to model video in terms of the induced affective response.
Álvaro García-Faura, Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María, Rubén San-Segundo-Hernández
Web Intell.3
2018 Exploiting visual saliency for assessing the impact of car commercials upon viewers
Fernando Fernández Martínez, René Arnulfo García-Hernández, Miguel Angel Fernandez-Torres, Iván González-Díaz 0001, Álvaro García-Faura, Fernando Díaz-de-María
Multim. Tools Appl.1
2017 Emotion and attention: predicting electrodermal activity through video visual descriptors
abstract
This paper contributes to the field of affective video content analysis through the novel employment of electrodermal activity (EDA) measurements as ground truth for machine learning algorithms. The variation of the electrical properties of the skin, known as EDA, is a psychophysiological indicator widely used in medicine, psychology and neuroscience which can be considered a somatic marker of the emotional and attentional reaction of subjects towards stimuli. One of its main advantages is that the recorded information is not biased by the cognitive process of giving an opinion or a score to characterize the subjective perception. In this work, we predict the levels of emotion and attention, derived from EDA records, by means of a small set of low-level visual descriptors computed from the video stimuli. Linear regression experiments show that our descriptors predict significantly well the sum of emotion and attention levels, reaching a coefficient of determination R2 = 0.25. This result sets a promising path for further research on the prediction of emotion and attention from videos using EDA.
Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María
WI2
2016 Feature extraction from smartphone inertial signals for human activity segmentation
Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Roberto Barra-Chicote, Fernando Fernández Martínez, José Manuel Pardo
Signal Process.4
2016 Comparing visual descriptors and automatic rating strategies for video aesthetics prediction
Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María
Signal Process. Image Commun.2
2015 Towards a robust affect recognition: Automatic facial expression recognition in 3D faces
Amal Azazi, Syaheerah L. Lutfi, Ibrahim Venkat, Fernando Fernández Martínez
Expert Syst. Appl.4
2015 Succeeding metadata based annotation scheme and visual tips for the automatic assessment of video aesthetic quality in car commercials
Fernando Fernández Martínez, Alejandro Hernández-García, Fernando Díaz-de-María
Expert Syst. Appl.1
2014 A web-based application for the management and evaluation of tutoring requests in PBL-based massive laboratories
abstract
One important steps in a successful project-based-learning methodology (PBL) is the process of providing the students with a convenient feedback that allows them to keep on developing their projects or to improve them. However, this task is more difficult in massive courses, especially when the project deadline is close. Besides, the continuous evaluation methodology makes necessary to find ways to objectively and continuously measure students' performance without increasing excessively instructors' work load. In order to alleviate these problems, we have developed a web service that allows students to request personal tutoring assistance during the laboratory sessions by specifying the kind of problem they have and the person who could help them to solve it. This service provides tools for the staff to manage the laboratory, for performing continuous evaluation for all students and for the student collaborators, and to prioritize tutoring according to the progress of the student's project. Additionally, the application provides objective metrics which can be used at the end of the subject during the evaluation process in order to support some students' final scores. Different usability statistics and the results of a subjective evaluation with more than 330 students confirm the success of the proposed application.
Luis Fernando D'Haro, Fernando Fernández Martínez, Ricardo de Córdoba, Juan Manuel Montero-Martínez
FIE2
2013 NEMOHIFI: an affective HiFi agent
abstract
This demo concerns a recently developed prototype of an emotionally-sensitive autonomous HiFi Spoken Conversational Agent, called NEMOHIFI. The baseline agent was developed by the Speech Technology Group (GTH) and has recently been integrated with an emotional engine called NEMO (Need-inspired Emotional Model) to enable it to adapt to users' emotion and respond to the users using appropriate expressive speech. NEMOHIFI controls and manages the HiFi audio system, and for end users, its functions equate a remote control, except that instead of clicking, the user interacts with the agent using voice. A pairwise comparison between the baseline (non-adaptive) and NEMO-HIFI showed that the latter was not only statistically substantially preferred by users to the former, but they are also significantly more satisfied with it than the former.
Syaheerah L. Lutfi, Fernando Fernández Martínez, Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Juan Manuel Montero-Martínez
ICMI2
2013 On the dynamic adaptation of language models based on dialogue information
Juan Manuel Lucas, Javier Ferreiros, Fernando Fernández Martínez, Julián D. Echeverry-Correa, Syaheerah L. Lutfi
Expert Syst. Appl.3
2013 A satisfaction-based model for affect recognition from conversational features in spoken dialog systems
Syaheerah L. Lutfi, Fernando Fernández Martínez, Juan Manuel Lucas, Lorena Lopez-Lebon, Juan Manuel Montero-Martínez
Speech Commun.2
2012 Investigating Verbal Intelligence Using the TF-IDF Approach
Kseniya Zablotskaya, Fernando Fernández Martínez, Wolfgang Minker
LREC2
2012 Relating Dominance of Dialogue Participants with their Verbal Intelligence Scores
Kseniya Zablotskaya, Umair Rahim, Fernando Fernández Martínez, Wolfgang Minker
LREC3
2012 Estimating Adaptation of Dialogue Partners with Different Verbal Intelligence
Kseniya Zablotskaya, Fernando Fernández Martínez, Wolfgang Minker
SIGDIAL Conference2
2012 Text categorization methods for automatic estimation of verbal intelligence
Fernando Fernández Martínez, Kseniya Zablotskaya, Wolfgang Minker
Expert Syst. Appl.1
2012 Towards building intelligent speech interfaces through the use of more flexible, robust and natural dialogue management solutions
abstract
In this paper a Bayesian Networks-based solution for dialogue modelling is presented. This solution is combined with carefully designed contextual information handling strategies. With the purpose of validating these solutions, and introducing a spoken dialogue system for controlling a Hi-Fi audio system as the selected prototype, a real-user evaluation has been conducted. Two different versions of the prototype are compared. Each version corresponds to a different implementation of the algorithm for the management of the actuation order, the algorithm for deciding the proper order to carry out the actions required by the user. The evaluation is carried out in terms of a battery of both subjective and objective metrics collected from speakers interacting with the Hi-Fi audio box through predefined scenarios. Defined metrics have been specifically adapted to measure: first, the usefulness and the actual relevance of the proposed solutions, and, secondly, their joint performance through their intelligent combination mainly measured as the level achieved with regard to the user satisfaction. A thorough and comprehensive study of the main differences between both approaches is presented. Two-way analysis of variance (ANOVA) tests are also included to measure the effects of both: the system used and the type of scenario factors, simultaneously. Finally, the effect of bringing this flexibility, robustness and naturalness into our home dialogue system is also analyzed through the results obtained. These results show that the intelligence of our speech interface has been well perceived, highlighting its excellent ease of use and its good acceptance by users, therefore validating the approached dialogue management solutions and demonstrating that a more natural, flexible and robust dialogue is possible thanks to them.
Fernando Fernández Martínez, Javier Ferreiros, Juan Manuel Lucas, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández, Ricardo de Córdoba
Interact. Comput.1
2012 Design, development and field evaluation of a Spanish into sign language translation system
Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Ricardo de Córdoba, Valentín Sama Rojo, Fernando Fernández Martínez, Luis Fernando D'Haro, Verónica López-Ludeña, D. Sánchez
Pattern Anal. Appl.5
2011 Evaluation of a User-adapted Spoken Language Dialogue System - Measuring the Relevance of the Contextual Information Sources
Juan Manuel Lucas, Fernando Fernández Martínez, G. Dragos Rada, Syaheerah L. Lutfi, Javier Ferreiros
ICAART (1)2
2010 HIFI-AV: An Audio-visual Corpus for Spoken Language Human-Machine Dialogue Research in Spanish
Fernando Fernández Martínez, Juan Manuel Lucas, Roberto Barra-Chicote, Javier Ferreiros, Javier Macías Guarasa
LREC1
2009 A Bayesian NETWORKS approach for dialog modeling: The fusion BN
abstract
Bayesian networks, BNs, are suitable for mixed-initiative dialog modeling allowing a more flexible and natural spoken interaction. This solution can be applied to identify the intention of the user considering the concepts extracted from the last utterance and the dialog context. Subsequently, in order to make a correct decision regarding how the dialog should continue, unnecessary, missing, wrong, optional and required concepts have to be detected according to the inferred goals. This information is useful to properly drive the dialog prompting for missing concepts, clarifying for wrong concepts, ignoring unnecessary concepts and retrieving those required and optional. This paper presents a novel BNs approach where a single BN is obtained from N goal-specific BNs through a fusion process. The new fusion BN enables a single concept analysis which is more consistent with the whole dialog context.
Fernando Fernández Martínez, Javier Ferreiros, Ricardo de Córdoba, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández, José Manuel Pardo
ICASSP1
2009 Acoustic emotion recognition using dynamic Bayesian networks and multi-space distributions
abstract
In this paper we describe the acoustic emotion recognition system built at the Speech Technology Group of the Universidad Politecnica de Madrid (Spain) to participate in the INTERSPEECH 2009 Emotion Challenge. Our proposal is based on the use of a Dynamic Bayesian Network (DBN) to deal with the temporal modelling of the emotional speech information. The selected features (MFCC, F0, Energy and their variants) are modelled as different streams, and the F0 related ones are integrated under a Multi Space Distribution (MSD) framework, to properly model its dual nature (voiced/unvoiced). Experimental evaluation on the challenge test set, show a 67.06%and 38.24% of unweighted recall for the 2 and 5-classes tasks respectively. In the 2-class case, we achieve similar results compared with the baseline, with a considerable less number of features. In the 5-class case, we achieve a statistically significant 6.5% relative improvement
Roberto Barra-Chicote, Fernando Fernández Martínez, Syaheerah L. Lutfi, Juan Manuel Lucas, Javier Macías Guarasa, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández, José Manuel Pardo
INTERSPEECH2
2009 Using dialogue-based dynamic language models for improving speech recognition
abstract
We present a new approach to dynamically create and manage different language models to be used on a spoken dialogue system. We apply an interpolation based approach, using several measures obtained by the DialogueManager to decide what LM the system will interpolate and also to estimate the interpolation weights. We propose to use not only semantic information (the concepts extracted from each recognized utterance), but also information obtained by the dialogue manager module (DM), that is, the objectives or goals the user wants to fulfill, and the proper classification of those concepts according to the inferred goals. The experiments we have carried out show improvements over word error rate when using the parsed concepts and the inferred goals from a speech utterance for rescoring the same utterance.
Juan Manuel Lucas, Fernando Fernández Martínez, Javier Ferreiros
INTERSPEECH2
2008 Evaluation of a spoken dialogue system for controlling a Hifi audio system
abstract
In this paper a Bayesian Networks, BNs, approach to dialogue modelling [1] is evaluated in terms of a battery of both subjective and objective metrics. A significant effort in improving the contextual information handling capabilities of the system has been done. Consequently, besides typical dialogue measurement rates for usability like task or dialogue completion rates, dialogue time, etc. we have included a new figure measuring the contextuality of the dialogue as the number of turns where contextual information is helpful for dialogue resolution. The evaluation is developed through a set of predefined scenarios according to different initiative styles and focusing on the impact of the user's level of experience.
Fernando Fernández Martínez, Juan Blázquez, Javier Ferreiros, Roberto Barra-Chicote, Javier Macías Guarasa, Juan Manuel Lucas
SLT1
2008 Speech to sign language translation system for Spanish
Rubén San-Segundo-Hernández, Roberto Barra-Chicote, Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Javier Ferreiros, Juan Manuel Lucas, Javier Macías Guarasa, Juan Manuel Montero-Martínez, José Manuel Pardo
Speech Commun.5
2007 Language identification based on n-gram frequency ranking
abstract
We present a novel approach for language identification based on a text categorization technique, namely an n-gram frequency ranking. We use a Parallel phone recognizer, the same as in PPRLM, but instead of the language model, we create a ranking with the most frequent n-grams, keeping only a fraction of them. Then we compute the distance between the input sentence ranking and each language ranking, based on the difference in relative positions for each n-gram. The objective of this ranking is to be able to model reliably a longer span than PPRLM, namely 5-gram instead of trigram, because this ranking will need less training data for a reliable estimation. We demonstrate that this approach overcomes PPRLM (6 % relative improvement) due to the inclusion of 4gram and 5-gram in the classifier. We present two alternatives: ranking with absolute values for the number of occurrences and ranking with discriminative values (11% relative improvement). Index Terms: Language Identification, n-gram frequency ranking, text categorization, PPRLM
Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Javier Macías Guarasa, Javier Ferreiros
INTERSPEECH3
2007 Language identification using several sources of information with a multiple-Gaussian classifier
abstract
We present several innovative techniques that can be applied in a PPRLM system for language identification (LID). To normalize the scores, eliminate the bias in the scores and improve the classifier, we compared the bias removal technique (up to 19 % relative improvement (RI)) and a Gaussian classifier (up to 37 % RI). Then, we include additional sources of information in different feature vectors of the Gaussian classifier: the sentence acoustic score (11% RI), the average acoustic score for each phoneme (11 % RI), and the average duration for each phoneme (7.8 % RI). The use of a multiple-Gaussian classifier with 4 feature vectors meant an additional 15.1 % RI. Using 4 feature vectors instead of just PPRLM provides a 26.1 % RI. Finally, we include additional acoustic HMMs of the same language with success (10 % relative improvement). We will show how all these improvements have been mostly additive.
Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Juan Manuel Montero-Martínez, Roberto Barra-Chicote
INTERSPEECH3
2005 Speech interface for controlling an hi-fi audio system based on a Bayesian belief networks approach for dialog modeling
abstract
This paper presents the development of a speech interface for controlling a high fidelity system from natural language sentences. A Bayesian Belief Network approach is proposed for dialog modeling. This solution is applied to infer the user’s goals corresponding to the processed utterances. Subsequently, from the inferred goals, missing or spurious concepts are automatically detected. This is used to drive the dialog prompting for missing concepts and clarifying for spurious concepts allowing more flexible and natural dialogs. A dialog strategy which makes use of the dialog history and the system’s state is also presented.
Fernando Fernández Martínez, Javier Ferreiros, Valentín Sama Rojo, Juan Manuel Montero-Martínez, Rubén San-Segundo-Hernández, Javier Macías Guarasa, Rafael García
INTERSPEECH1
2005 New word-level and sentence-level confidence scoring using graph theory calculus and its evaluation on speech understanding
abstract
A lot of work has been devoted to the estimation of confidence measures for speech recognizers. In the quite extended case where a word-graph speech recognizer is in use, we will present new confidence measures employing the graph theory that shows us how to estimate some interesting characteristics about the different paths through the graph that constitute the recognition solutions, without the need of expanding them all. We will take advantage of some of these features to generate confidence scores both at the word and sentence level. We will also compare this new confidence scoring to more traditional ones and will find similar behavior with less computational load and with an increase in the simplicity of the approach that will lead to more generalization power of the confidence estimation to different applications of the recognizer.
Javier Ferreiros, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Luis Fernando D'Haro, Valentín Sama Rojo, Roberto Barra-Chicote, Pedro Mellén
INTERSPEECH3
2004 Language identification techniques based on full recognition in an air traffic control task
abstract
Automatic language identification has become an important issue in recent years in speech recognition systems.In this paper, we present the work done in language identification for an air traffic control speech recognizer for continuous speech.The system is able to distinguish between Spanish and English.We present several language identification techniques based on full recognition that improve the baseline results obtained using the most commonly known "PPRLM" technique.We have in our database some task specific critical problems for language identification like non native speakers, extremely spontaneous speech or Spanish-English mix in the same sentence.We confirm that PPRLM is quite sensible to those problems and that a technique based on a Bayesian classifier is the one with the best performance in spite of its higher computational cost.
Ricardo de Córdoba, Javier Ferreiros, Valentín Sama Rojo, Javier Macías Guarasa, Luis Fernando D'Haro, Fernando Fernández Martínez
INTERSPEECH6
2004 Implementation of dialog applications in an open-source voiceXML platform
abstract
In this paper, we study the approach followed to use the VoiceXML standard in a dialog system platform already available in our group. As VoiceXML interpreter we have chosen OpenVXI, an open source portable solution where we can make the modifications needed to adapt the solution to the characteristics of our recognition and synthesis modules; so we will emphasize the changes that we have had to make in such interpreter. Besides, we review some relevant modules in our platform and their capabilities, highlighting the use of standards in them, as SSML for the text-to-speech system and JSGF for the specification of grammars for recognition. Finally, we discuss several ideas regarding the limitations detected in VoiceXML.
Fernando Fernández Martínez, Valentín Sama Rojo, Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba, Juan Manuel Montero-Martínez
INTERSPEECH1