EDBT 2026 Demo / reviewers in the wild / expert
Luis Fernando D'Haro
dblp:57/1419 · also Luis F. D'Haro, Luis Fernando D'Haro-Enríquez
· DBLP profile ↗
55ranked-venue papers
17as first author
15since 2021 · last 2026
0000-0002-3411-7384ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 11 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DREAM: A Multicultural Multimodal Dataset Linking Dialogues and Realistic Image Sequences
Juan Mallo, Marcos Estecha-Garitagoitia, Ricardo de Córdoba, Luis Fernando D'Haro |
LREC | 4 |
| 2026 | Speech emotion recognition using multimodal LLMs and quality-controlled TTS-based data augmentation for Iberian languagesabstractThis work proposes the usage of multimodal large language models for speech emotion recognition (SER) on low-resource settings for Iberian languages. Given the existing low amount of annotated SER data for other languages, we also propose a pipeline for generating high-quality synthetic data using existing emotional text-to-speech (TTS) and their cloning capabilities. Specifically, we design a selective, quality-controlled TTS pipeline combining LLM-ensemble translation with self-verification and expressive voice cloning, followed by automatic ASR-WER, speaker-similarity, and emotion filters. This approach introduces a novel filtering strategy that ensures synthetic data reliability. The resulting data include MSP-MEA, 1 1 https://huggingface.co/datasets/jaimebellver/SER-MSPMEA-Spanish . a synthetic Spanish extension of MSP-Podcast. Building on our previous multimodal SER framework, we compare the usage of frozen LLM as classifiers with an MLP baseline and evaluate classical versus TTS-based augmentation across five corpora (IEMOCAP, MEACorpus, EMS, VERBO, AhoEmo3). The best configuration (W2v-BERT-2 → attentive pooling → frozen Bloomz-7b1) improves mean F1 by +4.9 point increase over a MLP head baseline. Among augmentation techniques, Mix-up remains the most robust overall, while TTS achieves competitive performance, surpassing traditional data augmentation techniques on EMS and VERBO. These results indicate that carefully filtered TTS data can complement classical perturbations, providing a viable, dataset-dependent strategy for multilingual SER. Code, models, and datasets are publicly released. 2 2 https://github.com/jaimebs2/SpeechFactory . • Propose SER with frozen LLMs over acoustic embeddings on 5 multilingual corpora. • Release a quality-controlled TTS data augmentation pipeline for emotional speech. • Show that TTS-based augmentation can match or surpass classical methods. • Release trained models and MSP-MEA, the first synthetic Spanish extension of MSP-Podcast. Jaime Bellver-Soler, Anmol Guragain, Samuel Ramos-Varela, Ricardo de Córdoba, Luis Fernando D'Haro |
Comput. Speech Lang. | 5 |
| 2024 | A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsabstractAutomatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis. Chen Zhang 0020, Luis Fernando D'Haro, Yiming Chen 0010, Malu Zhang, Haizhou Li 0001 |
AAAI | 2 |
| 2024 | Understanding and Improving the Next Generation of Conversational Systems: Trends, Challenges and Future
Luis Fernando D'Haro |
ICAART | 1 |
| 2024 | CVQA: Culturally-diverse Multilingual Visual Question Answering BenchmarkabstractVisual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field. Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji |
NeurIPS | 70 |
| 2024 | Evaluating emotional and subjective responses in synthetic art-related dialogues: A multi-stage framework with large language modelsabstractThe appearance of Large Language Models (LLM) has implied a qualitative step forward in the performance of conversational agents, and even in the generation of creative texts. However, previous applications of these models in generating dialogues neglected the impact of ‘hallucinations’ in the context of generating synthetic dialogues, thus omitting this central aspect in their evaluations. For this reason, we propose an open-source and flexible framework called GenEvalGPT framework: a comprehensive multi-stage evaluation strategy utilizing diverse metrics. The objective is two-fold: first, the goal is to assess the extent to which synthetic dialogues between a chatbot and a human align with the specified commands, determining the successful creation of these dialogues based on the provided specifications; and second, to evaluate various aspects of emotional and subjective responses. Assuming that dialogues to be evaluated were synthetically produced from specific profiles, the first evaluation stage utilizes LLMs to reconstruct the original templates employed in dialogue creation. The success of this reconstruction is then assessed in a second stage using lexical and semantic objective metrics. On the other hand, crafting a chatbot’s behaviors demands careful consideration to encompass a diverse range of interactions it is meant to engage in. Synthetic dialogues play a pivotal role in this context, as they can be deliberately synthesized to emulate various behaviors. This is precisely the objective of the third stage: evaluating whether the generated dialogues adhere to the required aspects concerning emotional and subjective responses. To validate the capabilities of the proposed framework, we applied it to recognize whether the chatbot exhibited one of two distinct behaviors in the synthetically generated dialogues: being emotional and providing subjective responses, or remaining neutral. This evaluation will encompass traditional metrics and automatic metrics generated by the LLM. In our use case of art-related dialogues, our findings reveal that the capacity to recover templates or profiles is more effective for information or profile items that are objective and factual, in contrast to those related to mental states or subjective facts. For the emotional and subjective behavior assessment, rule-based metrics achieved a 79% of accuracy in detecting emotions or subjectivity (anthropic), and an 82% on the LLM automatic metrics. The combination of these metrics and stages could help to decide which of the generated dialogues should be maintained depending on the applied policy, which could vary from preserving between 57% to 93% of the initial dialogues. Cristina Luna Jiménez, Manuel Gil-Martín, Luis Fernando D'Haro, Fernando Fernández Martínez, Rubén San-Segundo-Hernández |
Expert Syst. Appl. | 3 |
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 26 |
| 2023 | PoE: A Panel of Experts for Generalized Automatic Dialogue AssessmentabstractChatbots are expected to be knowledgeable across multiple domains, e.g. for daily chit-chat, exchange of information, and grounding in emotional situations. To effectively measure the quality of such conversational agents, a model-based automatic dialogue evaluation metric (ADEM) is expected to perform well across multiple domains. Despite significant progress, existing ADEMs tend to perform well only on data that are similar to its training data (overfit to its training domain). This calls for a domain-generalized metric that can assess dialogues of different characteristics. To this end, we propose aPanel of Experts(PoE), a multitask network that consists of a shared transformer encoder and a collection of lightweight adapters. The shared encoder captures the general knowledge of dialogues across domains, while each adapter specializes in one specific domain and serves as a domain expert. To validate the idea, we construct a high-quality multi-domain dialogue dataset leveraging data augmentation and pseudo-labeling. The PoE network is comprehensively assessed on 16 dialogue evaluation datasets spanning a wide range of dialogue domains. It achieves state-of-the-art performance in terms of mean Spearman correlation over all the evaluation datasets. It exhibits better zero-shot generalization than existing state-of-the-art ADEMs and the ability to easily adapt to new domains with few-shot transfer learning. Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationabstractChatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure the quality of such conversational agents, a dialogue evaluator is expected to conduct assessment across domains as well. However, most of the state-of-the-art automatic dialogue evaluation metrics (ADMs) are not designed for multi-domain evaluation. We are motivated to design a general and robust framework, MDD-Eval, to address the problem. Specifically, we first train a teacher evaluator with human-annotated data to acquire a rating skill to tell good dialogue responses from bad ones in a particular domain and then, adopt a self-training strategy to train a new evaluator with teacher-annotated multi-domain data, that helps the new evaluator to generalize across multiple domains. MDD-Eval is extensively assessed on six dialogue evaluation benchmarks. Empirical results show that the MDD-Eval framework achieves a strong performance with an absolute improvement of 7% over the state-of-the-art ADMs in terms of mean Spearman correlation scores across all the evaluation benchmarks. Chen Zhang 0055, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou Li 0001 |
AAAI | 2 |
| 2022 | FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationabstractRecent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment 1 .However, they either perform turn-level evaluation or look at a single dialogue quality dimension.One would expect a good evaluation metric to assess multiple quality dimensions at the dialogue level.To this end, we are motivated to propose a multi-dimensional dialogue-level metric, which consists of three sub-metrics with each targeting a specific dimension.The submetrics are trained with novel self-supervised objectives and exhibit strong correlations with human judgment for their respective dimensions.Moreover, we explore two approaches to combine the sub-metrics: metric ensemble and multitask learning.Both approaches yield a holistic metric that significantly outperforms individual sub-metrics.Compared to the existing state-of-the-art metric, the combined metrics achieve around 16% relative improvement on average across three high-quality dialoguelevel evaluation benchmarks. Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001 |
EMNLP | 2 |
| 2022 | Phonotactic Language Recognition Using A Universal Phoneme Recognizer and A Transformer ArchitectureabstractIn this paper, we describe a phonotactic language recognition model that effectively manages long and short n-gram input sequences to learn contextual phonotactic-based vector embeddings. Our approach uses a transformer-based encoder that integrates a sliding window attention to attempt finding discriminative short and long cooccurrences of language dependent n-gram phonetic units. We then evaluate and compare the use of different phoneme recognizers (Brno and Allosaurus) and sub-unit tokenizers to help select the more discriminative n-grams. The proposed architecture is evaluated using the Kalaka-3 database that contains clean and noisy audio recordings for very similar languages (i.e. Iberian languages, e.g., Spanish, Galician, Catalan). We provide results using the Cavg and accuracy metrics used in NIST evaluations. The experimental results show that our proposed approach outperforms by 21% of relative improvement to the best system presented in the Albayzin LR competition. Luis Fernando D'Haro, Marcos Estecha-Garitagoitia, Christian Salamea Palacios |
ICASSP | 2 |
| 2021 | DynaEval: Unifying Turn and Dialogue Level EvaluationabstractChen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Zhang 0055, Yiming Chen 0010, Luis Fernando D'Haro, Yan Zhang 0004, Thomas Friedrichs, Grandee Lee, Haizhou Li 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Editorial: Special Issue on the Eighth Dialog System Technology Challenge
Seokhwan Kim, Hannes Schulz, R. Chulaka Gunasekara, Chiori Hori, Abhinav Rastogi, Luis Fernando D'Haro |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | D-Score: Holistic Dialogue Evaluation Without ReferenceabstractIn artistic gymnastics, difficulty score or D-score is used for judging performance. Starting from zero, an athlete earns points from different aspects such as composition requirement, difficulty, and connection between moves. The final score is a composition of the quality of various performance indicators. Similarly, when evaluating dialogue responses, human judges generally follow a number of criteria, among which language fluency, context coherence, logical consistency, and semantic appropriateness are on top of the agenda. In this paper, we propose an automatic dialogue evaluation framework called D-score that resembles the way gymnastics is evaluated. Following the four human judging criteria above, we devise a range of evaluation tasks and model them under a multi-task learning framework. The proposed framework, without relying on any human-written reference, learns to appreciate the overall quality of human-human conversations through a representation that is shared by all tasks without over-fitting to individual task domain. We evaluate D-score by performing comprehensive correlation analyses with human judgement on three dialogue evaluation datasets, among which two are from past DSTC series, and benchmark against state-of-the-art baselines. D-score not only outperforms the best baseline by a large margin in terms of system-level Spearman correlation but also represents an important step towards explainable dialogue scoring. Chen Zhang 0055, Grandee Lee, Luis Fernando D'Haro, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Overview of the seventh Dialog System Technology Challenge: DSTC7
Luis Fernando D'Haro, Koichiro Yoshino, Chiori Hori, Tim K. Marks, Lazaros Polymenakos, Jonathan K. Kummerfeld, Michel Galley, Xiang Gao 0011 |
Comput. Speech Lang. | 1 |
| 2019 | Joint Learning of Word and Label Embeddings for Sequence Labelling in Spoken Language UnderstandingabstractWe propose an architecture to jointly learn word and label embeddings for slot filling in spoken language understanding. The proposed approach encodes labels using a combination of word embeddings and straightforward word-label association from the training data. Compared to the state-of-the-art methods, our approach does not require label embed-dings as part of the input and therefore lends itself nicely to a wide range of model architectures. In addition, our architecture computes contextual distances between words and labels to avoid adding contextual windows, thus reducing memory footprint. We validate the approach on established spoken dialogue datasets and show that it can achieve state-of-the-art performance with much fewer trainable parameters. Jiewen Wu, Luis Fernando D'Haro, Nancy F. Chen, Pavitra Krishnaswamy, Rafael E. Banchs |
ASRU | 2 |
| 2019 | Automatic evaluation of end-to-end dialog systems with adequacy-fluency metrics
Luis Fernando D'Haro, Rafael E. Banchs, Chiori Hori, Haizhou Li 0001 |
Comput. Speech Lang. | 1 |
| 2018 | Multi-Modal Robot Apprenticeship: Imitation Learning Using Linearly Decayed DMP+ in a Human-Robot Dialogue SystemabstractRobot learning by demonstration gives robots the ability to learn tasks which they have not been programmed to do before. The paradigm allows robots to work in a greater range of real-world applications in our daily life. However, this paradigm has traditionally been applied to learn tasks from a single demonstration modality. This restricts the approach to be scaled to learn and execute a series of tasks in a real-life environment. In this paper, we propose a multi-modal learning approach using DMP+ with linear decay integrated in a dialogue system with speech and ontology for the robot to learn seamlessly through natural interaction modalities (like an apprentice) while learning or re-learning is done on the fly to allow partial updates to a learned task to reduce potential user fatigue and operational downtime in teaching. The performance of new DMP+ with linear decay system is statistically benchmarked against state-of-the-art DMP implementations. A gluing demonstration is also conducted to show how the system provides seamless learning of multiple tasks in a flexible manufacturing set-up. Yan Wu 0002, Luis Fernando D'Haro, Rafael E. Banchs, Keng Peng Tee |
IROS | 3 |
| 2017 | Evaluation of the Neurological State of People with Parkinson's Disease Using i-Vectors
Nicanor García, Juan Rafael Orozco-Arroyave, Luis Fernando D'Haro, Najim Dehak, Elmar Nöth |
INTERSPEECH | 3 |
| 2016 | A Multimodal Control Architecture for Autonomous Unmanned Aerial VehiclesabstractWe present our preliminary work on a multimodal control architecture that enables an operator to manage an autonomous Unmanned Aerial Vehicle (UAV) through high level tasks in an indoors environment. The intelligence embedded in our architecture is able to decode these tasks into low level instructions that a UAV is able to execute. Our system allows the user to operate the UAV through speech, text or keyboard/mouse input, all presented in a web based graphical user interface that can be accessed from any Internet powered device. Marco A. Gutiérrez 0002, Luis Fernando D'Haro, Rafael E. Banchs |
HAI | 2 |
| 2016 | A Web-based Platform for Collection of Human-Chatbot InteractionsabstractOver recent years, the world has seen multiple uses for conversational agents. Chatbots has been implemented into ecommerce systems, such as Amazon Echo's Alexa [1]. Businesses and organizations like Facebook are also implementing bots into their applications. While a number of amazing chatbot platform exists, there are still difficulties in creating data-driven-systems as they large amount of data is needed for development and training. This paper we describe an advanced platform for evaluating and annotating human-chatbot interactions, its main features and goals, as well as the future plans we have for it. Lue Lin, Luis Fernando D'Haro, Rafael E. Banchs |
HAI | 2 |
| 2016 | Automatic Correction of ASR Outputs by Using Machine Translation
Luis Fernando D'Haro, Rafael E. Banchs |
INTERSPEECH | 1 |
| 2016 | The fifth dialog state tracking challengeabstractDialog state tracking - the process of updating the dialog state after each interaction with the user - is a key component of most dialog systems. Following a similar scheme to the fourth dialog state tracking challenge, this edition again focused on human-human dialogs, but introduced the task of cross-lingual adaptation of trackers. The challenge received a total of 32 entries from 9 research groups. In addition, several pilot track evaluations were also proposed receiving a total of 16 entries from 4 groups. In both cases, the results show that most of the groups were able to outperform the provided baselines for each task. Seokhwan Kim, Luis Fernando D'Haro, Rafael E. Banchs, Jason D. Williams, Matthew Henderson, Koichiro Yoshino |
SLT | 2 |
| 2016 | Frequency features and GMM-UBM approach for gait-based person identification using smartphone inertial signals
Rubén San-Segundo-Hernández, Ricardo de Córdoba, Javier Ferreiros, Luis Fernando D'Haro |
Pattern Recognit. Lett. | 4 |
| 2015 | Automatic fill-the-blank question generator for student self-assessmentabstractToday's educational systems require students to recall and apply major concepts from study material to perform competently in assessments. Crucial to achieving this is practice and self-assessment through questions. The crafting of such questions can be time consuming for teachers while questions from external sources, e.g. assessment books, might not be tailored to suit students' study materials. As such, we present RevUP: an automated system for Gap-Fill Question Generation (GFQG) from educational texts. Example GFQs generated by RevUP from a high-school biology textbook are as follows. · Endocrine signals generated by the hypothalamus regulate hormone secretion by the _____. (a) posterior pituitary (b) anterior pituitary (c) thyroid gland (d) pineal gland Our system factors the problem of generating good questions into 3 parts: 1) Selecting topically important sentences to ask about 2) Identifying which part of the resulting sentence to choose as the gap 3) Crafting effective distractors to be part of the set of options to confuse the learner The immediate advantage of this service is that it provides the tools that make it easy for teachers to quickly generate and edit questions for pop quizzes and worksheets from their lecture notes. In the paper, we will provide a detailed description of the methodology used to generate questions. We will also provide preliminary usability statistics from high school students where they qualified the usability and the negative aspects of the application. Rafael E. Banchs, Luis Fernando D'Haro |
FIE | 3 |
| 2015 | Conversational agent and management tools for conference and tourism domain
Luis Fernando D'Haro, Seokhwan Kim, Rafael E. Banchs |
INTERSPEECH | 1 |
| 2015 | Adequacy-Fluency Metrics: Evaluating MT in the Continuous Space Model FrameworkabstractThis work extends and evaluates a two-dimensional automatic evaluation metric for machine translation, which is designed to operate at the sentence level. The metric is based on the concepts of adequacy and fluency, aiming at decoupling both semantic and syntactic components of the translation process to provide a more balanced view on translation quality. These two elements are independently evaluated by using continuous space and$n$-gram language modeling frameworks, respectively. Two different implementations are evaluated: a monolingual version that fully operates on the target language side, and a cross-language version that has the main advantage of not requiring reference translations. Both implementations are evaluated by comparing their performance with state-of-the-art automatic metrics over a dataset involving five different European languages. Rafael E. Banchs, Luis Fernando D'Haro, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | A web-based application for the management and evaluation of tutoring requests in PBL-based massive laboratoriesabstractOne important steps in a successful project-based-learning methodology (PBL) is the process of providing the students with a convenient feedback that allows them to keep on developing their projects or to improve them. However, this task is more difficult in massive courses, especially when the project deadline is close. Besides, the continuous evaluation methodology makes necessary to find ways to objectively and continuously measure students' performance without increasing excessively instructors' work load. In order to alleviate these problems, we have developed a web service that allows students to request personal tutoring assistance during the laboratory sessions by specifying the kind of problem they have and the person who could help them to solve it. This service provides tools for the staff to manage the laboratory, for performing continuous evaluation for all students and for the student collaborators, and to prioritize tutoring according to the progress of the student's project. Additionally, the application provides objective metrics which can be used at the end of the subject during the evaluation process in order to support some students' final scores. Different usability statistics and the results of a subjective evaluation with more than 330 students confirm the success of the proposed application. Luis Fernando D'Haro, Fernando Fernández Martínez, Ricardo de Córdoba, Juan Manuel Montero-Martínez |
FIE | 1 |
| 2014 | Extended phone log-likelihood ratio features and acoustic-based i-vectors for language recognitionabstractThis paper presents new techniques with relevant improvements added to the primary system presented by our group to the Albayzin 2012 LRE competition, where the use of any additional corpora for training or optimizing the models was forbidden. In this work, we present the incorporation of an additional phonotactic subsystem based on the use of phone log-likelihood ratio features (PLLR) extracted from different phonotactic recognizers that contributes to improve the accuracy of the system in a 21.4% in terms of Cavg(we also present results for the official metric during the evaluation, Fact). We will present how using these features at the phone state level provides significant improvements, when used together with dimensionality reduction techniques, especially PCA. We have also experimented with applying alternative SDC-like configurations on these PLLR features with additional improvements. Also, we will describe some modifications to the MFCC-based acoustic i-vector system which have also contributed to additional improvements. The final fused system outperformed the baseline in 27.4% in Cavg. Luis Fernando D'Haro, Ricardo de Córdoba, Christian Salamea Palacios, Julián D. Echeverry-Correa |
ICASSP | 1 |
| 2014 | Language recognition using phonotactic-based shifted delta coefficients and multiple phone recognizersabstractA new language recognition technique based on the application of the philosophy of the Shifted Delta Coefficients (SDC) to phone log-likelihood ratio features (PLLR) is described. The new methodology allows the incorporation of long-span phonetic information at a frame-by-frame level while dealing with the temporal length of each phone unit. The proposed features are used to train an i-vector based system and tested on the Albayzin LRE 2012 dataset. The results show a relative improvement of 33.3% in Cavg in comparison with different state-of-the-art acoustic i-vector based systems. On the other hand, the integration of parallel phone ASR systems where each one is used to generate multiple PLLR coefficients which are stacked together and then projected into a reduced dimension are also presented. Finally, the paper shows how the incorporation of state information from the phone ASR contributes to provide additional improvements and how the fusion with the other acoustic and phonotactic systems provides an important improvement of 25.8% over the system presented during the competition. Luis Fernando D'Haro, Ricardo de Córdoba, Christian Salamea Palacios, Javier Ferreiros |
INTERSPEECH | 1 |
| 2013 | Low-resource language recognition using a fusion of phoneme posteriorgram counts, acoustic and glottal-based i-vectorsabstractThis paper presents a description of our system for the Albayzin 2012 LRE competition. One of the main characteristics of this evaluation was the reduced number of available files for training the system, especially for the empty condition where no training data set was provided but only a development set. In addition, the whole database was created from online videos and around one third of the training data was labeled as noisy files. Our primary system was the fusion of three different i-vector based systems: one acoustic system based on MFCCs, a phonotactic system using trigrams of phone-posteriorgram counts, and another acoustic system based on RPLPs that improved robustness against noise. A contrastive system that included new features based on the glottal source was also presented. Official and post-evaluation results for all the conditions using the proposed metrics for the evaluation and the Cavgmetric are presented in the paper. Luis Fernando D'Haro, Ricardo de Córdoba, Miguel Ánguel Caraballo, José Manuel Pardo |
ICASSP | 1 |
| 2012 | Phonotactic Language Recognition using i-vectors and Phoneme Posteriogram Counts
Luis Fernando D'Haro, Ondrej Glembek, Oldrich Plchot, Pavel Matejka, Mehdi Soufifar, Ricardo de Córdoba, Jan Cernocký |
INTERSPEECH | 1 |
| 2012 | Patrol Team Language Identification System for DARPA RATS P1 EvaluationabstractThis paper describes the language identification (LID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We show that techniques originally developed for LID on telephone speech (e.g., for the NIST language recognition evaluations) remain effective on the noisy RATS data, provided that careful consideration is applied when designing the training and development sets. In addition, we show significant improvements from the use of Wiener filtering, neural network based and language dependent i-vector modeling, and fusion. Index Terms: language identification, noisy speech. 1. Pavel Matejka, Oldrich Plchot, Mehdi Soufifar, Ondrej Glembek, Luis Fernando D'Haro, Karel Veselý, Frantisek Grézl, Jeff Z. Ma, Spyridon Matsoukas, Najim Dehak |
INTERSPEECH | 5 |
| 2012 | Application of backend database contents and structure to the design of spoken dialog services
Luis Fernando D'Haro, Ricardo de Córdoba, Juan Manuel Montero-Martínez, Javier Ferreiros, José Manuel Pardo |
Expert Syst. Appl. | 1 |
| 2012 | Design, development and field evaluation of a Spanish into sign language translation system
Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Ricardo de Córdoba, Valentín Sama Rojo, Fernando Fernández Martínez, Luis Fernando D'Haro, Verónica López-Ludeña, D. Sánchez |
Pattern Anal. Appl. | 6 |
| 2011 | Design and evaluation of acceleration strategies for speeding up the development of dialog applications
Luis Fernando D'Haro, Ricardo de Córdoba, Rubén San-Segundo-Hernández, Javier Ferreiros, José Manuel Pardo |
Speech Commun. | 1 |
| 2009 | Strategies for accelerating the design of dialogue applications using heuristic information from the backend databaseabstractNowadays, current commercial and academic platforms for developing spoken dialogue applications lack of acceleration strategies based on using heuristic information from the contents or structure of the backend database in order to speed up the definition of the dialogue flow.In this paper we describe our attempts to take advantage of these information sources using the following strategies: the quick creation of classes and attributes to define the data model structure, the semi-automatic generation and debugging of database access functions, the automatic proposal of the slots that should be preferably requested using mixed-initiative forms or the slots that are better to request one by one using directed forms, and the generation of automatic state proposals to specify the transition network that defines the dialogue flow.Subjective and objective evaluations confirm the advantages of using the proposed strategies to simplify the design, and the high acceptance of the platform and its acceleration strategies. Luis Fernando D'Haro, Ricardo de Córdoba, Rubén San-Segundo-Hernández, Javier Macías Guarasa, José Manuel Pardo |
INTERSPEECH | 1 |
| 2009 | Speeding Up the Design of Dialogue Applications by Using Database Contents and Structure Information
Luis Fernando D'Haro, Ricardo de Córdoba, Juan Manuel Lucas, Roberto Barra-Chicote, Rubén San-Segundo-Hernández |
SIGDIAL Conference | 1 |
| 2008 | Language model adaptation for a speech to sign language translation system using web frequencies and a MAP frameworkabstractThis paper presents a successful technique for creating a new language model (LM) that adapts the original target LM used by a machine translation (MT) system.This technique is especially useful for situations where there are very scarce resources for training the target side (Spanish Sign Language (LSE) in our case) in order to properly estimate the target LM, the Sign Language Model (SLM), used by the MT system.The technique uses information from the source language, Spanish in our task, and from the phrase-based translation matrix in order to create a new LM, estimated using web frequencies, which adapts the counts of the SLM through the Maximum A Posteriori method (MAP).The corpus consists of common used sentences spoken by an officer when assisting people in applying for, or renewing, the National Identification Document.The proposed technique allows relative reductions of 15.5% on perplexity and 2.7% on WER for translation, which are close to half the maximum performance obtainable when only the LM is optimized. Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba, Jan Bungeroth, Daniel Stein, Hermann Ney |
INTERSPEECH | 1 |
| 2008 | Speech to sign language translation system for Spanish
Rubén San-Segundo-Hernández, Roberto Barra-Chicote, Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Javier Ferreiros, Juan Manuel Lucas, Javier Macías Guarasa, Juan Manuel Montero-Martínez, José Manuel Pardo |
Speech Commun. | 4 |
| 2007 | A Multimodal Interface for Access to Content in the Home
Michael Johnston, Luis Fernando D'Haro, Michelle Levine, Bernard Renger |
ACL | 2 |
| 2007 | Language identification based on n-gram frequency rankingabstractWe present a novel approach for language identification based on a text categorization technique, namely an n-gram frequency ranking. We use a Parallel phone recognizer, the same as in PPRLM, but instead of the language model, we create a ranking with the most frequent n-grams, keeping only a fraction of them. Then we compute the distance between the input sentence ranking and each language ranking, based on the difference in relative positions for each n-gram. The objective of this ranking is to be able to model reliably a longer span than PPRLM, namely 5-gram instead of trigram, because this ranking will need less training data for a reliable estimation. We demonstrate that this approach overcomes PPRLM (6 % relative improvement) due to the inclusion of 4gram and 5-gram in the classifier. We present two alternatives: ranking with absolute values for the number of occurrences and ranking with discriminative values (11% relative improvement). Index Terms: Language Identification, n-gram frequency ranking, text categorization, PPRLM Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Javier Macías Guarasa, Javier Ferreiros |
INTERSPEECH | 2 |
| 2007 | Language identification using several sources of information with a multiple-Gaussian classifierabstractWe present several innovative techniques that can be applied in a PPRLM system for language identification (LID). To normalize the scores, eliminate the bias in the scores and improve the classifier, we compared the bias removal technique (up to 19 % relative improvement (RI)) and a Gaussian classifier (up to 37 % RI). Then, we include additional sources of information in different feature vectors of the Gaussian classifier: the sentence acoustic score (11% RI), the average acoustic score for each phoneme (11 % RI), and the average duration for each phoneme (7.8 % RI). The use of a multiple-Gaussian classifier with 4 feature vectors meant an additional 15.1 % RI. Using 4 feature vectors instead of just PPRLM provides a 26.1 % RI. Finally, we include additional acoustic HMMs of the same language with success (10 % relative improvement). We will show how all these improvements have been mostly additive. Ricardo de Córdoba, Luis Fernando D'Haro, Fernando Fernández Martínez, Juan Manuel Montero-Martínez, Roberto Barra-Chicote |
INTERSPEECH | 2 |
| 2007 | Evaluation of alternatives on speech to sign language translationabstractThis paper evaluates different approaches on speech to sign language machine translation. The framework of the application focuses on assisting deaf people to apply for the passport or related information. In this context, the main aim is to automatically translate the spontaneous speech, uttered by an officer, into Spanish Sign Language (SSL). In order to get the best translation quality, three alternative techniques have been evaluated: a rule-based approach, a phrase-based statistical approach, and a approach that makes use of stochastic finite state transducers. The best speech translation experiments have reported a 32.0 % SER (Sign Error Rate) and a 7.1 BLEU (BiLingual Evaluation Understudy) including speech recognition errors. Rubén San-Segundo-Hernández, Alicia Pérez, Daniel Ortiz-Martínez, Luis Fernando D'Haro, M. Inés Torres, Francisco Casacuberta |
INTERSPEECH | 4 |
| 2006 | Prosodic and Segmental Rubrics in Emotion IdentificationabstractIt is well known that the emotional state of a speaker usually alters the way she/he speaks. Although all the components of the voice can be affected by emotion in some statistically-significant way, not all these deviations from a neutral voice are identified by human listeners as conveying emotional information. In this paper we have carried out several perceptual and objective experiments that show the relevance of prosody and segmental spectrum in the characterization and identification of four emotions in Spanish. A Bayes classifier has been used in the objective emotion identification task. Emotion models were generated as the contribution of every emotion to the build-up of a universal background emotion codebook. According to our experiments, surprise is primarily identified by humans through its prosodic rubric (in spite of some automatically-identifiable segmental characteristics); while for anger the situation is just the opposite. Sadness and happiness need a combination of prosodic and segmental rubrics to be reliably identified Roberto Barra-Chicote, Juan Manuel Montero-Martínez, Javier Macías Guarasa, Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba |
ICASSP (1) | 4 |
| 2006 | A Spanish speech to sign language translation system for assisting deaf-mute peopleabstractThis paper describes the first experiments of a speech to sign language translation system in a real domain. The developed system is focused on the sentences spoken by an officer when assisting people in applying for, or renewing the National Identification Document (NID) and the Passport. This system translates officer explanations into sign language for deafmute people. The translation system is composed by a speech recognizer (for decoding the spoken utterance into a word sequence), a natural language translator (for converting a word sequence into a sequence of gestures belonging to the sign language), and a 3D avatar animation module (for playing the gestures). The field experiments have reported a 27.2 % GER (Gesture Error Rate) and a 0.62 BLEU Rubén San-Segundo-Hernández, Roberto Barra-Chicote, Luis Fernando D'Haro, Juan Manuel Montero-Martínez, Ricardo de Córdoba, Javier Ferreiros |
INTERSPEECH | 3 |
| 2006 | Error Analysis of Statistical Machine Translation Output
David Vilar, Jia Xu 0004, Luis Fernando D'Haro, Hermann Ney |
LREC | 3 |
| 2006 | An advanced platform to speed up the design of multilingual dialog applications for multiple modalities
Luis Fernando D'Haro, Ricardo de Córdoba, Javier Ferreiros, Stefan W. Hamerich, Volker Schless, Basilis Kladis, Volker Schubert, Otilia Kocsis, Stefan Igel, José Manuel Pardo |
Speech Commun. | 1 |
| 2005 | New word-level and sentence-level confidence scoring using graph theory calculus and its evaluation on speech understandingabstractA lot of work has been devoted to the estimation of confidence measures for speech recognizers. In the quite extended case where a word-graph speech recognizer is in use, we will present new confidence measures employing the graph theory that shows us how to estimate some interesting characteristics about the different paths through the graph that constitute the recognition solutions, without the need of expanding them all. We will take advantage of some of these features to generate confidence scores both at the word and sentence level. We will also compare this new confidence scoring to more traditional ones and will find similar behavior with less computational load and with an increase in the simplicity of the approach that will lead to more generalization power of the confidence estimation to different applications of the recognizer. Javier Ferreiros, Rubén San-Segundo-Hernández, Fernando Fernández Martínez, Luis Fernando D'Haro, Valentín Sama Rojo, Roberto Barra-Chicote, Pedro Mellén |
INTERSPEECH | 4 |
| 2004 | Language identification techniques based on full recognition in an air traffic control taskabstractAutomatic language identification has become an important issue in recent years in speech recognition systems.In this paper, we present the work done in language identification for an air traffic control speech recognizer for continuous speech.The system is able to distinguish between Spanish and English.We present several language identification techniques based on full recognition that improve the baseline results obtained using the most commonly known "PPRLM" technique.We have in our database some task specific critical problems for language identification like non native speakers, extremely spontaneous speech or Spanish-English mix in the same sentence.We confirm that PPRLM is quite sensible to those problems and that a technique based on a Bayesian classifier is the one with the best performance in spite of its higher computational cost. Ricardo de Córdoba, Javier Ferreiros, Valentín Sama Rojo, Javier Macías Guarasa, Luis Fernando D'Haro, Fernando Fernández Martínez |
INTERSPEECH | 5 |
| 2004 | Strategies to reduce design time in multimodal/multilingual dialog applicationsabstractIn this paper, we present a complete platform for the semiautomatic generation of human-machine dialog systems, that using as input a description of the database of the service, a flow model with the different states of the final application and a guided interaction step by step with the designer’s intervention, generates dialogs to access the service data in different languages and two modalities, speech and web, simultaneously. We describe in detail several strategies that have been followed to reduce the time needed to do the design using the mentioned information. We also address important issues in dialog applications as mixed initiative and overanswering dialogs, confirmation handling and how to provide the user long lists of information. Luis Fernando D'Haro, Ricardo de Córdoba, Rubén San-Segundo-Hernández, Juan Manuel Montero-Martínez, Javier Macías Guarasa, José Manuel Pardo |
INTERSPEECH | 1 |
| 2004 | Implementation of dialog applications in an open-source voiceXML platformabstractIn this paper, we study the approach followed to use the VoiceXML standard in a dialog system platform already available in our group. As VoiceXML interpreter we have chosen OpenVXI, an open source portable solution where we can make the modifications needed to adapt the solution to the characteristics of our recognition and synthesis modules; so we will emphasize the changes that we have had to make in such interpreter. Besides, we review some relevant modules in our platform and their capabilities, highlighting the use of standards in them, as SSML for the text-to-speech system and JSGF for the specification of grammars for recognition. Finally, we discuss several ideas regarding the limitations detected in VoiceXML. Fernando Fernández Martínez, Valentín Sama Rojo, Luis Fernando D'Haro, Rubén San-Segundo-Hernández, Ricardo de Córdoba, Juan Manuel Montero-Martínez |
INTERSPEECH | 3 |
| 2004 | The GEMINI platform: semi-automatic generation of dialogue applicationsabstractThe EC funded research project GEMINI (Generic Environment for Multilingual Interactive Natural Interfaces) has two main objectives: On the one hand the development and implementation of a platform able to produce user-friendly interactive multilingual and multi-modal dialogue interfaces to databases with a minimum of human effort and on the other hand the demonstration of the platform’s efficiency through the development of two different applications using this platform. The platform consists of different assistants that help the user to semi-automatically generate dialogue applications. Its open and modular architecture simplifies the adaptability of generated applications to different use cases. 1. Stefan W. Hamerich, Volker Schless, Basilis Kladis, Volker Schubert, Otilia Kocsis, Stefan Igel, Ricardo de Córdoba, Luis Fernando D'Haro, José Manuel Pardo |
INTERSPEECH | 8 |
| 2003 | Revisiting scenarios and methods for variable frame rate analysis in automatic speech recognitionabstractIn this paper we present a revision and evaluation of some of the main methods used in variable frame rate (VFR) analysis, applied to speech recognition systems. The work found in the literature in this area usually deals with restricted conditions and scenarios and we have revisited the main algorithmic alternatives and evaluated them under the same experimental framework, so that we have been able to establish objective considerations for each of them, selecting the most adequate strategy. We also show till what extent VFR analysis is useful in its three main application scenarios, namely “reduction of computational load”, “improve acoustic modelling ” and “handling additive noise conditions in the time domain”. From our evaluation on a difficult telephone large vocabulary task, we establish that VFR analysis does not significantly improve the results obtained using the traditional fixed frame rate analysis (FFR), except when additive noise is present in the database and specially for low SNRs. 1. Javier Macías Guarasa, J. Ordóñez, Juan Manuel Montero-Martínez, Javier Ferreiros, Ricardo de Córdoba, Luis Fernando D'Haro |
INTERSPEECH | 6 |